An intelligent decision-making method and architecture for a UAV communication system

Through the layered intelligent decision-making method, the resource management problems of the UAV communication system are decoupled, which reduces the computational complexity, adapts to the dynamic environment, and improves the strategic performance of the UAV communication system.

CN115580885BActive Publication Date: 2025-07-22UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211168943.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-25
Publication Date
2025-07-22
Estimated Expiration
2042-09-25

AI Technical Summary

Technical Problem

The resource management architecture of traditional drone communication systems requires accurate environmental status prior information, high computational complexity, and deteriorate decision performance when environmental changes, making it difficult to adapt to large-scale drone communication systems.

Method used

Hierarchical intelligent decision-making methods are adopted, including decision analysis layer, task planning layer, algorithm design, collaborative control layer and task execution layer. Through the interaction between intelligent algorithms and the environment, decoupling joint optimization problems as sub-optimization problems, design corresponding intelligent algorithms and map them into physical behaviors.

Benefits of technology

It reduces the computational complexity, is suitable for dynamic environments, improves policy performance, does not require specific environmental parameters, and is suitable for large-scale drone communication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115580885B_ABST
    Figure CN115580885B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of wireless communication technologies, and particularly relates to an intelligent decision-making method and architecture for a drone communication system. The present invention provides a hierarchical intelligent decision-making method and architecture, including a decision analysis layer, a task planning layer, an algorithm design and cooperative control (ADCC) layer, and a task execution layer; in this method, the resource management problem to be solved is to establish a joint optimization problem by the decision analysis layer, the task planning layer transforms the complex joint optimization problem into several sub-optimization problems that can be solved by existing intelligent algorithms, the ADCC layer designs specific intelligent algorithms to provide solution strategies for these sub-optimization problems, and finally, the strategies are transformed into physical actions by the task execution layer to interact with the environment; the intelligent decision-making architecture of the drone communication system proposed by the present invention solves problems through intelligent algorithms, so the computational complexity is relatively low, and specific environmental parameters are not required. Only through continuous interaction with the environment can the strategy performance be continuously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to an intelligent decision-making method and architecture for a drone communication system. Background Art

[0002] In recent years, with the increasing number of wireless communication users and the demand for ultra-high reliability and high quality of service (QoS) in wireless communication, these have posed challenges to traditional fixed terrestrial communication networks. Therefore, as a means to address this challenge, drone communication has attracted extensive attention worldwide. Compared with terrestrial communication networks, drone communication is much less affected by small-scale fading and shadow fading, has better line-of-sight (LoS) characteristics, and the mobility of drones enables them to provide real-time communication services for users in emergency situations, with more flexible networking. Therefore, the cost of enhancing the required links is lower, and efficient, reliable, and economical communication services can be provided for users.

[0003] In a drone communication system, wireless communication resources are extremely precious. Therefore, resource management is of great significance for drone communication. For example, trajectory design can make full use of the high mobility of drones, dispatch the drones above the users who need communication, and effectively utilize the LoS link characteristics of drone communication to reduce path loss and provide efficient and reliable wireless communication services for users; user association technology enables users to select suitable drones, reduce path loss, and access the wireless network without overloading the drones; resource block allocation technology can effectively and fully utilize these time-frequency resource blocks, occupy channels with good conditions, and avoid the impact of malicious interference on transmission performance; power control enables drones to successfully transmit information with appropriate power, improve the endurance of drone communication, and at the same time avoid interfering with other users. Ground communication facilities can obtain stable and sufficient energy supply from the power grid, while drones can only be powered by the fuel or batteries they carry. Moreover, in addition to communication energy consumption, drones also require additional energy to generate thrust to maintain flight. Therefore, drone communication has more stringent requirements for power control.

[0004] Resource allocation in UAV communication needs to consider many factors, and it is difficult to solve this problem. With the increasing demand for air communication, many organizations have conducted research on resource management technologies in UAV communication. The traditional UAV resource management architecture requires accurate and complete prior information about the environmental state, which is difficult to obtain in actual scenarios. In addition, the traditional architecture needs to establish a convex optimization problem, while most resource management problems are non-convex. Moreover, the computational complexity of the traditional architecture increases with the high-order terms of the number of users and UAVs, making it unsuitable for resource management in large-scale UAV communication systems. Furthermore, the traditional resource management architecture plans strategies in advance based on existing information, and when the environment changes, the decision-making performance will be greatly degraded. Summary of the Invention

[0005] To solve the deficiencies of the prior art and address the resource management problem in UAV communication systems, the present invention provides a hierarchical intelligent decision-making method and architecture, including a decision analysis layer, a task planning layer, an Algorithm Design and Cooperative Control (ADCC) layer, and a task execution layer. In this method, the resource management problem to be solved is formulated as a joint optimization problem by the decision analysis layer. The task planning layer transforms the complex joint optimization problem into several sub-optimization problems that can be solved by existing intelligent algorithms. The ADCC layer designs specific intelligent algorithms to provide solution strategies for these sub-optimization problems. Finally, the task execution layer transforms the strategies into physical actions to interact with the environment. The intelligent decision-making architecture of the UAV communication system proposed by the present invention solves problems through intelligent algorithms, so it has a low computational complexity and does not require specific environmental parameters. It can continuously improve the strategy performance only through continuous interaction with the environment.

[0006] To facilitate the description of the content of the present invention, in the UAV communication system described in the present invention, a single UAV can act as a communication relay to provide a service for forwarding information between a transceiver pair with a long distance or lacking a line-of-sight path. Multiple UAVs can form a UAV cluster base station, and the UAVs in the cluster can communicate reliably to ensure the cooperation between UAVs. On the premise of cooperation, the UAV cluster acts as an air base station to provide communication services for ground users.

[0007] The technical solution of the present invention is as follows:

[0008] An intelligent decision-making method for a UAV communication system, where the UAV communication system uses a UAV cluster base station composed of multiple UAVs to provide communication services for ground users, and a single UAV is used as a communication relay or base station, and the intelligent decision-making method synchronizes the cognitive process and the decision-making process;

[0009] The specific steps of the cognitive function are as follows:

[0010] R1. The task execution layer perceives environmental information and transmits it to the ADCC layer and the decision-making unit of this layer;

[0011] R2. The ADCC layer obtains the information of the UAVs performing tasks, including UAV coordinates, flight speed, flight direction, etc., and transmits the UAV information and environmental information to the task planning layer and the decision-making unit of this layer together;

[0012] R3. The task planning layer calculates the dynamic reachable performance of the system based on the environmental information and UAV information, and transmits the calculated performance, environmental information and UAV information to the decision analysis layer and the decision-making unit of this layer together;

[0013] R4. The decision analysis layer analyzes the current situation based on the environmental information, UAV information and reachable performance, judges whether the current decision-making performance meets the expectations, and if not, how to improve it.

[0014] The specific steps of the decision-making function are as follows:

[0015] S1. The decision analysis layer formulates corresponding optimization problems and the constraint conditions to be satisfied according to the tasks to be completed;

[0016] S2. The task planning layer plans a reasonable task solution according to the formulated optimization problems, and decouples the joint optimization problems into sub-problems that can be solved by specific algorithms;

[0017] S3. For the disassembled sub-problems, on the basis of UAV cooperation, specific algorithms for solving these sub-problems are designed, and strategies are output under the current environmental state and UAV state;

[0018] S4. Map the strategies of the UAVs into physical behaviors and then interact with the environment.

[0019] Furthermore, the specific steps of formulating the optimization problems and their constraint conditions in S1 are as follows:

[0020] S11. Describe the current communication scenario, and judge whether the UAV communication mode is relay mode or base station mode; if it is relay mode, go to S131; otherwise, go to S12;

[0021] S12. Judge whether it is communication between UAVs and ground users or communication between UAVs in the UAV base station mode; if it is communication between UAVs and ground users, go to S131; otherwise, go to S132;

[0022] S131. Take the maximum efficiency ratio as the optimization goal and go to S14;

[0023] S132. Take the minimum outage probability as the optimization goal and go to S14;

[0024] S14. Design decision variables for the task to be completed, and go to S15;

[0025] S15. Add constraint conditions, and go to S16;

[0026] S16. After the optimization objective, decision variables, and constraint conditions are clear, formulate a joint optimization problem.

[0027] Furthermore, the specific steps of the planning solution described in S2 are as follows:

[0028] S21. Analyze the search space complexity of the joint optimization problem, and go to S22;

[0029] S22. Judge whether the optimization problem can be directly solved by the reinforcement learning algorithm; if so, go to S26; otherwise, go to S23;

[0030] S23. Judge whether the decision variables of the optimization problem can be decoupled; if so, go to S241; otherwise, go to S242;

[0031] S241. Decouple the variables of the joint optimization problem, and transform the joint optimization problem into several sub-optimization problems with lower dimensions, and go to S251;

[0032] S242. Perform pre-training to obtain a sub-optimal solution, and go to S252;

[0033] S251. Plan the solution to the sub-optimization problem, and go to S26;

[0034] S252. Starting from the sub-optimal solution, plan the search scheme in the policy domain, and go to S26;

[0035] S26. Based on the above processing flow, formulate a solution with reinforcement learning as the main body.

[0036] Furthermore, the specific steps of the planning solution described in S3 are as follows:

[0037] S31. Judge whether the optimization problem consists of multiple sub-problems; if so, go to S32; otherwise, go to S33;

[0038] S32. Design an intelligent algorithm for the problem that cannot be solved by reinforcement learning to reduce the dimension of the decision variables, and go to S33;

[0039] S33. Judge whether it is a UAV swarm problem; if so, go to S341; otherwise, go to S342;

[0040] S341. Design a multi-agent reinforcement learning algorithm, and go to S351;

[0041] S342. Design a single-agent reinforcement learning algorithm, and go to S352;

[0042] S351. The reinforcement learning algorithm is trained by the central unit, and then go to S36;

[0043] S352. The reinforcement learning algorithm is trained by the UAV itself, and then go to S37;

[0044] S36. The reinforcement learning strategy is updated periodically by the central unit, transmitted to the UAV, and then go to S37;

[0045] S37. The UAV formulates the strategy to be executed according to the environmental state and UAV information.

[0046] Furthermore, the specific steps of the UAV policy mapping in S4 are as follows:

[0047] S41. Whether the UAV policy needs to continue training; if yes, go to S421; otherwise, go to S422;

[0048] S421. The UAV executes the exploration strategy and maps it to physical behavior;

[0049] S422. The UAV executes the greedy strategy and maps it to physical behavior.

[0050] For the above method, the present invention proposes an intelligent decision-making architecture for a new UAV communication system, as shown in Figure 1 shown, which has 4 levels, including a decision analysis layer, a task planning layer, an ADCC layer, and a task execution layer, and has cognitive and decision-making functions;

[0051] The decision analysis layer is used to execute step S1 in the decision-making process and step R4 in the cognitive process. Specifically, in the decision-making process, it receives the resource management task, establishes an optimization problem according to the resource management task, determines the constraint conditions, and sends the established optimization problem to the task planning layer; in the cognitive process, it receives the current system dynamic reachable performance, environmental information, and UAV information sent by the task planning layer, and judges whether the performance index of the current UAV resource management decision meets the preset goal according to the resource management task. If not, optimization is carried out;

[0052] The task planning layer is used to execute step S2 in the decision-making process and step R3 in the cognitive process. Specifically, in the decision-making process, according to the received optimization problem, it plans the solution of the task, decouples the joint optimization problem into sub-problems to be solved by specific algorithms and sends them to the ADCC layer; in the cognitive process, it receives the UAV information and environmental information sent by the ADCC layer, calculates the system dynamic reachable performance, and transmits the calculated performance, environmental information, and UAV information to the decision analysis layer and the decision-making unit of this layer;

[0053] The ADCC layer is used to execute step S3 in the decision-making process and step R2 in the cognitive process. Specifically, in the decision-making process, according to sub-problems and on the basis of UAV collaboration, specific algorithms for solving these sub-problems are designed, and a UAV control strategy is output to the task execution layer under the current environmental state and UAV state; in the cognitive process, UAV information is acquired and environmental information sensed by the UAV in the task execution layer is received.

[0054] The task execution layer is used to execute step S4 in the decision-making process and step R1 in the cognitive process. Specifically, in the decision-making process, according to the received control strategy, the strategy of the UAV is mapped into physical behavior and then interacts with the environment; in the cognitive process, environmental information is acquired and sent to the ADCC layer.

[0055] The beneficial effects of the present invention are:

[0056] The present invention can achieve intelligent decision-making of UAV communication resources through a hierarchical method: the UAV decision analysis layer constructs a reasonable optimization problem according to high-performance communication requirements, the task planning layer simplifies the complex optimization problem to promote the convergence of the method of the present invention, the ADCC layer designs corresponding intelligent algorithms for the simplified problem, can continuously interact with the environment, learn through trial and error according to the feedback of the environment, so as to improve the performance of its own strategy, and the task execution layer maps the strategy into physical behavior; compared with the traditional management method of UAV communication resources, the present invention does not need to obtain specific environmental information for each one, does not need to construct a convex optimization problem, nor does it need to plan the strategy in advance, and is more suitable for dynamic environments, especially dynamic interference environments; and the present invention has a low computational complexity and strong convergence, which is of great significance for high-performance communication of UAVs and has strong application value. Brief Description of the Drawings

[0057] Figure 1 is the overall hierarchical schematic diagram of the intelligent architecture of the UAV communication system;

[0058] Figure 2 is the processing flow chart of the decision analysis layer;

[0059] Figure 3 is the processing flow chart of the task planning layer;

[0060] Figure 4 is the processing flow chart of the ADCC layer;

[0061] Figure 5 is the processing flow chart of the task execution layer. Specific Embodiments

[0062] The technical solution of the present invention will be described in detail below in conjunction with the embodiments and the drawings.

[0063] Taking the resource management problem of the communication link between the UAV swarm base station and ground users as an example; the UAV swarm needs to jointly optimize user association, flight trajectory, transmit power, and resource block allocation strategies to provide high-performance communication services for ground users.

[0064] As Figure 1 shown, the specific steps of the cognitive scheme adopted in this embodiment are as follows:

[0065] R1. The task execution layer senses environmental information: The UAV senses the interference time-frequency pattern, collects the status information of ground users, and transmits it to the ADCC layer and the decision-making unit of this layer;

[0066] R2. The ADCC layer obtains the information of the UAVs performing tasks: The UAV shares its own coordinates, flight speed, flight direction, and the sensed environmental information with other UAVs and the central unit, and transmits them to the task planning layer;

[0067] R3. The task planning layer calculates the ratio that meets the communication rate requirement and the total transmit power of the UAV swarm according to information such as environmental feedback, and then calculates the performance of the user UAV decision according to the formulated objective function, and transmits it to the decision analysis layer and the decision-making unit of this layer;

[0068] R4. The decision analysis layer determines whether the performance of the current decision can meet the indicators, and if not, how to adjust the parameters.

[0069] As Figure 1 shown, the specific steps of the decision-making scheme adopted in this embodiment are as follows:

[0070] S1. The decision analysis layer formulates an optimization problem according to this task and determines the constraint conditions. As Figure 2 shown, it specifically includes:

[0071] S11. Judge that the UAV communication mode is the base station mode, and go to S12;

[0072] S12. Judge that it is the communication between the UAV and ground users in the UAV base station mode, and go to S131;

[0073] S131. Take the maximum efficiency ratio as the optimization goal. Specifically, the UAV swarm needs to use as low transmit power as possible to meet the communication rate requirements of users, and go to S14;

[0074] S14. Design the decision variables as: a. user association, b. flight trajectory, c. transmit power, and d. resource block allocation, and go to S15;

[0075] S15. The added constraints are as follows: a. Each user can only be provided with communication services by one drone. b. A resource block can only be uniquely allocated to one user. c. There is a constraint on the maximum flight speed of the drone. d. There is a constraint on the flight altitude of the drone. e. There is a collision avoidance constraint. f. There is a maximum transmission power constraint. g. There is a constraint on the transmission rate received by the ground user. Then go to S16;

[0076] S16. Integrate the optimization objective, decision variables, and constraints to formulate a joint optimization problem.

[0077] S2. According to the formulated optimization problem, the task planning layer plans a reasonable task solution, as Figure 3 shown, specifically including:

[0078] S21. The dimension of the decision variables of this joint optimization problem is very high. Even if only the user association problem is considered, there are possibilities, where is the number of drones, and is the number of users. Therefore, this optimization problem is a non-deterministic polynomial (NP) hard problem; this optimization problem also contains both discrete variables and continuous variables, and there is a strong non-linearity between the objective function and the decision variables. Therefore, it is also a mixed integer non-linear programming (MINLP) problem;

[0079] S22. Judge that it is difficult to guarantee the convergence of the algorithm directly using reinforcement learning for NP hard problems and MINLP. Then go to S23;

[0080] S23. Use the method of decoupling decision variables to reduce the dimension of decision variables. Then go to S241;

[0081] S241. Decouple the joint optimization problem into: a. Clustering users into clusters. b. The problem of drone-user cluster association. c. The problem of flight trajectory and power control. d. The problem of resource block allocation. Then go to S251;

[0082] S251. Use an unsupervised clustering algorithm to solve problem a in S241, use a matching algorithm to solve problem b, use a reinforcement learning algorithm suitable for continuous decision variables to solve problem c, and use a reinforcement learning algorithm suitable for discrete decision variables to solve problem d. Then go to S26;

[0083] S26. Integrate the method of decoupling decision variables to formulate a solution with reinforcement learning as the main body and other intelligent algorithms as the auxiliary.

[0084] S3. On the basis of drone cooperation, design algorithms to solve these sub-problems and output strategies, asFigure 4 As shown in the figure, specifically including:

[0085] S31. If the optimization problem consists of multiple sub-problems, go to S32;

[0086] S32. Use the Expectation-Maximization (EM) algorithm to solve the user clustering problem, and use the Kuhn-Munkres (KM) algorithm to solve the association problem between drones and clusters, then go to S33

[0087] S33. This problem is a drone swarm problem, go to S341;

[0088] S341. Design the Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm to solve the flight trajectory and power control problem, and design the Multi-Agent Advantage Actor-Critic (MAA2C) algorithm to solve the resource block allocation problem, then go to S351;

[0089] S351. Use the collected experience in the central unit to train MATD3 and MAA2C, then go to S36;

[0090] S36. The trained flight trajectory, power control, and resource block allocation strategies are periodically transmitted from the central unit to the drones, then go to S37;

[0091] S37. The drones input the environmental state and drone information into the MATD3 algorithm and the MAA2C algorithm to formulate the strategies to be executed;

[0092] S4. Map the strategies of the drones into physical behaviors, and then interact with the environment, such as Figure 5 As shown in the figure, specifically including:

[0093] S41. If the reinforcement learning algorithm training is not completed, go to S421; otherwise, go to S422;

[0094] S421. The drones select the current optimal strategy with a certain probability, explore other strategies with a certain probability, and map them into physical behaviors;

[0095] S422. The drones select the current optimal strategy and map it into physical behaviors.

[0096] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0097] In the present invention, unless otherwise clearly specified and defined, the terms such as "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0098] In the present invention, unless otherwise clearly specified and defined, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0099] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

Claims

1. An intelligent decision-making method for a UAV communication system, where the UAV communication system uses a UAV cluster base station composed of multiple UAVs to provide communication services for ground users, and a single UAV is used as a communication relay or base station; characterized in that, The intelligent decision-making method synchronizes the cognitive process and the decision-making process; The cognitive process includes: R1. Obtain environmental information through a drone, specifically, the drone performing the task senses the interference time-frequency pattern and collects the status information of ground users; R2. Obtain the information of the drone performing the task, including the drone coordinates, flight speed, and flight direction. At the same time, the drone performing the task shares the obtained environmental information with other drones in the drone cluster; R3. Calculate the performance of the current drone resource management decision according to the environmental information and the drone information, specifically, obtain the ratio of meeting the communication rate requirement and the total transmission power of the drone cluster from the environmental information and the drone information, and then obtain the performance of the drone resource management decision according to the set objective function; R4. According to the set task objective, judge whether the performance index of the current drone resource management decision meets the preset objective. If not, perform optimization; The decision-making process includes: S1. According to the set task objective, establish an optimization problem and determine the constraint conditions, specifically including: S11. Judge whether the drone communication mode is the relay mode or the base station mode; if it is the relay mode, go to S131; otherwise, go to S12; S12. Judge whether it is communication from the drone to the ground user or communication between drones in the drone base station mode; if it is communication from the drone to the ground user, go to S131, otherwise, go to S132; S131. Take the maximum efficiency ratio as the optimization objective and go to S14; S132. Take the minimum outage probability as the optimization objective and go to S14; S14. Design decision variables for the task to be completed; S15. Add constraint conditions; S16. After the optimization objective, decision variables, and constraint conditions are clear, formulate a joint optimization problem; S2. According to the formulated optimization problem, formulate a solution, specifically including: S21. Analyze the search space complexity of the joint optimization problem; S22. Judge whether this optimization problem can be directly solved by the reinforcement learning algorithm; if so, go to S26, otherwise, go to S23; S23. Judge whether the decision variables of the optimization problem can be decoupled; if so, go to S241, otherwise, go to S242; S241. Decouple the variables of the joint optimization problem, transform the joint optimization problem into several sub-optimization problems with lower dimensions, and go to S251; S242. Perform pre-training to obtain a sub-optimal solution and go to S252; S251. Plan the solution to the sub-optimization problem and go to S26; S252. Take the sub-optimal solution as the starting point and plan the search scheme in the strategy domain and go to S26; S26. Obtain a solution with reinforcement learning as the main body; S3. On the basis of drone cooperation, design a processing method for the solution to obtain a drone control strategy, specifically including: S31. Judge whether the optimization problem consists of multiple sub-problems; if so, go to S32, otherwise, go to S33; S32. Design an intelligent algorithm for the problem not solved by reinforcement learning to reduce the dimension of the decision variables and go to S33; S33. Judge whether it is a drone cluster problem; if so, go to S341, otherwise, go to S342; S341. Design a multi-agent reinforcement learning algorithm and go to S351; S342. Design a single-agent reinforcement learning algorithm and go to S352; S351. The reinforcement learning algorithm is trained by the control center and go to S36; S352. The reinforcement learning algorithm is trained by the UAV itself and go to S37; S36. The reinforcement learning policy is updated periodically by the control center and transmitted to the UAV, then go to S37; S37. The UAV formulates the control strategy to be executed according to the environmental state and UAV information; S4. Map the UAV's policy to physical behavior and then interact with the environment, specifically including: S41. Whether the UAV policy needs to continue training; if so, go to S421; otherwise, go to S422; S421. The UAV executes the exploration policy and maps it to physical behavior; S422. The UAV executes the greedy policy and maps it to physical behavior; The intelligent decision-making method of the UAV communication system makes intelligent resource management decisions based on the intelligent decision-making architecture. The intelligent decision-making architecture includes a decision analysis layer, a task planning layer, an algorithm design and cooperative control ADCC layer, and a task execution layer; The decision analysis layer is used to execute step S1 in the decision-making process and step R4 in the cognitive process. Specifically, in the decision-making process, it receives the resource management task, establishes an optimization problem based on the resource management task, determines the constraints, and sends the established optimization problem to the task planning layer; in the cognitive process, it receives the current system dynamic reachable performance, environmental information, and UAV information sent by the task planning layer, and judges whether the performance index of the current UAV resource management decision meets the preset goal according to the resource management task. If not, optimization is carried out; The task planning layer is used to execute step S2 in the decision-making process and step R3 in the cognitive process. Specifically, in the decision-making process, according to the received optimization problem, it plans the solution of the task, decouples the joint optimization problem into sub-problems solved by specific algorithms and sends them to the ADCC layer; in the cognitive process, it receives the UAV information and environmental information sent by the ADCC layer, calculates the system dynamic reachable performance, and transmits the calculated performance, environmental information, and UAV information to the decision analysis layer and the decision-making unit of this layer together; The ADCC layer is used to execute step S3 in the decision-making process and step R2 in the cognitive process. Specifically, in the decision-making process, according to the sub-problems, on the basis of UAV cooperation, it designs specific algorithms to solve these sub-problems and outputs the UAV control strategy to the task execution layer under the current environmental state and UAV state; in the cognitive process, it obtains the UAV information and receives the environmental information sensed by the UAV in the task execution layer; The task execution layer is used to execute step S4 in the decision-making process and step R1 in the cognitive process. Specifically, in the decision-making process, according to the received control strategy, it maps the UAV's policy to physical behavior and then interacts with the environment; in the cognitive process, it obtains the environmental information and sends it to the ADCC layer.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle path collaborative planning method and device based on hierarchical reinforcement learning

    CN109992000A

  • Unmanned aerial vehicle network hovering position optimization method based on multi-agent deep reinforcement learning

    CN111786713A