Optimization Control Method for Multi-UAV Communication System Based on Intrinsic Curiosity Mechanism

By applying reinforcement learning methods based on the intrinsic curiosity mechanism in the UAV communication system, efficient exploration and solar energy collection of drones in urban environments are achieved, problems of endurance and service paths are solved, and the quality and efficiency of communication services are significantly improved.

CN118647032BActive Publication Date: 2025-06-10SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410632869.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-06-10
Estimated Expiration
2044-05-21

AI Technical Summary

Technical Problem

How drones better explore and collect solar energy in urban environments to improve their endurance and determine the best way to provide communication services to users.

Method used

Using reinforcement learning method based on the intrinsic curiosity mechanism, the optimal centralized control strategy of the drone is realized by transforming optimization problems into Markov decision-making process. This method divides the training process into an exploration-intensive stage and a development-intensive stage, and uses the intrinsic curiosity module to promote autonomous exploration of drones.

Benefits of technology

The exploration capabilities of the drone are significantly improved, allowing it to effectively locate the light area and obtain energy in complex environments, solving the service path problem, and the simulation results show that the reward is 20% higher than the baseline method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118647032B_ABST
    Figure CN118647032B_ABST
Patent Text Reader

Abstract

The present invention discloses an optimization control method for a multi-UAV communication system based on an intrinsic curiosity mechanism. The goal is to learn an optimal multi-UAV centralized control strategy, enabling the UAVs to find lighting areas in an urban environment through curiosity-driven exploration and collect energy to continuously and stably provide communication services for users. First, a multi-UAV centralized control strategy based on reinforcement learning (RL) is proposed to maximize the cumulative communication service score. In the proposed framework, the curiosity reward generated by the intrinsic curiosity module (ICM) can serve as an internal incentive signal, allowing the UAVs to explore the environment without any prior knowledge. Second, a two-stage exploration protocol is proposed for easy practical implementation. The method of the present invention can obtain a higher cumulative communication service score in the exploitation-intensive stage, obtain a more accurate service path, and can well handle the exploration-exploitation trade-off.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of path planning, and particularly relates to an optimization control method for a multi-UAV communication system based on an intrinsic curiosity mechanism. Background Art

[0002] Unmanned Aerial Vehicles (UAVs) are becoming increasingly popular as an important part of future wireless communication systems. UAVs equipped with radio transceivers can act as mobile base stations to replace ground base stations (BSs), providing improved radio connections and cutting-edge on-demand services for ground users, while significantly reducing the infrastructure required during network development, thus significantly reducing deployment costs. Applications of UAVs include network extension and enhancement, rescue missions, crowd / traffic monitoring, mobile edge computing, cached content distribution, etc. Currently, comprehensive and in-depth research has been conducted on many aspects of UAV communication, including energy management, computing offloading, path planning, and radio resource allocation. Communication problems in UAV systems are usually transformed into optimization problems, and classical methods define them as mixed integer non-convex problems. Generally, iterative techniques decompress the initial NP-hard problem into a set of smaller problems. In UAV-based communication, the complexity of on-demand services, terrain undulations, and the mobility of UAVs cause characteristics such as user distribution, wireless channel conditions, and network topology to change over time.

[0003] Some research in the field of UAV communication systems has been studying energy efficiency issues. Power control in UAV-supported ultra-dense networks performs better compared to traditional fixed infrastructure-based networks. A coverage model considering multi-UAV systems for energy-efficient communication uses the K-means clustering method to cluster ground devices and then serves them with UAVs. This UAV deployment method can collect useful knowledge from Internet of Things devices through the uplink, thereby improving efficiency and energy conservation. The problem of maximizing the computing rate of a UAV-supported mobile edge computing (MEC) system considering both partial and binary computing offloading modes. A UAV BSs optimal placement algorithm to maximize ground base station coverage and improve network energy efficiency. The power minimization problem in UAV networks allows iterative optimization of the UAV unit boundaries and positions. If the boundaries of a given unit are given, the facility location framework is used to determine the UAV positions. However, the above research work has not studied the role played by energy harvesting in the service path problem of multi-UAV communication systems.

[0004] The energy-neutral drone Internet increases connectivity by overcoming energy limitations, thus providing continuous operation and longer endurance. To power the drones and reduce the energy gap between energy harvesting and energy consumption, wireless energy transfer is explored. Considering the energy causality constraint, the problem of resource allocation can be formulated as a non-convex optimization problem. This method aims to optimize the average throughput of a device-to-device (D2D) network assisted by drones. When calculating the outage probability, multiple urban environment parameters are considered. To extend the endurance of drones, wireless energy harvesting is proposed. To optimize network throughput, the duality of the information-theoretic uplink and downlink channels and dirty paper coding are considered. Considering a drone-enabled dual-user wireless energy transfer system, the drone provides a certain endurance for charging multiple energy receivers. All the above-mentioned work focuses on harvesting energy from radio frequency sources rather than using renewable energy such as solar energy.

[0005] An energy management framework using solar drones is applicable to cellular heterogeneous networks (HetNets). The goal of this framework is to jointly determine the optimal routes of ground satellite navigation systems and drones to reduce the overall energy consumption of the network. The drones can use the harvested solar energy or charging stations to supplement the energy of the batteries. The energy management framework in HetNets supported by drones is equipped with devices that capture solar energy and charge from fixed charging stations. To optimize the overall throughput of the solar drone system within a given time, an effective strategy must be designed. However, most of them require ground base stations and fixed charging stations. To reduce the constraints on drones and enable them to obtain energy by independently exploring and finding illuminated areas in the environment, the present invention proposes a new method for solving the service path problem of multi-drone communication systems.

[0006] The curiosity mechanism, as a promising enlightenment method, has been widely studied in reinforcement learning to solve sparse reward tasks. Curiosity-based methods are more effective and exploratory in reinforcement learning and can potentially shorten the convergence time of reinforcement learning training. For example, PathFinder is a practical human-computer interface software exploration framework that selects actions based on a curiosity-based reinforcement learning framework to discover additional unrecognized states. A web testing framework based on reinforcement learning uses a unique curiosity-driven reward function to effectively explore various web application behaviors. In some scenarios where the agent receives sparse external rewards, using the prediction error of the forward model as an intrinsic reward guides some exploration of the agent. So far, no research has utilized the curiosity mechanism to handle the service path problem in energy-limited multi-drone communication systems.

[0007] With the latest progress in machine learning, the problem of drone communication has a good answer in reinforcement learning. Reinforcement learning agents have powerful capabilities to make continuous judgments in an environment independent of the environmental model by constantly interacting with the surrounding environment and obtaining experience from the interactions. The research on drone communication systems based on reinforcement learning mainly focuses on contact avoidance control techniques for path planning. There is little research on how drones independently explore and find lighting areas in the city and utilize solar energy to improve their endurance. Due to the complexity of the urban environment, it is usually difficult to find lighting areas. In densely populated metropolitan areas, high-rise buildings tower into the clouds, often dividing the lighting areas into smaller parts. The lighting areas are usually located in areas with relatively few people. Therefore, the location of the lighting areas cannot be known in advance. A major challenge in improving the endurance of drones is to ensure that drones conduct sufficient autonomous exploration to find the location of the lighting areas and continue to provide communication services for user gathering places after replenishing energy. The high-dimensional state space and continuous action space are another difficulty. The existing deep deterministic policy gradient (DDPG) technology makes it challenging to find the optimal policy in the urban environment. Therefore, traditional methods are difficult to keep up with the complex environment as it becomes more diverse in the future. Currently, there is an urgent need for a multi-drone communication system based on autonomous detection. And the present invention proposes an optimized control method for a multi-drone communication system based on an intrinsic curiosity mechanism that can well solve the above problems. Summary of the Invention

[0008] Technical Problem: The technical problem to be solved by the present invention and the goal to be achieved.

[0009] The technical problem to be solved by the present invention is: how drones can better explore and collect solar energy to improve their endurance in the urban environment and determine the best way to provide communication services for users. The goal to be achieved by the present invention is to realize the optimal centralized control strategy of drones through reinforcement learning. When the energy of the drone is about to run out, this strategy enables the drone to actively explore and find the lighting area to obtain energy. In order to enable the drone to mainly aim at providing communication services while fully exploring, the entire reinforcement learning training process is divided into two stages, namely the exploration-intensive stage and the exploitation-intensive stage. The exploration-intensive stage focuses on exploring the environment and ensuring rapid coverage of all environmental details, while the exploitation-intensive stage focuses more on effectively formulating the best strategy.

[0010] Technical Solution: The complete technical means and methods of the present invention.

[0011] To achieve the above object, the technical solution of the present invention is as follows:

[0012] An optimization control method for a multi-UAV communication system based on an intrinsic curiosity mechanism, comprising the following steps:

[0013] First, design a multi-UAV to multi-user communication system.

[0014] In a target area of L×L, a group of N U UAVs in the swarm S U provide services to a group of N u ground users S u Most users are concentrated in several hotspots in the target area, and the remaining users are evenly distributed throughout the target area. The UAVs fly at a fixed height H in the target area, and each UAV has a height-directed antenna. Therefore, the transmission power of the UAV is concentrated within an angle of θ and provides the minimum required throughput to the ground users. The service range of the UAV to the ground is circular, and its radius is If a user is not within the service range of the UAV, no service is provided. The UAVs consume energy during flight, but can replenish energy when flying in the illuminated area. If there are buildings ahead, the UAVs need to detour.

[0015] The calculation methods for various performance indicators in the multi-UAV to multi-user communication system are defined as follows:

[0016] All UAVs have a backhaul link (such as a satellite link) to connect them to an external network. Since there is no spectrum overlap between the UAV backhaul link and the UAV user link, there is no mutual interference. The path loss from UAV k to ground user u is denoted as PL ku and is calculated by the following formula:

[0017]

[0018] where ω represents the spectral center frequency allocated to user u, d ku represents the three-dimensional distance between user u and UAV k, c represents the speed of light, and η represents the additional path loss of the line-of-sight link. dB represents decibel, which is the unit of the relative difference in signal strength. The signal-to-noise ratio SINR ku from UAV k to user u is calculated by the following formula:

[0019]

[0020] In the formula P t represents the power spectral density of the UAV, n e represents the power spectral density of the ambient noise, represents the set of UAVs covering user u.

[0021] R uis the minimum throughput required by each user. Drone k will only serve user u when and the following formula is satisfied:

[0022] W ku log 2 (1 + SINR ku ) ≥ R u

[0023] where W ku represents the bandwidth allocated by drone k to user u. From the above formula, it can be seen that each user only connects to the drone with sufficient bandwidth and the highest signal-to-noise ratio.

[0024] Within the time step t, the energy consumption of drone k consists of three parts: (i) E F represents the energy consumption of drone k flying a distance at a fixed speed v within time t; (ii) E S represents the energy consumption of the drone signal transmitted on the backhaul link and the drone-user link; (iii) E O represents the operating cost energy of the drone, which is proportional to time t. The energy consumption of drone k is calculated by the following formula:

[0025]

[0026] The initial battery power of drone k is denoted as The time segment of an episode is divided into N T time steps. Each drone flies at a fixed speed v and can fly in any direction. Within the time step t, the flight distance of drone k is denoted as The horizontal flight energy consumption P of the drone hor is calculated by the following formula:

[0027]

[0028] where the weight W of the drone is in Newtons, the air density is denoted by ρ, and the rotary wing area of the drone is denoted by D. Due to the influence of the horizontal flight speed, the energy consumption of the drone during horizontal flight is lower than that during hovering.

[0029] When drones fly in the illumination area, the energy they harvest is determined by many factors, and the most important factor is the illumination radiation intensity. The relationship between the solar radiation intensity and the harvested energy is calculated by the following formula:

[0030]

[0031] where \(I(t)\) is the light radiation intensity, and its threshold is denoted as \(K\). c , \(\eta\) c is the efficiency. When exceeding this threshold, the efficiency can be roughly expressed as a constant.

[0032] Combining and analyzing the above multi-UAV multi-user communication system model, the present invention constructs a constrained optimization problem involving UAV trajectory optimization and communication resource allocation, aiming to enable the agent to obtain an optimal multi-UAV communication service strategy that maximizes the communication service score, while restricting the movement of UAVs in the risk area, so as to achieve a balance between maximizing communication capacity and minimizing cumulative risk. The formalization of the constrained optimization problem is as follows.

[0033]

[0034]

[0035] where represents the horizontal coordinate of UAV \(k\) when flying at time step \(t\). represents the total number of users successfully served within time step \(t\). Considering the number of users served by the UAV and the necessary throughput, \(R\) t reflects the communication service quality within time step \(t\). If user \(u\) is successfully served within time step \(t\), then define as 1, otherwise 0. The number, position, and battery status of the UAVs together constitute the parameters of the UAVs. The parameters of the UAVs and several other parameters, such as user distribution and spectrum access rules, jointly determine the magnitude of. The exponent \(\beta>0\) is a factor measuring the overall user satisfaction of the agent with respect to the number of users successfully served.

[0036] Based on the reinforcement learning framework, the present invention transforms the optimization problem into a Markov decision process and designs the three elements of reinforcement learning, namely the state space \(S\) t , the action space \(A\) t , and the reward function \(r\); the specific design is as follows:

[0037] The state space includes the position and remaining battery power of each UAV. The position of the UAV largely determines the number of users successfully served. Within time step \(t\leq N\) T , since the flight altitude of the UAV is fixed during service, only two-dimensional coordinates need to be considered within the target area. The maneuverability of the UAV is restricted, that is At time step \(t\), the remaining battery power of the UAV is expressed as The battery power of drones has a direct impact on their actions. Generally speaking, the state space cardinality of the proposed exploration method is 3N U , which is expressed as follows:

[0038]

[0039] The action space includes the moving distances and directions of all drones. To enable the central agent to manage the movements of all drones, an action A is formed at time step t t , to reflect the collective movement of the drones. The moving distance and the moving direction constitute the action of drone k at time step t At each time step t, a drone can move a maximum distance d in any direction max or hover. Generally speaking, the state space cardinality of the proposed exploration method is 2N U , which is expressed as follows

[0040]

[0041] The reward function is divided into two parts, namely, external reward and internal reward.

[0042] The external reward within time step t is represented by , which is a function representing the instantaneous communication service score within time step t (i.e., R t ), and is specifically expressed as follows:

[0043]

[0044] In the above formula, the instantaneous communication service score is divided by because better convergence may occur if the absolute value of remains within 1. In addition, when β > 1, the reward difference between different values is amplified, which can prompt the drones to take proactive measures when the energy is about to run out. However, β cannot be too large because in the initial simulation, β ≥ 3 will result in a lower final convergence value. In addition, within N T time steps, the present invention ultimately maximizes the cumulative reward by maximizing the cumulative communication service score. For each drone that crosses the boundary, runs out of energy, or collides with an obstacle, giving a negative reward as a penalty is an alternative reward function design. F in the above formula is the penalty flag, and the size of the penalty flag varies with the number of drones that cross the boundary, run out of energy, or collide with an obstacle. If a drone leaves the target area during training, the current movement is cancelled. Execute the current action A under the current state S t ​t will result in a negative reward. During the training process, the negative reward strongly "opposes" the positive reward to improve the convergence performance and speed. Throughout the training process, the ratio of positive incentives to negative incentives is constantly changing, making it increasingly difficult to achieve satisfactory convergence and imposing higher computational requirements. In addition to the external rewards obtained from the interaction between the drone and the environment, the intrinsic curiosity module (ICM) can also provide specific internal rewards for the agent.

[0045] The internal reward within time step t is denoted by . The curiosity exploration mechanism performs well in solving the sparse reward problem, which is achieved by enhancing the agent's autonomous exploration ability. In the environment, the agent can achieve a full understanding of the environment through curiosity-driven exploration. The ICM consists of two subsystems: an intrinsic reward generator that generates an internal incentive signal driven by curiosity; and a policy that generates a series of actions to optimize the reward signal. The goal of training the policy subsystem during the exploration-intensive phase is to maximize the total value r of the two rewards at time step t. t There are two models in the ICM, namely the forward model and the inverse model.

[0046] The forward model consists of a set of fully connected layers. The input to these layers is the state vector obtained by concatenating s t with a t . It takes a t and s t as inputs to predict the state at time step t + 1. The function g is called the forward model and is specifically as follows:

[0047]

[0048] where the predicted estimate of the state s t+1 is denoted as The parameters for optimizing the neural network are defined as θ F , and the aim is to minimize the following loss function:

[0049]

[0050] The inverse model uses a series of convolutional layers to map the state vector, which connects s t and s t+1 . After each convolutional layer is an ELU non-linearity. To predict the potential action process, the vector is input into a fully connected layer and then exits from this layer. Training a deep neural network is equivalent to learning a function f, defined as:

[0051]

[0052] Among them, action a t The predicted estimated value is denoted as Define the parameters of the optimized neural network as θ I , and its optimization equation is:

[0053]

[0054] Among them, the loss function L I is used to measure the actual action a t and the predicted action The difference between. The tuple (s t , s t+1 ) is obtained by interacting with the environment using the current policy π(s). The function f is called the inverse model.

[0055] To provide an intrinsic reward based on curiosity, the forward loss and the backward loss are co-optimized respectively, and the prediction error generated by the forward model is fully utilized as the intrinsic curiosity reward for training the agent's policy. The internal reward The calculation formula is as follows:

[0056]

[0057] The present invention uses the DDPG-ICM algorithm to solve the optimal policy. The goal of ICM is to enable the agent to explore the environment completely autonomously, but the problem is that the agent with ICM becomes significantly more active. In practice, as long as the agent fully explores the surrounding environment and finds the location of the lighting area, the intrinsic curiosity reward is no longer needed. To ensure an effective exploration process using the proposed DDPG-ICM algorithm, a two-stage exploration protocol is proposed. The entire reinforcement learning training process is divided into an exploration-intensive stage and a exploitation-intensive stage. The exploration-intensive stage focuses more on environmental exploration and ensures that all environmental details are quickly covered using the proposed method. The exploitation-intensive stage focuses more on effectively and precisely formulating the optimal policy.

[0058] Train the network using the DDPG-ICM algorithm in the exploration-intensive stage. The actor network and the critic network have their respective target networks with the same structure, denoted as and Q′ respectively. The task of the actor network is to improve the performance of the agent by refining the deterministic policy, while the goal of the critic network is to estimate the action value function of the policy received from the actor network. Therefore, the state S t (or S t+1 ) is the input of the actor network (or the target actor network), and the action A t (or A t+1 ) is its output. At the same time, the output of the critic network (or the target critic network) is based on St (or S t+1 ) and A t (or A t+1 )'s action value function approximates Q(S t , A t |θ Q )(or Q′(S t+1 , A t+1 |θ Q′ ). Store the experience in the experience replay buffer ε with capacity M. To train the network, randomly sample a mini-batch of size G << M from this buffer. In the exploration-intensive phase, actions with higher curiosity values are preferentially selected.

[0059] This scenario-based training strategy has several advantages. First, an effective exploration method is random initialization or exploratory start. Through the collaboration of random initialization and exploratory start, the exploration of the state space can be further promoted. Second, by restricting the state range, the search efficiency can be improved, and the state values can be conveniently normalized, which is beneficial to network training. Third, the penalty for deviating from the acceptable range helps to adopt a sensible learning policy.

[0060] After the exploration-intensive phase, a curiosity-driven exploration strategy can be obtained, which can enable the drone to comprehensively explore the environment. However, due to the existence of the inherent curiosity reward, this strategy makes the drone highly active during actual flight, resulting in inaccurate service paths. Therefore, the exploitation-intensive phase is extended to improve performance.

[0061] The exploitation-intensive phase is designed to be implemented during actual flight. In addition, the exploitation-intensive phase and the exploration-intensive phase have the same actions, states, and network structures. However, starting from the trained exploitation-intensive network, these networks are trained to ensure a consistent initial exploration strategy, rather than randomly initializing the network weights. At the same time, the inherent curiosity reward is removed from the reward function. The remaining steps are the same as those in the exploration-intensive phase. First, randomly initialize the actor network and the critic network Q. Initialize their target networks and Q′. In addition, the experience replay buffer ε is also initialized. The actor network observes the environmental state S t , and then determines the action A t . Before the agent executes this action, it makes a preliminary decision at S t . The drone will only stop moving when it leaves the target area, runs out of energy, or collides with an obstacle; otherwise, they will execute the action A t . After that, the state enters S t+1 , and the agent obtains r t as the reward. There is a threshold NThre is used to determine the stage of the training process. When the number of training episodes is less than or equal to the threshold N Thre , the training is in the exploration-intensive stage. At this time, the reward value within the time step t When the number of training episodes is greater than the threshold N Thre , the training is in the exploitation-intensive stage, and the reward value within the time step t Finally, the network is trained to improve the performance of the system. Once the training conditions are met, a small batch of samples is randomly drawn from the experience replay buffer ε. Finally, the actor network, the critic network, and their corresponding target networks are updated respectively.

[0062] Beneficial effects: The benefits brought by the present invention and the achieved indicators.

[0063] Through curiosity-driven exploration, the present invention significantly enhances the exploration ability of the UAV, enabling the UAV to locate the light area and obtain energy in a vast exploration environment, and successfully handle the service path problem in a vast and complex environment. In addition, a two-stage exploration protocol is provided for the DDPG-ICM method to be implemented in practical applications. After generating the exploration method in the first stage, the intrinsic curiosity reward is removed in the second stage to develop the optimal centralized control strategy, enabling the UAV to accurately find the appropriate service path. The simulation results show that this method significantly improves the exploration ability of the multi-UAV communication system, and the reward in the exploitation-intensive stage is 20% higher than that of the baseline method. More critically, the present invention provides a potential centralized control method for the service path problem of multi-UAV communication systems with high-dimensional state spaces and continuous action spaces. Brief Description of the Drawings

[0064] Figure 1 It is a schematic diagram of the multi-UAV-to-user communication system environment.

[0065] Figure 2 It is a schematic diagram of the interaction between the intrinsic curiosity module and the environment.

[0066] Figure 3 It is a schematic diagram of the simulation experiment scenario. Among them Figure 3 (a) is a schematic diagram of Scenario 1, Figure 3 (b) is a schematic diagram of Scenario 2.

[0067] Figure 4 It is a schematic diagram of the comparison of the simulation results of the proposed method and the baseline method in Scenario 1. Among them Figure 4 (a) is a schematic diagram of the performance comparison of UAV 1, UAV 2, and UAV 3 in the case of low initial energy under the two methods; Figure 4 (b) is a schematic diagram of the performance comparison of UAV 4 and UAV 5 in the case of sufficient initial energy under the two methods. Figure 4(c) Schematic diagram of the trajectory of the UAV during the convergence stage under the baseline method (DDPG algorithm); Figure 4 (d) Schematic diagram of the trajectory of the UAV during the exploration-intensive stage under the proposed method; Figure 4 (e) Schematic diagram of the trajectory of the UAV during the exploitation-intensive stage under the proposed method. Figure 4 (c), Figure 4 (d), Figure 4 In (e), the blue curve represents the flight trajectory of UAV 1, the orange curve represents the flight trajectory of UAV 2, the green curve represents the flight trajectory of UAV 3, the red curve represents the flight trajectory of UAV 4, and the purple curve represents the flight trajectory of UAV 5.

[0068] Figure 5 Schematic diagram for comparing the simulation results of the proposed method and the baseline method in Scenario 2. Among them Figure 5 (a) Schematic diagram for comparing the performance of UAV 1, UAV 2, and UAV 3 under the two methods in the case of low initial energy; Figure 5 (b) Schematic diagram for comparing the performance of UAV 4 and UAV 5 under the two methods in the case of sufficient initial energy. Figure 5 (c) Schematic diagram of the trajectory of the UAV during the convergence stage under the baseline method (DDPG algorithm); Figure 5 (d) Schematic diagram of the trajectory of the UAV during the exploration-intensive stage under the proposed method; Figure 5 (e) Schematic diagram of the trajectory of the UAV during the exploitation-intensive stage under the proposed method. Figure 5 (c), Figure 5 (d), Figure 5 In (e), the blue curve represents the flight trajectory of UAV 1, the orange curve represents the flight trajectory of UAV 2, the green curve represents the flight trajectory of UAV 3, the red curve represents the flight trajectory of UAV 4, and the purple curve represents the flight trajectory of UAV 5. Detailed implementation method

[0069] The following further clarifies the present invention in conjunction with the accompanying drawings and specific implementation methods. It should be understood that the following specific implementation methods are only used to illustrate the present invention and not to limit the scope of the present invention.

[0070] As shown in the figure, the optimization control method for a multi-UAV communication system based on the intrinsic curiosity mechanism according to the present invention includes the following steps:

[0071] S1: Design a multi-UAV to multi-user communication system.

[0072] In the target area of L×L, a group of N U UAVs S U transmit to a group of N u ground users Su Provide services. The vast majority of users are concentrated in several hotspots within the target area, and the remaining users are evenly distributed throughout the target area. The drones fly at a fixed altitude H within the target area, and each drone has a highly directional antenna. Therefore, the transmission power of the drone is concentrated within an angle of θ and provides the minimum required throughput to the ground users. The service area of the drone on the ground is circular, and its radius is If a user is not within the service range of the drone, no service is provided. The drones consume energy during flight, but can replenish energy when flying in the illuminated area. If there is a building ahead, the drone needs to detour.

[0073] S2: Define the calculation methods for various performance metrics in a multi-drone to multi-user communication system.

[0074] All drones have a backhaul link (such as a satellite link) to connect them to an external network. Since there is no spectrum overlap between the drone backhaul link and the drone-user link, there is no mutual interference. The path loss from drone k to ground user u is denoted as PL ku and is calculated by the following formula: where ω represents the center frequency of the spectrum allocated to user u, d ku represents the three-dimensional distance between user u and drone k, c represents the speed of light, and η represents the additional path loss of the line-of-sight link. The signal-to-noise ratio SINR of drone k to user u ku is calculated by the following formula:

[0075]

[0076] In the formula P t represents the power spectral density of the drone, n e represents the power spectral density of the ambient noise, represents the set of drones covering user u.

[0077] R u is the minimum throughput required for each user. Drone k will only serve user u when and the following condition is satisfied:

[0078] W ku log 2 (1 + SINR ku ) ≥ R u

[0079] In the formula, W ku represents the bandwidth allocated by drone k to user u. From the above formula, it can be seen that each user only connects to the drone with sufficient bandwidth and the highest signal-to-noise ratio.

[0080] During the time step \(t\), the energy consumption of drone \(k\) consists of three parts: (i) \(E\) F represents the energy consumption of drone \(k\) flying a distance at a fixed speed \(v\) within time \(t\) ; (ii) \(E\) S represents the energy consumption for the transmission of the drone signal on the backhaul link and the drone - user link; (iii) \(E\) O represents the operating cost energy of the drone, which is proportional to the time \(t\). The energy consumption of drone \(k\) is calculated by the following formula:

[0081]

[0082] The initial battery level of drone \(k\) is denoted as The time segment of an episode is divided into \(N\) T time steps. Each drone flies at a fixed speed \(v\) and can fly in any direction. During the time step \(t\), the flying distance of drone \(k\) is denoted as The horizontal flight energy consumption \(P\) of the drone hor is calculated by the following formula:

[0083]

[0084] where the weight \(W\) of the drone is in Newtons, the air density is denoted by \(\rho\), and the rotating wing area of the drone is denoted by \(D\). Due to the influence of the horizontal flight speed, the energy consumption of the drone during horizontal flight is lower than that during hovering.

[0085] When the drones fly in the illumination area, the energy they harvest is determined by many factors, and the most important factor is the illumination radiation intensity. The relationship between the solar radiation intensity and the harvested energy is calculated by the following formula:

[0086]

[0087] where \(I(t)\) is the illumination radiation intensity, and its threshold is denoted as \(K\) c , \(\eta\) c is the efficiency. When exceeding this threshold, the efficiency can be roughly expressed as a constant.

[0088] S3: Construct a constrained optimization problem involving drone trajectory optimization and communication resource allocation, aiming to obtain an optimal multi - drone communication service strategy that maximizes the communication service score for the agent, while restricting the movement of drones in the risk area, thus achieving a balance between maximizing the communication capacity and minimizing the cumulative risk. The constrained optimization problem is formalized as follows:

[0089]

[0090]

[0091] Wherein, represents the horizontal coordinate of the UAV k during flight at time step t, represents the total number of users successfully served within time step t. Considering the number of users served by the UAV and the required throughput, R t reflects the quality of communication service within time step t. If user u is successfully served within time step t, then define as 1, otherwise 0. The number, position, and battery status of the UAVs together constitute the parameters of the UAVs. The parameters of the UAVs and several other parameters, such as user distribution and spectrum access rules, jointly determine the magnitude of. The exponent β > 0 is a factor measuring the overall user satisfaction of the agent with respect to the number of users successfully served.

[0092] S4: Based on the reinforcement learning framework, transform the optimization problem into a Markov decision process, and design the three elements of reinforcement learning, namely the state space S t , the action space A t , and the reward function r; the specific design is as follows:

[0093] The state space includes the position and remaining battery power of each UAV. The position of the UAV largely determines the number of users successfully served. At time step t ≤ N T Since the flight altitude of the UAV is fixed during service, only two-dimensional coordinates need to be considered Within the target area, the mobility of the UAV is restricted, that is At time step t, the remaining battery power of the UAV is expressed as The battery power of the UAV has a direct impact on their actions. Generally speaking, the state space cardinality of the proposed exploration method is 3N U , which is expressed as follows:

[0094]

[0095] The action space includes the moving distance and moving direction of all UAVs. In order for the central agent to manage the movement of all UAVs, an action A t is formed at time step t to reflect the collective movement of the UAVs. The moving distance and the moving direction constitute the action of UAV k at time step t At each time step t, a UAV can move a maximum distance d in any directionmax Hovering. Generally speaking, the state space cardinality of the proposed exploration method is 2N U , which is expressed as follows

[0096]

[0097] The reward function is divided into two parts, namely external reward and internal reward.

[0098] The external reward within the time step t is represented by , which is a function representing the instantaneous communication service score within the time step t (i.e., R t ), and is specifically expressed as follows:

[0099]

[0100] In the above formula, the instantaneous communication service score is divided by because if the absolute value of remains within 1, better convergence may occur. In addition, when β > 1, the reward difference between different values is amplified, which can prompt the UAV to take proactive measures when the energy is about to run out. However, β cannot be too large because in the initial simulation, β ≥ 3 will lead to a lower final convergence value. In addition, within N T time steps, the present invention finally maximizes the cumulative reward by maximizing the cumulative communication service score. For each UAV that crosses the boundary, runs out of energy, or collides with an obstacle, giving a negative reward as a penalty is an alternative reward function design. F in the above formula is a penalty flag, and the size of the penalty flag changes with the number of UAVs that cross the boundary, run out of energy, or collide with an obstacle. If a UAV leaves the target area during the training process, the current movement is cancelled. Executing the current action A t in the current state S t will result in a negative reward. During the training process, the negative reward strongly "opposes" the positive reward to improve the convergence performance and convergence speed. During the entire training process, the ratio of positive incentives to negative incentives is constantly changing, making it increasingly difficult to achieve satisfactory convergence and the computational requirements are also increasing. In addition to the external reward obtained from the interaction between the UAV and the environment, the intrinsic curiosity module (ICM) can also provide specific internal rewards for the agent.

[0101] The internal reward within the time step t is represented by It is shown that the curiosity exploration mechanism performs well in solving the sparse reward problem, which is achieved by enhancing the agent's autonomous exploration ability. In the environment, the agent can achieve a full understanding of the environment through curiosity-driven exploration. ICM consists of two subsystems: an intrinsic reward generator that produces an internal incentive signal driven by curiosity; and a policy that generates a series of actions to optimize the reward signal. The goal of training the policy subsystem in the exploration-intensive phase is to maximize the total value r of the two rewards at time step t. t There are two models in ICM, namely the forward model and the inverse model.

[0102] The forward model consists of a set of fully connected layers. The input to these layers is the state vector obtained by concatenating s t with a t . It takes a t and s t as inputs to predict the state at time step t+1. The function g is called the forward model and is specifically as follows:

[0103]

[0104] where the predicted estimated value of the state s t+1 is denoted as The parameters for optimizing the neural network are defined as θ F , and its purpose is to minimize the following loss function:

[0105]

[0106] The inverse model uses a series of convolutional layers to map the state vector, which connects s t and s t+1 . After each convolutional layer is an ELU non-linearity. To predict the potential action process, the vector is input into a fully connected layer and then exits from this layer. Training the deep neural network is equivalent to learning a function f, defined as:

[0107]

[0108] where the predicted estimated value of the action a t is denoted as The parameters for optimizing the neural network are defined as θ I , and its optimization equation is:

[0109]

[0110] where the loss function L I is used to measure the difference between the actual action a t and the predicted action . The tuple (s t ,s t+1) is obtained by interacting with the environment using the current policy π(s). The function f is called the inverse model.

[0111] To provide curiosity-based intrinsic rewards, the forward loss and the inverse loss are co-optimized respectively, making full use of the prediction error generated by the forward model as the intrinsic curiosity reward for training the agent's policy, the internal reward The calculation formula is as follows:

[0112]

[0113] S5: Solve the optimal policy using the DDPG-ICM algorithm. The goal of ICM is to enable the agent to explore the environment completely autonomously. However, the problem is that the agent with ICM becomes significantly more active. In practice, as long as the agent fully explores the surrounding environment and finds the location of the lighting area, the intrinsic curiosity reward is no longer needed. To ensure an effective exploration process using the proposed DDPG-ICM algorithm, a two-stage exploration protocol is proposed. The entire reinforcement learning training process is divided into an exploration-intensive stage and an exploitation-intensive stage. The exploration-intensive stage focuses more on environmental exploration and ensures that all environmental details are quickly covered using the proposed method. While the exploitation-intensive stage focuses more on effectively and precisely formulating the optimal policy.

[0114] Train the network using the DDPG-ICM algorithm in the exploration-intensive stage. The actor network and the critic network have their respective target networks with the same structure, denoted as and Q′ respectively. The task of the actor network is to improve the performance of the agent by refining the deterministic policy, while the goal of the critic network is to estimate the action value function of the policy received from the actor network. Therefore, the state S t (or S t+1 ) is the input of the actor network (or the target actor network), and the action A t (or A t+1 ) is its output. At the same time, the output of the critic network (or the target critic network) is the action value function approximation Q(S t (or S t+1 ) and A t (or A t+1 ) based on Q(S t , A t |θ Q )(or Q′(S t+1 , A t+1 |θ Q′ ). Store the experiences in the experience replay buffer pool ε with a capacity of M. To train the network, randomly sample a mini-batch of size G << M from this buffer pool. In the exploration-intensive stage, actions with higher curiosity values are preferentially selected.

[0115] This scenario-based training strategy has several advantages. First, an effective exploration method is random initialization or exploratory start. Through the collaboration of random initialization and exploratory start, the exploration of the state space can be further promoted. Second, by restricting the state range, the search efficiency can be improved, and the state values can be conveniently normalized, which is beneficial to network training. Third, the penalty for deviating from the acceptable range helps to adopt a sensible learning policy.

[0116] After the exploration-intensive phase, a curiosity-driven exploration strategy can be obtained, which can enable the UAV to comprehensively explore the environment. However, due to the existence of the inherent curiosity reward, this strategy makes the UAV highly active during actual flight, resulting in inaccurate service paths. Therefore, the exploitation-intensive phase is extended to improve performance.

[0117] The exploitation-intensive phase is designed to be implemented during actual flight. In addition, the exploitation-intensive phase and the exploration-intensive phase have the same actions, states, and network structures. However, starting from the trained exploitation-intensive network, these networks are trained to ensure a consistent initial exploration strategy instead of randomly initializing the network weights. At the same time, the inherent curiosity reward is removed from the reward function. The remaining steps are the same as those in the exploration-intensive phase. First, the actor network is randomly initialized and the critic network Q. Their target networks and Q′ are initialized. In addition, the experience replay buffer ε is also initialized. The actor network observes the environmental state S t , and then determines the action A t . Before the agent executes this action, it makes a preliminary decision at S t . The UAV will only stop moving when it leaves the target area, runs out of energy, or collides with an obstacle; otherwise, they will execute the action A t . After that, the state enters S t+1 , and the agent obtains r t as the reward. There is a threshold N Thre used to determine the stage of the training process. When the training episode is less than or equal to the threshold N Thre , the training is in the exploration-intensive phase. At this time, the reward value within the time step t When the training episode is greater than the threshold N Thre , the training is in the exploitation-intensive phase, and the reward value within the time step t Finally, the network is trained to improve the performance of the system. Once the training conditions are met, a small batch of samples is randomly drawn from the experience replay buffer ε. Finally, the actor network, the critic network, and their corresponding target networks are updated respectively.

Claims

1. An optimization control method for multi-UAV communication system based on intrinsic curiosity mechanism, characterized by: The following steps are involved: S1: Design a multi-UAV to multi-user communication system; S2: Calculate the performance indicators in the multi-UAV to multi-user communication system, including the path loss from UAV to ground users, the signal-to-noise ratio from UAV to users, the energy consumption of UAV, the energy consumption of UAV horizontal flight, and the relationship between solar radiation intensity and harvested energy; S3: Based on the above performance indicators, a constrained optimization problem involving UAV trajectory optimization and communication resource allocation is constructed; S4: Based on the reinforcement learning framework, the optimization problem is transformed into a Markov decision process, and the three elements of reinforcement learning are designed; S5: Use DDPG-ICM algorithm to solve the optimal strategy; The three elements of reinforcement learning in step S4 include the state space S t , action space A t , reward function r; the details are as follows: The state space contains the position and remaining battery power of each drone; at time step t≤N T Only two-dimensional coordinates are considered At time step t, the remaining battery power of the drone is expressed as The battery charge of drones has a direct impact on their actions; overall, the cardinality of the state space of the proposed exploration method is 3N U , which is expressed as follows: The action space contains the moving distance and moving direction of all drones; the action A is formed at the time step t t , to reflect the collective movement of drones; moving distance and moving direction Constitutes the action of drone k in time step t At each time step t, a drone can move a maximum distance d in any direction max Hover; the state space cardinality of the proposed exploration method is 2N U , which is expressed as follows The reward function is divided into two parts, namely external rewards and internal rewards; The external reward in time step t is It is a function representing the instantaneous communication service score R within time step t. t , specifically expressed as follows: In the above formula, the instant communication service score quilt This is because if If the absolute value of is kept within 1, better convergence may occur; in addition, when β>1, different The reward difference between the values ​​is amplified, which can encourage the drone to take proactive measures when the energy is about to run out; however, β cannot be too large, because in the initial simulation, β ≥ 3 will lead to a lower final convergence value; in addition, in N T Within the time step, the cumulative reward is maximized by maximizing the cumulative communication service score; for each drone that crosses the boundary, runs out of energy, or collides with an obstacle, giving a negative reward as a penalty is an alternative reward function design; F in the above formula is a penalty flag, and the size of the penalty flag varies with the number of drones that cross the boundary, run out of energy, or collide with obstacles; if the drone leaves the target area during training, the current movement is canceled; in the current state S t Execute the current action A t Will result in negative rewards; In addition to the external rewards obtained by the interaction between the drone and the environment, the intrinsic curiosity module ICM can also provide specific internal rewards for the agent; The internal reward in time step t is Representation; ICM consists of two subsystems: an intrinsic reward generator that produces an internal motivational signal driven by curiosity; and a policy that produces a sequence of actions that optimize the reward signal. The goal of training the policy subsystem during the exploration-intensive phase is to make the total value of the two rewards r at time step t t Maximization; There are two models in ICM, namely the forward model and the inverse model; The forward model consists of a set of fully connected layers whose input is s t with a t The state of the connection is a t and t As input, we predict the state at time step t+1; the function g is called the forward model and is as follows: Among them, the state s t+1 The predicted estimated value of The parameters of the optimized neural network are defined as θ F , whose purpose is to minimize the following loss function: The inverse model uses a series of convolutional layers to map the state vector, which concatenates s t and t+1 ; Each convolutional layer is followed by an ELU nonlinearity; To predict a potential course of action, the vector is fed into a fully connected layer and then exits from this layer; Deep neural network training is equivalent to learning a function f, defined as: Among them, action a t The predicted estimated value of Define the parameters of the optimized neural network as θ I , and its optimization equation is: The loss function L I To measure actual movement t With predictive action The difference between tuples (s t ,s t+1 ) is obtained by using the current strategy π(s) to interact with the environment; the function f is called the inverse model; Make full use of the prediction error generated by the forward model as an intrinsic curiosity reward for training the agent's strategy, the internal reward The calculation formula is as follows: In step S5, the DDPG-ICM algorithm is used to train the network in the exploration-intensive phase; the actor network and critic network have their own target networks with the same structure, denoted as μ' and Q' respectively; the task of the actor network is to improve the performance of the agent by refining the deterministic policy, while the goal of the critic network is to estimate the action-value function of the policy received from the actor network; therefore, the state S t or S t+1 is the input of the actor network or the target actor network, action A t or A t+1 is its output; at the same time, the output of the critic network or the target critic network is based on S t or S t+1 and A t or A t+1 The action value function approximation Q(S t ,A t |θ Q ) or Q'(S t+1 ,A t+1 |θ Q′ ); store the experience in an experience replay buffer pool ε with a capacity of M; randomly extract small batches of size G < < M from the buffer pool to train the network; in the exploration-intensive phase, actions with higher curiosity values ​​are preferred; The development-intensive phase is designed to be implemented in real flight; moreover, the development-intensive phase and the exploration-intensive phase have the same actions, states, and network structures; however, starting from the trained development-intensive networks, these networks are trained to ensure a consistent initial exploration strategy instead of randomly initializing the network weights; at the same time, the intrinsic curiosity reward is removed from the reward function; the remaining steps are the same as the exploration-intensive phase; first, the actor network μ and the critic network Q are randomly initialized; their target networks μ' and Q' are initialized; in addition, the experience replay buffer ε is also initialized; the actor network observes the environment state S t , then determine action A t ; Before the agent does this, it t The drones will only stop moving if they leave the target area, run out of energy, or collide with an obstacle; otherwise, they will perform action A. t ; After that, the state enters S t+1 , the agent obtains r t As a bonus; there is a threshold N Thre Used to determine the stage of the training process; when the training round is less than or equal to the threshold N Thre When , the training is in the exploration-intensive stage, at this time, the reward value within time step t When the training round is greater than the threshold N Thre When , the training is in the development-intensive stage, and the reward value within time step t Finally, the network is trained to improve the performance of the system; once the training conditions are met, a small batch of samples is randomly drawn from the experience replay buffer ε; finally, the actor network, the critic network and their corresponding target networks are updated separately.

2. The multi-UAV communication system optimization control method based on intrinsic curiosity mechanism according to claim 1 is characterized by: The step S1 specifically includes: In the L×L target area, a set of N U Swarm of drones U To a group of N u Ground users N u Provide services; the vast majority of users are concentrated in several hot spots in the target area, and the remaining users are evenly distributed throughout the target area; the drones fly at a fixed altitude H in the target area, and each drone has a highly directional antenna, so the drone's transmission power is concentrated within the angle θ and provides the required minimum throughput to ground users; the drone's service range to the ground is circular, with a radius of If the user is not within the service range of the drone, no service will be provided; the drone consumes energy during flight, but energy can be replenished when flying in an illuminated area; if there is a building in front, the drone needs to bypass it.

3. The multi-UAV communication system optimization control method based on intrinsic curiosity mechanism according to claim 2 is characterized by: The step S2 specifically includes: All drones have a backhaul link connecting them to the external network. Since there is no spectrum overlap between the drone backhaul link and the drone user link, no mutual interference will occur. The path loss from drone k to ground user u is denoted as PL ku , calculated by the following formula: Where ω represents the center frequency of the spectrum allocated to user u, d ku represents the three-dimensional distance between user u and drone k, c represents the speed of light, η represents the additional path loss of the line-of-sight link, dB represents decibel, which is the unit of relative difference in signal strength; SINR of drone k to user u ku Calculated by the following formula: In the formula P t represents the UAV power spectrum density, n e represents the power spectral density of the ambient noise, represents the set of drones covering user u; R u is the minimum throughput required by each user; UAV k only The service will be provided to user u only when the following conditions are met: W ku log2(1+SINR ku )≥R u Where W ku represents the bandwidth allocated by drone k to user u. From the above formula, we can see that each user is only connected to the drone that has sufficient bandwidth and provides the highest signal-to-noise ratio; In time step t, the energy consumption of UAV k is It consists of three parts: (i) E F It means that the UAV k flies at a fixed speed v within time t. energy consumption; (ii) E S represents the energy consumption of the UAV signal transmitted on the backhaul link and the UAV-user link; (iii) E O represents the operating cost energy of the drone, which is proportional to time t; the energy consumption of drone k Calculated by the following formula: The initial battery charge of drone k is An episode's time segment is divided into N T time steps; each drone flies at a fixed speed v. Fly in any direction; within the time step t, the flight distance of UAV k is recorded as Energy consumption of drone horizontal flight P hor Calculated by the following formula: In the formula The weight of the drone is W in Newton, the air density is ρ, and the rotor area of ​​the drone is D. Due to the influence of the horizontal flight speed, the energy consumption of the drone during horizontal flight is lower than that during hovering. When drones fly in an illuminated area, the energy they gain is determined by many factors, the most important of which is the intensity of light radiation; the relationship between solar radiation intensity and harvested energy is calculated by the following formula: Where I(t) is the light radiation intensity, and its threshold is K c , η c is the efficiency; when this threshold is exceeded, the efficiency can be expressed as a constant.

4. The multi-UAV communication system optimization control method based on intrinsic curiosity mechanism according to claim 3 is characterized by: The step S3 specifically includes: In the formula, represents the horizontal coordinate of UAV k when flying in time step t, represents the total number of users successfully served within time step t; considering the number of users served by the drone and the necessary throughput, R t reflects the communication service quality within time step t; if user u is successfully served within time step t, then define is 1, otherwise it is 0; the number, location and battery status of drones together constitute the parameters of drones; the parameters of drones and several other parameters together determine The index β>0 is a factor that measures the agent’s overall user satisfaction based on the number of users successfully served.

Citation Information

Patent Citations

  • Unmanned aerial vehicle autonomous route planning algorithm based on state decomposition in complex environment

    CN116225055A

  • Multi-agent complex system reinforcement learning exploration method based on staged curiosity

    CN117709438A