Track planning and resource allocation method for multi-hop unmanned aerial vehicle with integrated communication and sensing
By combining multi-hop drone networks and ISAC systems, the flight trajectory and resource allocation of drones are optimized, and the energy consumption limitations of traditional drone systems in large-scale detection scenarios are solved, and the longer communication distance and larger detection range are achieved, and spectrum utilization and detection accuracy are improved.
Patent Information
- Application Number
- CN202510868823.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The traditional ISAC-based UAV collaborative detection system is limited by the UAV energy consumption problem in large-scale detection scenarios, with short communication distance and small detection range, making it difficult to apply to large-scale detection scenarios.
Combining the multi-hop drone network and ISAC system, the flight trajectory, power and channel allocation of the drone is optimized through the SAC algorithm, and the multi-hop drone network is used to expand the detection range and improve spectrum utilization.
The communication distance of the drone is expanded, the detection range and spectrum utilization of the system are improved, the communication energy consumption is reduced, and the detection accuracy and robustness of the system are enhanced.
Smart Images

Figure CN120371019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and more specifically, to a multi-hop UAV trajectory planning and resource allocation method for integrated communication and sensing. Background Art
[0002] As a flexible and efficient communication and sensing platform, the UAV network has become an important research direction in the fields of intelligent transportation, Internet of Things, environmental monitoring, etc. The multi-hop UAV network can expand the signal coverage range through relay transmission to achieve long-distance communication and sensing functions. The integrated sensing and communication technology (ISAC) is abbreviated as integrated communication and sensing. By performing sensing and communication in parallel in the same network, it further improves the resource utilization rate and function integration degree of the system. In addition, ISAC reduces the weight and complexity of hardware devices, which is suitable for UAVs to reduce the load when performing complex tasks. Therefore, researching how to combine ISAC technology with multi-hop UAV networks is of great significance for improving network performance and intelligence level.
[0003] The traditional UAV cooperative detection system based on ISAC consists of multiple UAVs equipped with communication and sensing units, a data fusion center, and multiple detection and sensing targets to be detected. Each UAV performs radar detection tasks and transmits the detection data to the fusion center through a communication link. This system optimizes the flight trajectory, channel, and power allocation for each UAV. The goal is to maximize the success rate of the detection tasks that the UAV cluster can perform during flight while meeting the communication quality. Through the deep reinforcement learning algorithm, each UAV autonomously plans its flight path to ensure coverage of all targets while avoiding collisions with other UAVs. Each UAV dynamically allocates the total power between radar detection and communication to meet the dual requirements of the system for communication data rate and radar detection accuracy. However, due to the UAV energy consumption problem, the coverage range of the system is small and the communication distance is short, making it difficult to apply to large-scale detection scenarios.
[0004] Chinese Patent Document CN116704823A discloses a method for resource allocation and UAV trajectory planning in an ISAC system. The system consists of M UAVs, K base stations, N targets, and a control center. The UAVs are equipped with communication and sensing modules. Starting from the initial positions, they fly to directly above the targets to sense target information and transmit the sensing tasks to the base stations, where the MEC servers deployed at the base stations process the tasks. The base stations transmit the position information of the UAVs to the control center. The control center generates control input commands according to the status of the UAVs and transmits them to the UAVs through the base stations. In this patent, the problem of the small movement range of individual UAVs is solved by the base station relay method in which the UAVs send sensing information to the base stations and the base stations forward it to the control center. However, it brings problems such as high energy consumption of the base stations, poor system robustness, and low portability. Summary of the Invention
[0005] The object of the present invention is to design and develop a multi-hop UAV trajectory planning and resource allocation method for integrated communication and sensing, which combines a multi-hop UAV network with an ISAC system and combines with the SAC algorithm, increasing the communication distance of the UAVs, expanding the detection range of the system, and improving the spectrum utilization rate of the system.
[0006] The technical solution provided by the present invention is as follows: A multi-hop UAV trajectory planning and resource allocation method for integrated communication and sensing includes the following steps: Step 1: Randomly deploy a UAV cluster in the detection area; Step 2: Based on the Cartesian coordinate system, the fusion center obtains the current position information; Wherein the current position information includes the position coordinates of the fusion center, the UAV cluster, and all detection targets; Step 3: The fusion center plans the next trajectory of the UAV cluster according to the current position information, constructs an optimization problem, and converts the optimization problem into a Markov decision process, and solves the optimization problem through the SAC algorithm to obtain the flight direction, flight distance, power, and channel of each UAV; Wherein, the optimization problem is: ; Step 4: Send the instructions of the fusion center to each UAV node according to the number of hops specified in the multi-hop network; Step 5: Each-hop UAV receives and fuses the instructions assigned by the center, and moves according to the planned trajectory and conducts target detection according to the assigned power; Step 6: Each-hop UAV forwards the current sensing information to the previous-hop node through the multi-hop network along the assigned channel until the fusion center receives all detection information; Step 7: Repeat the above detection process until the specified number of detections is reached, and end the current round of communication perception.
[0007] Preferably, the detection target At time The effective number of detections satisfies: ; Wherein, Represents up to time The number of UAVs that have successfully sensed the detection target And transmitted information.
[0008] Preferably, the Jain fairness index up to time Satisfies: ; Wherein, Is the Average detection frequency of the detection target at time .
[0009] Preferably, the position of the UAV At time Satisfies: ; Wherein, Is the abscissa of the position of the UAV At time , Is the ordinate of the position of the UAV At time .
[0010] Preferably, the signal-to-noise ratio of the UAV At time Satisfies: ; Wherein, Is the power allocated by the UAV To the detection function, Is the bandwidth of the UAV, Is the transmitting antenna gain, Is the receiving antenna gain, Is the operating wavelength, Is the target Radar cross section, Is the Boltzmann constant, Is the UAV And the detection target Distance between, Is the effective noise temperature, Is the radar noise factor, To detect loss.
[0011] Preferably, the communication link data rate of the drone satisfies: ; wherein, is the signal-to-noise ratio of the drone on the channel .
[0012] Preferably, the signal-to-noise ratio of the drone on the channel satisfies: ; wherein, is the communication noise factor, is the channel allocation coefficient of the drone at time , , is the average channel power gain between the drone and the previous-hop device , is the power allocated by the drone to the communication function, is the power allocated by the drone to the communication function.
[0013] Preferably, the average channel power gain between the drone and the previous-hop device satisfies: ; wherein, is a parameter of the traditional radar model, is the distance between the drone and its previous-hop device at time , is the attenuation factor of the line-of-sight link, is the attenuation factor of the non-line-of-sight link, is the line-of-sight probability between the drone and its previous-hop device at time , is the non-line-of-sight probability between the drone and its previous-hop device at time ; The distance between the drone and its previous-hop device at time The line-of-sight probability between them satisfies: ; In the formula, is the first constant, is the second constant, is the elevation angle between the UAV and its previous-hop device ; The non-line-of-sight probability between the UAV at time and its previous-hop device satisfies: ; The elevation angle between the UAV and its previous-hop device satisfies: ; In the formula, is the distance between the UAV and its previous-hop device , is the height of the previous-hop device of the UAV , is the height of the UAV at the position at time .
[0014] Preferably, the conversion of the optimization problem into a Markov decision process specifically includes: Action space: , ; State space: , ; Reward function: ; Among them, is the action of the UAV at time , is the observation space of the UAV at time , is the detection reward, is the incentive reward, is the set of detection targets; The detection reward satisfies: ; The incentive reward satisfies: ; In the formula, is the scaling factor, is the attenuation coefficient of the detection reward, is the UAV is a binary variable indicating whether the detection target is effectively detected .
[0015] Preferably, the SAC algorithm specifically includes the following steps: Step 1, initialize the Actor network parameters , the first Critic network parameters , the second Critic network parameters , the first target Critic network parameters and the second target Critic network parameters , and copy the two Critic network parameters to the two target Critic network parameters correspondingly. At the same time, initialize the experience replay buffer and the number of iterations; Step 2, initialize the number of iterations and time slots, and the number of training time slots in each round of iteration is 100; Step 3, obtain from the environment the state of the time slot and input it into the Actor network to obtain the action of the time slot ; Step 4, the UAV cluster executes the action to obtain the state of the next time slot and the reward of the current time slot ; Step 5, store in the experience replay buffer; Step 6, if the number of experiences stored in the experience replay buffer is greater than the preset minimum sampling batch size, randomly extract a set of experience samples from the experience replay buffer . For each sample, input the state of the time slot and the action of the time slot into the two Critic networks respectively to obtain Q values, and input the state at the next moment into the Actor network to generate the action at the next moment . Then, calculate the target return values of and through the first target Critic network and the second target Critic network respectively, and obtain the target Q value accordingly: ; ; In the formula, is the target of the time slot value is the target return value output by the first target Critic network is the target return value output by the second target Critic network is the discount factor is the entropy temperature parameter is the Actor network is the logarithm of the probability density of the action output by the Actor network Step 7: Update the parameters of the two Critic networks by gradient descent to minimize the loss values of the two Critic networks: ; In the formula, is the loss value of the Critic network represents the expected value of calculating the loss function for all experience samples sampled from the experience replay buffer ; Step 8: Update the Critic network parameters : ; In the formula, is the updated Critic network parameter is the learning rate is the loss function for the th critic network parameter ; Step 9: Calculate the loss function of the Actor network and the gradient of the loss function with respect to the Actor network parameter ; Among them, the loss function of the Actor network is: ; In the formula, represents the expectation of the joint distribution of the state and the action ; The gradient of the loss function with respect to the Actor network parameter is: ; In the formula, is the gradient of the policy distribution is the gradient with respect to the action is the action with respect to the parameter ; Step 10: Update the Actor network parameters, entropy temperature parameter, and the parameters of the two target Critic networks: ; ; ; wherein, is the updated Actor network parameter, is the learning rate, is the gradient of the policy function with respect to the policy network parameters, is the updated entropy temperature parameter, is the target entropy, is the updated target Critic network parameter, is the proportionality factor for soft update; Step 11: If , then directly jump to Step 3; If and , then directly jump to Step 2; If and , then the algorithm ends.
[0016] Advantages of the present invention: (1) A multi-hop UAV trajectory planning and resource allocation method for integrated communication and sensing designed and developed by the present invention uses a multi-hop UAV network to solve problems such as short communication distance and small detection and sensing range in traditional integrated communication and sensing systems, and intelligently allocates the power and channels of UAVs for communication and detection and sensing according to the current environment, improving the energy usage efficiency of UAVs and at the same time improving the spectrum utilization rate in the integrated communication and sensing system; (2) The multi-hop UAV trajectory planning and resource allocation method for integrated communication and sensing designed and developed by the present invention improves the spectrum efficiency by jointly optimizing power, channels, and the flight trajectories of UAV relays. On the basis of ensuring the detection quality of the original technical objectives, it expands the communication distance of UAVs, reduces communication energy consumption, enables more limited energy to be used for detecting targets, and improves the detection accuracy of the system; (3) A multi-hop UAV trajectory planning and resource allocation method for integrated communication and sensing designed and developed by the present invention combines the SAC algorithm with MDP to maintain the diversity of policies while ensuring convergence, thereby improving the robustness and convergence speed of learning. Description of the Drawings
[0017] Figure 1 is a schematic flow chart of the multi-hop UAV trajectory planning and resource allocation method for integrated communication and sensing according to the present invention.
[0018] Figure 2 Schematic diagram of the integrated communication and sensing multi-hop UAV trajectory planning and resource allocation method described in the present invention.
[0019] Figure 3 Schematic diagram of the curves of the number of iterations and the average reward of the environment of the SAC algorithm described in the present invention and other algorithms.
[0020] Figure 4 Schematic diagram of the curves of the number of iterations and the average effective detection number of the SAC algorithm described in the present invention and other algorithms. Detailed implementation manners
[0021] The following further describes the present invention in detail with reference to the accompanying drawings of the specification, so that those skilled in the art can implement it according to the text of the specification.
[0022] As Figure 1 shown, a method for integrated communication and sensing multi-hop UAV trajectory planning and resource allocation provided by the present invention includes: Step 1: As Figure 2 shown, randomly deploy a UAV cluster in the detection area; Among them, the detection area includes detection targets, UAVs; Step 2: As Figure 1 shown, based on the Cartesian coordinate system, the fusion center obtains the current position information; Among them, the position information includes the position coordinates of the fusion center itself, the UAV cluster, and all detection targets; Step 3: The fusion center plans the next trajectory of the UAV cluster according to the current position information, constructs an optimization problem, and allocates the flight direction, flight distance, power for sensing, and communication channels for each UAV; Among them, UAV moves at time with as the length and as the angle, where , assuming that all UAVs fly at the same height, given the current position coordinates of UAV as , where is the abscissa of the position of UAV at time , is the ordinate of the position of UAV at time , is the ordinate of the position of UAV at time The height of the position; at the next moment The position can be expressed by the following formula: ; Collision should be prevented between any two drones to avoid additional losses. Let the minimum safe distance between two drones be , then there is: ; Wherein, is the distance between drone and drone : ; In the formula, is the position coordinate of drone at time , drone at time position coordinate; The drone must also operate within a predefined detection area. Let the boundary of the detection area be , , the relationship between the position of the drone and the boundary of the detection area is as follows: ; In addition to the flight mission, each drone needs to detect the target and transmit the detection information to the fusion center through the communication link. After the fusion center processes the detection data, it sends the control instruction to the drone. Each drone uses the beam to perform dual functions: the main beam is used for target perception, and the secondary beam is used for communication. The target detection model of the radar is as follows: The total power of the drone is , , representing the power distribution coefficient. The power allocated to the main beam is , and the power allocated to the secondary beam is . The drone includes a total bandwidth and channels, and the bandwidth of each channel is . Assuming , in order to avoid interference between sensing and communication, each drone is assigned a dedicated sensing channel. Therefore, the set of channels allocated for communication is , the binary variable indicates whether the communication channel is allocated to drone , where , , is the set of drones, , specifically, if the communication channel is assigned to the UAV , then , otherwise .
[0023] Within the detection area, each UAV emits electromagnetic waves into the environment and identifies potential targets by analyzing the reflected signals. The UAV estimates the distance based on the round-trip time delay of the signals and uses statistical hypothesis testing to determine whether the reflected signals correspond to actual targets. Since reflection, diffraction, and scattering generate multipath interference, selective interference mitigation techniques are required to improve the detection performance. In this case, it is assumed that the distance estimation is mainly affected by environmental noise, which is simulated as an additive white Gaussian noise process. To quantify the accuracy of distance estimation, the Cramér-Rao lower bound (CRLB) is used, which provides a theoretical lower bound for the variance of an unbiased estimator. According to modern radar theory, the CRLB of frequency and phase parameters is inversely proportional to the signal-to-noise ratio (SNR), indicating that increasing the SNR can effectively reduce the variance of target detection.
[0024] UAV at time The SNR is expressed as:[[]] ; In the formula, is the power allocated to the detection function by the UAV , is the bandwidth of the UAV, is the transmitting antenna gain, is the receiving antenna gain, is the operating wavelength, is the target 's radar cross section, is the Boltzmann constant, is the UAV and the detection target the distance between them, is the effective noise temperature, is the radar noise figure, is the detection loss; And the operating wavelength satisfies: ; In the formula, is the speed of light, is the frequency of the radar.
[0025] In this embodiment, the effective noise temperature is .
[0026] Therefore, to ensure the detection quality, the UAV at time detects the target The following conditions must be satisfied: ; In the formula, is the minimum signal-to-noise ratio when the UAV ensures effective detection.
[0027] The probabilistic path loss model is used to describe the communication between the UAV and the fusion center, and between UAVs. Therefore, the line-of-sight (LoS) probability between the UAV and its previous-hop device at time is: ; In the formula, is the first constant, is the second constant, is the elevation angle between the UAV and its previous-hop device ; Among them, the elevation angle between the UAV and its previous-hop device satisfies: ; In the formula, is the height of the previous-hop device of the UAV , is the distance between the UAV and its previous-hop device ; The non-line-of-sight (NLoS) probability between the UAV and its previous-hop device at time is: ; Thus, the average channel power gain between the UAV and its previous-hop device is calculated as: ; In the formula, is the average channel power gain between the UAV and its previous-hop device , is a parameter of the traditional radar model, is the distance between the UAV and its previous-hop device , and represent the attenuation factors of the line-of-sight and non-line-of-sight links respectively; Based on the transmission power and channel gain, the unmanned aerial vehicle (UAV) on the channel the signal-to-noise ratio (SINR) calculation formula is: ; In the formula, is the signal-to-noise ratio of the UAV on the channel , is the communication noise factor, is the channel allocation coefficient of the UAV at time , , is the power allocated by the UAV to the communication function, is the power allocated by the UAV to the communication function.
[0028] According to the Shannon capacity formula, the communication link data rate of the UAV is: ; To ensure reliable data transmission, the minimum data rate of the UAV needs to meet the following conditions: ; To ensure the detection accuracy, each target must be sensed multiple times. The effective detection of the detection target at time is denoted as , and it satisfies: ; In the formula, represents the number of UAVs that have successfully sensed the detection target and transmitted information by time ; The condition requirement for effective detection is that for the UAV and time , it satisfies , , A measure of the detection effectiveness of a quantized target is the effective detection frequency of the target by time , which is expressed as: ; To address the potential geographical unfairness problem, that is, targets located in dense areas may be detected multiple times, while sparse targets located in edge areas may be ignored, the present invention introduces the Jain fairness index, which measures the detection fairness among all targets by time , and is expressed as: ; wherein is the Jain fairness index up to time ; is the average detection frequency of the th detection target at time ; The goal of this multi-hop ISAC system is to maximize the effective detection frequency and geographical fairness of all targets while ensuring the communication and detection quality within time slots. To this end, the system is restricted by four conditions: the safe distance of the UAVs, the regional scope, the minimum signal-to-noise ratio, and the minimum communication power. It is necessary to optimize the flight direction, flight distance, power allocation, and channel allocation of the UAVs. The corresponding optimization problem is formulated as: ; wherein is the number of time slots, is the Jain fairness index up to time ; is the effective detection times of the detection target at time ; is the number of detection targets, is the distance between UAV and UAV ; is the minimum safe distance between two UAVs, is the abscissa boundary of the detection area, is the abscissa of the position of UAV at time ; is the ordinate boundary of the detection area, is the ordinate of the position of UAV at time ; is the signal-to-noise ratio of UAV at time ; is the minimum signal-to-noise ratio threshold for UAV to reach effective detection, is the communication link data rate of UAV ; is the minimum communication link data rate threshold for UAV to reach effective detection, is the moving length of UAV at time ; , is the maximum moving length of UAV ; is the UAV The moving angle at time , is the power allocation coefficient of the UAV at time . is the number of the communication channel, is the set of channels allocated by the fusion center for communication, is the channel allocation coefficient of the UAV at time ; Among them, the constraints in (19e) in the optimization problem ensure that each UAV is exactly allocated one communication channel at each moment, preventing interference and optimizing bandwidth utilization. The constraints in formulas (19a), (19b), (19c), (19d) and (19e) respectively define the flight distance, direction, power allocation and channel allocation of the UAVs to ensure the optimization of system performance.
[0029] In this embodiment, the SAC algorithm is combined with MDP to solve the optimization problem; Among them, to solve the proposed problem model, the system is constructed as a Markov decision process (MDP), which is a widely used framework in reinforcement learning. The previously described radar sensing and communication environment is the learning environment, and the fusion center, as the decision maker, is responsible for trajectory planning, power allocation and channel allocation. It transmits decisions to each UAV through the network. This structure allows for a formulated representation of the decision process, enabling the UAVs to make optimal decisions based on their observations and past actions. The MDP is defined by the tuple , where is described as follows: Action space : The optimization problem includes determining the flight direction, flight distance, power allocation coefficient and channel allocation of each UAV. Therefore, the action of the UAV at time is defined as a vector of its flight parameters and resource allocation decisions, expressed as: ; The action space includes the actions of all UAVs in the system. Therefore, the system state at time is the set of the states of all UAVs .
[0030] State space : Each UAV makes local observations of the environment and collects information related to its current state and previous actions. These factors affect the current state. Therefore, the observation space of the UAV at time is defined as follows: ; Among them, and represent the position of the UAV at time . The remaining variables reflect the actions taken in the previous time slot; The state space includes the observation data of all UAVs in the system. Therefore, the system state at time is the set of the states of all UAVs .
[0031] Reward function : The reward function aims to improve the target detection efficiency and encourage geographical fairness at the same time. It consists of two main parts: detection reward and incentive reward; The described detection reward reflects the contribution of the UAV at time to the overall detection and perception task. This reward is aggregated by the fusion center for all the detection information of the UAVs to ensure the collaborative performance. Its calculation formula is: ; The described incentive reward is to further improve the system performance and ensure fairness. This reward will prevent the UAVs from repeatedly detecting the same target and encourage them to explore the targets that are less detected or not detected. If the UAV perceives the target within the time , then the incentive reward of this UAV is expressed as: ; In the formula, is the scaling factor, which is used to adjust the reward amplitude; is the discount factor, which is used to reduce the reward of the frequently detected targets, is the attenuation coefficient of the detection reward; is a binary variable, which represents whether the UAV effectively detects the detection target . If the UAV detects the target at time , then , otherwise .
[0032] To sum up, the calculation formula of the total reward is: ; This reward structure can not only encourage the UAVs to contribute to the overall system performance, but also ensure the fair detection of targets in the entire sensing area.
[0033] In reinforcement learning, Soft Actor-Critic (SAC) is an algorithm based on policy gradients. It adopts a reinforcement learning framework based on maximum entropy, aiming to optimize both the policy and the value function simultaneously to achieve efficient exploration and more stable learning. Compared with traditional policy gradient methods, SAC combines the Actor-Critic method with the principle of maximum entropy in continuous action spaces by introducing an entropy term, maximizing the expected cumulative reward and policy entropy, thus encouraging exploration and preventing convergence to suboptimal solutions. It is particularly suitable for solving the proposed formulation problems, encouraging the agent to perform more exploratory behaviors, thereby improving the robustness and convergence speed of learning.
[0034] The SAC algorithm includes an Actor network, a first Critic network, a second Critic network, a first target Critic network, and a second target Critic network, specifically including: (1) Initialize the parameters of the Actor network , the parameters of the first Critic network , the parameters of the second Critic network , the parameters of the first target Critic network and the parameters of the second target Critic network , and copy the parameters of the two Critic networks to the parameters of the two target networks respectively. At the same time, initialize the experience replay buffer and the number of iterations .
[0035] (2) Initialize the number of iterations and the time slot , and the number of time slots per round of training is 100, which also represents 100 detection times.
[0036] (3) In the data collection stage, the agent interacts with the environment to collect training data. Specifically, the agent obtains the current state from the environment and uses it as the input of the policy network (Actor). The Actor network outputs the action at the current moment based on this state. Subsequently, the UAV swarm executes this action to update the environment, and then returns the state at the next moment and the reward at the current moment. During the training process, the quadruple generated by each round of interaction is stored in the experience pool . After the experience pool collects 1024 pieces of data, it serves as a buffer to facilitate randomly sampling from it during subsequent training, thereby promoting the model to learn a more generalized policy.
[0037] (4) When updating the two Critic networks, a batch of experience samples is randomly drawn from the experience pool . The current state and action of each sample are used as the inputs of the two Critic networks respectively, and the expected return of the corresponding action, that is, value, is denoted as , . To train the Critic network, the target Critic network is needed to obtain the target return value. First, the next moment state is input into the Actor network to obtain the optimal action at the next moment state. Then, and are input into the two target Critic networks respectively to obtain the target return values, denoted as , . Select the minimum value among , as the value of the next state, and calculate the target value accordingly. The calculation formula for the target value in the time slot is: ; In the formula, is the discount factor, is the entropy temperature parameter, is the Actor network, is the logarithm form of the probability density of the action output by the Actor network; Then, the parameters of the first Critic network and the parameters of the second Critic network are updated using the gradient descent method. Specifically, by comparing the value predicted by the evaluation network and the target value , the mean square error can be calculated, and then the backpropagation algorithm is used to update the parameters and of the two Critic networks to minimize this loss value, that is: ; In the formula, is the loss value of the Critic network, represents the expected value of calculating the loss function for all experience samples sampled from the experience replay buffer ; (5) Update the parameters of the Critic network : ; In the formula, is the updated Critic network parameter, is the learning rate, is the loss function for the critic network parameter gradient; (6) Calculate the loss function of the Actor network and the gradient of the loss function with respect to the Actor network parameter gradient; Among them, the loss function of the Actor network is: ; In the formula, represents the expectation of the joint distribution of the state and the action ; The gradient of the loss function with respect to the Actor network parameter is: ; In the formula, is the gradient of the policy distribution, is the gradient with respect to the action gradient, is the action gradient with respect to the parameter gradient.
[0038] (7) Then update the parameters of the Actor network according to the gradient descent method: ; In the formula, is the updated Actor network parameter, is the learning rate, is the gradient of the policy function with respect to the policy network parameter; (8) In order to automatically adjust the entropy of the policy and achieve a dynamic balance between exploration and exploitation, the entropy temperature parameter needs to be updated: ; In the formula, is the updated entropy temperature parameter, is the target entropy; (9) Apply a soft update strategy to the first target Critic network and the second target Critic network , that is: ; In the formula, For the updated target Critic network parameters, is the scaling factor for soft update; (10) Increment the time slot by 1, and determine whether the number of time slots per round of training reaches 100; if not, return to step 3 to continue training; if so, determine whether all algorithm iteration times are completed; if not, increment the iteration count by 1, re-initialize the time slot , and return to step 3 to continue training; if all iteration times are completed, the algorithm ends.
[0039] Through continuous iteration of the above steps, the SAC algorithm can continuously optimize the policy network and value network during training, and finally obtain an effective decision-making strategy.
[0040] In this embodiment, the soft update scaling factor is 0.005, and the number of iterations is 15000.
[0041] As described above, the specific process of the SAC algorithm is shown in Table 1: Table 1 SAC algorithm process
[0042] Step 4: Send the instructions from the fusion center to each UAV node according to the specified number of hops in the multi-hop network; Step 5: Each hop of UAV receives and executes the instructions assigned by the fusion center, that is, moves according to the planned trajectory and performs target detection according to the assigned power; Step 6: Each hop of UAV forwards the current sensing information to the previous hop node through the multi-hop network along the assigned channel until the fusion center receives all the detection information; Step 7: Repeat the above detection process until the specified number of detections is reached, and end this round of communication sensing.
[0043] In this embodiment, it includes 1 fusion center, 3 UAVs and 35 targets. The number of available channels is 7. The movement range of all UAVs is restricted within the area meters, the flight altitude is 100 meters, the maximum flight distance of each UAV is 100 meters, the minimum safety distance between UAVs is 50 meters, and the altitude of the fusion center is 5 meters. For radar and communication functions, the maximum transmission power of each UAV is 0.1W, the bandwidth of each channel is 1MHz, and the operating frequency is 3GHz; The parameter settings related to the sensor are as follows: , , , , The parameter settings related to communication are as follows: , , , , . The thresholds for sensing and communication are set to and .
[0044] The present invention uses an advanced reinforcement learning algorithm, namely the SAC algorithm, which is more efficient than traditional convex optimization and swarm intelligence algorithms and has better training results. As shown in Figure 3 and Figure 4 , the superiority of the SAC algorithm is reflected in two aspects: the reward value of the system and the number of effective detections. During the training process, both the average reward and the average number of effective detections of the SAC algorithm are higher than those of other reinforcement learning algorithms, ensuring convergence while maintaining the diversity of the policy.
[0045] A multi-hop UAV trajectory planning and resource allocation method for integrated communication and sensing designed and developed by the present invention is different from the traditional ISAC and UAV combined system. It increases the UAV communication distance, expands the detection range of the system, improves the spectrum utilization rate of the system, better plays the advantages of integrated communication and sensing, has strong portability, only requires a fusion center and multiple UAVs, is easy to deploy and saves system overhead, and is more suitable for large-scale and energy-constrained collaborative detection environments.
[0046] Although the embodiments of the present invention have been disclosed above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the embodiments shown and described here.
Claims
1. A multisensory integrated multi-hop UAV trajectory planning and resource allocation method, characterized in that It includes the following steps: Step 1: Randomly deploy a drone swarm in the detection area; Step 2: Based on the Cartesian coordinate system, the fusion center obtains the current position information; wherein the current position information includes the position coordinates of the fusion center, the drone swarm, and all detection targets; Step 3: The fusion center plans the next trajectory of the drone swarm according to the current position information, constructs an optimization problem, converts the optimization problem into a Markov decision process, and solves the optimization problem through the SAC algorithm to obtain the flight direction, flight distance, power, and channel of each drone; wherein, the optimization problem is: ; Step 4: Send the instructions of the fusion center to each drone node according to the number of hops specified in the multi-hop network; Step 5: Each hop of drones receives and fuses the instructions assigned by the center, moves according to the planned trajectory, and performs target detection according to the assigned power; Step 6: Each hop of drones forwards the current perception information to the previous hop node through the multi-hop network along the assigned channel until the fusion center receives all detection information; Step 7: Repeat the above detection process until the specified number of detections is reached, and end this round of communication perception.
2. The integrated communication and sensing multi-hop UAV trajectory planning and resource allocation method according to claim 1, characterized in that The detection target At the moment The effective detection times satisfy: ; In the formula, represents the number of UAVs that have successfully detected the detection target and transmitted information by time 3. The integrated communication and sensing multi-hop UAV trajectory planning and resource allocation method according to claim 2, wherein, The Jain fairness index up to the time instant satisfies: ; Wherein, is the th average detection frequency of the detection target at the moment .
4. The integrated sensing and communication multi-hop UAV trajectory planning and resource allocation method according to claim 3, wherein, The drone At the moment The position satisfies: ; In the formula, is the abscissa of the position of the unmanned aerial vehicle at time , and is the ordinate of the position of the unmanned aerial vehicle at time .
5. The multi-hop UAV trajectory planning and resource allocation method for integrated communication and sensing as claimed in claim 4, wherein The drone At the moment The signal-to-noise ratio satisfies: ; Wherein, is the power allocated to the detection function of the drone ; is the bandwidth of the drone ; is the transmitting antenna gain ; is the receiving antenna gain ; is the operating wavelength ; is the radar cross section of the target ; is the Boltzmann constant ; is the effective noise temperature 6. The integrated communication and sensing multi-hop UAV trajectory planning and resource allocation method according to claim 5, wherein The drone has a communication link data rate that satisfies: ; Wherein, is the unmanned aerial vehicle on the channel signal-to-noise ratio.
7. The integrated communication and sensing multi-hop UAV trajectory planning and resource allocation method according to claim 6, characterized in that, The UAV on the channel has a signal-to-noise ratio that satisfies: ; Wherein, is the communication noise factor, is the UAV at time channel allocation coefficient, , is the UAV and the previous-hop device the average channel power gain between, is the UAV the power allocated to the communication function, is the UAV the power allocated to the communication function.
8. The integrated communication and sensing multi-hop UAV trajectory planning and resource allocation method according to claim 7, wherein The drone and the previous-hop device have an average channel power gain that satisfies: ; In the formula, are the parameters of the traditional radar model, is the distance between the UAV and its previous-hop device at time , is the attenuation factor of the line-of-sight link, is the attenuation factor of the non-line-of-sight link, is the line-of-sight probability between the UAV and its previous-hop device at time , is the non-line-of-sight probability between the UAV and its previous-hop device at time ; At the moment the drone and its previous-hop device the line-of-sight probability between them satisfies: ; In the formula, is the first constant, is the second constant, is the UAV and its previous-hop device the elevation angle between them; At the moment drone and its previous-hop device the non-line-of-sight probability between them satisfies: ; The drone and its previous-hop device have an elevation angle that satisfies: ; In the formula, is the distance between the UAV and its previous-hop device , is the height of the previous-hop device of the UAV , is the height of the UAV at the position at time .
9. The integrated communication and sensing multi-hop UAV trajectory planning and resource allocation method according to claim 8, characterized in that, Converting the optimization problem into a Markov decision process specifically includes: Action space: , ; State space: , ; Reward function: ; wherein, is the action of the UAV at time ; is the observation space of the UAV at time ; is the detection reward is the incentive reward is the set of detection targets; The detection reward satisfies: ; The incentive reward satisfies: ; In the formula, is the scaling factor, is the attenuation coefficient of the detection reward, is the drone whether the detection target is effectively detected is a binary variable.
10. The integrated communication and sensing multi-hop UAV trajectory planning and resource allocation method according to claim 9, characterized in that, The SAC algorithm specifically includes the following steps: Step 1, initialize the Actor network parameters , the first Critic network parameters , the second Critic network parameters , the first target Critic network parameters and the second target Critic network parameters , and copy the two Critic network parameters to the two target Critic network parameters correspondingly. At the same time, initialize the experience replay buffer and the number of iterations; Step 2: Initialize the number of iterations and time slots, and the number of training time slots in each round of iteration is 100; Step 3: Obtain from the environment the status of the time slot input it into the Actor network to obtain the action of the time slot ; Step 4, the drone swarm executes actions , to obtain the state of the next time slot and the reward of the current time slot ; Step 5, store in the experience replay buffer; Step 6: If the number of experiences stored in the experience replay buffer is greater than the preset minimum sampling batch size, then randomly draw a set of experience samples from the experience replay buffer . For each sample, input the state of the time slot and action of the time slot into two Critic networks respectively to obtain Q values, and input the state at the next moment into the Actor network to generate the action at the next moment . Then, calculate the target return values through the first target Critic network and the second target Critic network respectively for and respectively, and obtain the target Q values accordingly: ; Wherein, is the target value of the time slot, is the target return value output by the first target Critic network, is the target return value output by the second target Critic network, is the discount factor, is the entropy temperature parameter, is the Actor network, is the logarithmic form of the probability density of the action output by the Actor network; Step 7: Update the parameters of the two Critic networks by the gradient descent method to minimize the loss values of the two Critic networks: ; In the formula, is the loss value of the Critic network, represents the expected value of calculating the loss function for all experience samples sampled from the experience replay buffer; Step 8, update the parameters of the Critic network : ; Wherein, are the updated Critic network parameters, is the learning rate, is the loss function for the th critic network parameter gradient; Step 9, calculate the loss function of the Actor network and the gradient of the loss function with respect to the parameters of the Actor network ; wherein, the loss function of the Actor network is: ; wherein, represents the expectation of the joint distribution of states and actions; The gradient of the loss function with respect to the Actor network parameters is: ; In the formula, is the gradient of the policy distribution, is the gradient with respect to the action , is the action with respect to the parameter ; Step 10: Update the Actor network parameters, entropy temperature parameters, and the parameters of the two target Critic networks: ; ; ; In the formula, is the updated Actor network parameter, is the learning rate, is the gradient of the policy function with respect to the policy network parameter, is the updated entropy temperature parameter, is the target entropy, is the updated target Critic network parameter, is the scaling factor for soft update; Step 11. If , then directly jump to Step 3; If and , directly jump to Step 2; If and , the algorithm ends.
Citation Information
Patent Citations
Unmanned aerial vehicle intelligent trajectory planning and communication resource allocation method based on reinforcement learning
CN116704823A
Resource allocation method, system and equipment and computer medium
CN109769292A
Trajectory optimization and resource allocation method of unmanned aerial vehicle multi-hop relay communication system
CN110380773A
Radar and communication integrated unmanned aerial vehicle cooperative multi-target detection method
CN114679729A
Trajectory optimization and resource allocation method in multi-machine collaborative Internet of Vehicles
CN116233791A