A synaesthesia-integrated multi-hop UAV trajectory planning and resource allocation method
By optimizing drone trajectories and resource allocation through multi-hop drone networks and the SAC algorithm, the problems of short communication distance and small detection range in traditional ISAC systems are solved, large-scale detection and efficient spectrum utilization are achieved, and system performance and robustness are improved.
Patent Information
- Application Number
- CN202510868823.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The traditional ISAC-based UAV collaborative detection system is subject to the energy consumption problem of UAVs, resulting in short communication distance and small detection range, making it difficult to apply to large-scale detection scenarios, and the base station energy consumption and system robustness are poor.
A multi-hop UAV trajectory planning and resource allocation method with integrated synesthesia is designed. Combining the SAC algorithm, the communication and detection power and channels of the UAV are intelligently allocated through the multi-hop UAV network and the ISAC system. The Markov decision process is used to optimize the UAV flight trajectory and resource allocation, thereby expanding the detection range and improving the spectrum utilization.
It increases the communication distance of UAVs, expands the detection range, reduces communication energy consumption, improves the detection accuracy and spectrum utilization of the system, and enhances the robustness and convergence speed of learning.
Smart Images

Figure CN120371019B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of communication technology, and more particularly to a multi-hop UAV trajectory planning and resource allocation method integrating synaesthesia. Background Art
[0002] As a flexible and efficient communication and perception platform, drone networks have become an important research topic in fields such as intelligent transportation, the Internet of Things, and environmental monitoring. Multi-hop drone networks can extend signal coverage through relay transmission, enabling long-range communication and perception. Integrated Sensing and Communication (ISAC), also known as ISAC, further improves system resource utilization and functional integration by performing both perception and communication in parallel within the same network. Furthermore, ISAC reduces the weight and complexity of hardware, making it suitable for alleviating the load when drones perform complex missions. Therefore, research on how to integrate ISAC technology with multi-hop drone networks is crucial for improving network performance and intelligence.
[0003] The traditional ISAC-based UAV collaborative detection system consists of multiple UAVs equipped with communication and perception units, a data fusion center, and multiple targets to be detected. Each UAV performs radar detection tasks and transmits detection data to the fusion center through a communication link. The system optimizes the flight trajectory, channel, and power allocation of each UAV. The goal is to meet the communication quality while maximizing the success rate of the UAV cluster's detection tasks during flight. Through deep reinforcement learning algorithms, each UAV autonomously plans its flight path to ensure coverage of all targets while avoiding collisions with other UAVs. Each UAV dynamically allocates total power between radar detection and communication to meet the system's dual requirements for communication data rate and radar detection accuracy. However, due to the energy consumption of UAVs, the system has a small coverage range and short communication distance, making it difficult to apply to large-scale detection scenarios.
[0004] Chinese patent document CN116704823A discloses an ISAC system resource allocation and drone trajectory planning method. The system consists of M drones, K base stations, N targets, and a control center. The drones, equipped with communication and perception modules, start from a starting position, fly directly above the target to perceive target information, and transmit the perception task to the base station, where the MEC server performs task processing. The base station transmits the drone's location information to the control center. The control center generates control input commands based on the drone's status and transmits them to the drone via the base station. This patent solves the problem of the limited range of individual drones by using a base station relay method where drones send perception information to the base station, which then forwards it to the control center. However, this method introduces problems such as energy consumption of the base station, poor system robustness, and low portability. Summary of the Invention
[0005] The purpose of this invention is to design and develop a multi-hop UAV trajectory planning and resource allocation method with integrated synesthesia, which combines the multi-hop UAV network with the ISAC system and combines it with the SAC algorithm to increase the UAV communication distance, expand the detection range of the system, and improve the spectrum utilization of the system.
[0006] The technical solution provided by the present invention is:
[0007] A multi-hop UAV trajectory planning and resource allocation method with integrated synaesthesia includes the following steps:
[0008] Step 1: Randomly deploy drone clusters within the detection area;
[0009] Step 2: Based on the Cartesian coordinate system, the fusion center obtains the current position information;
[0010] The current location information includes the location coordinates of the fusion center, the UAV cluster, and all detected targets;
[0011] Step 3: The fusion center plans the next trajectory of the UAV cluster based on the current location information, constructs an optimization problem, and converts the optimization problem into a Markov decision process. The optimization problem is solved using the SAC algorithm to obtain the flight direction, flight distance, power, and channel of each UAV.
[0012] The optimization problem is:
[0013] ;
[0014] Step 4: Send the fusion center's instructions to each drone node according to the number of hops specified in the multi-hop network;
[0015] Step 5: Each hop UAV receives and integrates the instructions assigned by the center, moves according to the planned trajectory, and detects the target according to the allocated power;
[0016] Step 6: Each hop UAV forwards the current sensing information to the previous hop node along the assigned channel through the multi-hop network until the fusion center receives all the detection information;
[0017] Step 7: Repeat the above detection process until the specified number of detections is reached, ending this round of communication perception.
[0018] Preferably, the detection target At the moment The effective number of detections meets the following requirements:
[0019] ;
[0020] Where, Indicates the time Until now, the detection target has been successfully sensed and the number of drones transmitting information.
[0021] Preferably, the time The Jain fairness index so far satisfies:
[0022] ;
[0023] Where, For the The detected target at time The average detection frequency.
[0024] Preferably, the drone At the moment The position meets:
[0025] ;
[0026] Where, For drones At the moment The horizontal coordinate of the position, For drones At the moment The vertical coordinate of the position.
[0027] Preferably, the drone At the moment The signal-to-noise ratio satisfies:
[0028] ;
[0029] Where, For drones The power allocated to the detection function, is the bandwidth of the drone, is the transmitting antenna gain, is the receiving antenna gain, is the working wavelength, Target radar cross section, is the Boltzmann constant, It's a drone and detection targets The distance between is the effective noise temperature, is the radar noise figure, To detect loss.
[0030] Preferably, the drone The communication link data rate satisfies:
[0031] ;
[0032] Where, For drones In the channel The signal-to-noise ratio on .
[0033] Preferably, the drone In the channel The signal-to-noise ratio on satisfies:
[0034] ;
[0035] Where, is the communication noise factor, For drones At the moment The channel allocation coefficient, , For drones With the previous hop device The average channel power gain between For drones The power allocated to communication functions, For drones Power allocated to communication functions.
[0036] Preferably, the drone With the previous hop device The average channel power gain between satisfies:
[0037] ;
[0038] Where, are the parameters of the traditional radar model, For the moment drones The device on the previous hop The distance between is the attenuation factor of the line-of-sight link, is the attenuation factor of the non-line-of-sight link, For the moment drones The device on the previous hop The probability of sight distance between For the moment drones The device on the previous hop The probability of non-line-of-sight between
[0039] At the moment drones The device on the previous hop The probability of the line of sight between them satisfies:
[0040] ;
[0041] Where, is the first constant, is the second constant, For drones The device on the previous hop The elevation angle between
[0042] At the moment drones The device on the previous hop The non-line-of-sight probability between them satisfies:
[0043] ;
[0044] The drone The device on the previous hop The elevation angle between them satisfies:
[0045] ;
[0046] Where, For drones The device on the previous hop The distance between For drones Previous hop device height, For drones At the moment The height of the location.
[0047] Preferably, converting the optimization problem into a Markov decision process specifically includes:
[0048] Action Space: , ;
[0049] State Space: , ;
[0050] Reward function: ;
[0051] in, For drones At the moment The action of time, For drones At the moment The observation space, For detection rewards, To motivate rewards, is a collection of detection targets;
[0052] The detection reward satisfies:
[0053] ;
[0054] The incentive reward meets the following requirements:
[0055] ;
[0056] Where, is the scaling factor, is the decay coefficient of the detection reward, For drones Whether the detection target is effectively detected binary variable.
[0057] Preferably, the SAC algorithm specifically includes the following steps:
[0058] Step 1. Initialize Actor network parameters , First Critic network parameters , Second Critic Network Parameters , the first target Critic network parameters and the second target Critic network parameters , and copy the two critic network parameters to the two target critic network parameters, and initialize the experience replay buffer and the number of iterations;
[0059] Step 2: Initialize the number of iterations and time slots, and the number of training time slots in each iteration is 100;
[0060] Step 3: Get from the environment Time slot status Enter the Actor network and get Time slot action ;
[0061] Step 4: Drone swarm performs actions , get the state of the next time slot and the reward for the current time slot ;
[0062] Step 5: Stored in the experience replay buffer;
[0063] Step 6: If the number of experiences stored in the experience replay buffer is greater than the preset minimum sampling batch size, A set of experience samples are randomly selected from the set, and for each sample, Time slot status and Time slot action Input two Critic networks respectively, obtain the Q value, and the state of the next moment Input Actor network to generate the action at the next moment , and then 、 The target return value is calculated through the first target critic network and the second target critic network respectively, and the target Q value is obtained accordingly:
[0064] ;
[0065] Where, for Time slot target value, is the target return value output by the first target Critic network, is the target return value output by the second target Critic network, is the discount factor, is the entropy temperature parameter, For the Actor network, Output the logarithmic form of the probability density of the action for the Actor network;
[0066] Step 7. Update the parameters of the two critic networks by gradient descent to minimize the loss values of the two critic networks:
[0067] ;
[0068] Where, For the The loss value of the Critic network, To represent the experience replay buffer All experience samples obtained from Calculate the expected value of the loss function;
[0069] Step 8. Update Critic network parameters :
[0070] ;
[0071] Where, is the updated Critic network parameter, is the learning rate, is the loss function For the first critic network parameters gradient;
[0072] Step 9. Calculate the loss function of the Actor network and the effect of the loss function on the Actor network parameters gradient;
[0073] Among them, the loss function of the Actor network is:
[0074] ;
[0075] Where, Indicates status and actions The expectation of the joint distribution of
[0076] The loss function is used to determine the Actor network parameters The gradient of is:
[0077] ;
[0078] Where, is the gradient of the policy distribution, It's the action Gradient It's action Parameters gradient;
[0079] Step 10. Update the Actor network parameters, entropy temperature parameters, and two target Critic network parameters:
[0080] ;
[0081] ;
[0082] ;
[0083] Where, is the updated Actor network parameter, is the learning rate, is the gradient of the policy function to the policy network parameters, is the updated entropy temperature parameter, is the target entropy, After the update Target Critic network parameters, is the scaling factor for soft updates;
[0084] Step 11: If , then jump directly to step 3;
[0085] like and , then jump directly to step 2;
[0086] like and , the algorithm ends.
[0087] The beneficial effects of the present invention are:
[0088] (1) The present invention designs and develops a multi-hop UAV trajectory planning and resource allocation method for integrated communication and perception. This method utilizes a multi-hop UAV network to solve the problems of short communication distance and small detection and perception range in traditional integrated communication and perception systems. It also intelligently allocates the power and channels of the UAV for communication and detection and perception according to the current environment, thereby improving the energy efficiency of the UAV and the spectrum utilization rate in the integrated communication and perception system.
[0089] (2) The multi-hop UAV trajectory planning and resource allocation method designed and developed by the present invention improves spectrum efficiency by jointly optimizing power, channels, and flight trajectories of UAV repeaters. While ensuring the target detection quality of the original technology, it expands the communication distance of the UAV, reduces communication energy consumption, and allows limited energy to be used more for target detection, thereby improving the detection accuracy of the system.
[0090] (3) The present invention designs and develops a multi-hop UAV trajectory planning and resource allocation method that integrates synaesthesia. By combining the SAC algorithm with MDP, the method maintains the diversity of strategies while ensuring convergence, thereby improving the robustness and convergence speed of learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 This is a flow chart of the synaesthesia-integrated multi-hop UAV trajectory planning and resource allocation method of the present invention.
[0092] Figure 2 Schematic diagram of the synaesthesia-integrated multi-hop UAV trajectory planning and resource allocation method of the present invention.
[0093] Figure 3 Schematic diagram of the curve of the number of iterations and the average reward of the environment for the SAC algorithm described in the present invention and other algorithms.
[0094] Figure 4 Schematic diagram of the curve of the number of iterations and the average effective detection number of the SAC algorithm described in the present invention and other algorithms. DETAILED DESCRIPTION
[0095] The present invention will be further described below in detail with reference to the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.
[0096] like Figure 1 As shown, the present invention provides a multi-hop UAV trajectory planning and resource allocation method that integrates synaesthesia, including:
[0097] Step 1: Figure 2 As shown, a drone cluster is randomly deployed in the detection area;
[0098] The detection area includes detection targets, drones;
[0099] Step 2: Figure 1 As shown, based on the Cartesian coordinate system, the fusion center obtains the current position information;
[0100] The location information includes the location coordinates of the fusion center itself, the drone cluster, and all detected targets;
[0101] Step 3: The fusion center plans the next trajectory of the drone cluster based on the current location information, constructs an optimization problem, and allocates flight direction, flight distance, power for sensing, and channel for communication for each drone;
[0102] Among them, drones At the moment by For length, is the angle movement, where , assuming that all drones fly at the same altitude, given drone The current position coordinates are ,in, For drones At the moment The horizontal coordinate of the position, For drones At the moment The vertical coordinate of the position, For drones At the moment The height of the position; its The position can be expressed by the following formula:
[0103] ;
[0104] Any two drones should be prevented from colliding to avoid additional losses. The minimum safe distance between the two drones is , then:
[0105] ;
[0106] in, For drones With drones Distance between:
[0107] ;
[0108] Where, For drones At the moment The location coordinates of drones At the moment The location coordinates of
[0109] The drone must also operate within a predefined detection area, with the boundaries of the detection area being , ,The relationship between the position of the UAV and the boundary of the detection area is as follows:
[0110] ;
[0111] In addition to flight missions, each UAV also needs to detect targets and transmit the detection information to the fusion center via a communication link. After the fusion center processes the detection data, it sends control instructions to the UAV. Each UAV uses a beam to perform dual functions: the main beam is used for target perception, and the secondary beam is used for communication. The radar's target detection model is as follows:
[0112] The total power of the drone is , , represents the power allocation coefficient, and the power allocated to the main beam is , the power allocated to the sub-beam is , the total bandwidth of the drone includes and channels, and the bandwidth of each channel is , assuming ,In order to avoid interference between perception and communication, each UAV is assigned a dedicated perception channel, so the set of channels allocated to communication is , binary variable Indicates whether the communication channel Assigned to drone ,in , , For the collection of drones, Specifically, if the communication channel Assigned to drones ,but ,otherwise .
[0113] Within the detection area, each drone emits electromagnetic waves into the environment and identifies potential targets by analyzing the reflected signals. The drone estimates the distance based on the round-trip time delay of the signal and uses statistical hypothesis testing to determine whether the reflected signal corresponds to the actual target. Since reflection, diffraction and scattering will cause multipath interference, selective interference mitigation technology is needed to improve detection performance. In this case, it is assumed that the distance estimation is mainly affected by environmental noise, and the environmental noise is simulated as an additive Gaussian white noise process. In order to quantify the accuracy of the distance estimation, the Cramer-Rao lower bound (CRLB) is used, which provides a theoretical lower bound for the variance of the unbiased estimator. According to modern radar theory, the CRLB of frequency and phase parameters is inversely proportional to the signal-to-noise ratio (SNR), indicating that improving the SNR can effectively reduce the variance of target detection.
[0114] drones At the moment The SNR is expressed as:
[0115] ;
[0116] Where, For drones The power allocated to the detection function, is the bandwidth of the drone, is the transmitting antenna gain, is the receiving antenna gain, is the working wavelength, Target radar cross section, is the Boltzmann constant, It's a drone and detection targets The distance between is the effective noise temperature, is the radar noise figure, To detect loss;
[0117] And the working wavelength satisfies:
[0118] ;
[0119] Where, is the speed of light, is the frequency of the radar.
[0120] In this embodiment, the effective noise temperature is .
[0121] Therefore, in order to ensure the quality of detection, the UAV At the moment Detection target The following conditions must be met:
[0122] ;
[0123] Where, The minimum signal-to-noise ratio is required to ensure effective detection of the UAV.
[0124] The probabilistic path loss model is used to describe the communication between UAVs and the fusion center, and between UAVs. The device on the previous hop The probability of line of sight (LoS) between for:
[0125] ;
[0126] Where, is the first constant, is the second constant, For drones The device on the previous hop The elevation angle between
[0127] Among them, drones The device on the previous hop The elevation angle between them satisfies:
[0128] ;
[0129] Where, For drones Previous hop device height, For drones The device on the previous hop the distance between them;
[0130] drones The device on the previous hop The probability of non-line-of-sight (NLoS) between for:
[0131] ;
[0132] Therefore, drones With the previous hop device The average channel power gain between is calculated as:
[0133] ;
[0134] Where, For drones With the previous hop device The average channel power gain between are the parameters of the traditional radar model, For drones The device on the previous hop The distance between and denote the attenuation factors for line-of-sight and non-line-of-sight links, respectively;
[0135] Based on transmission power and channel gain, UAV In the channel The signal-to-noise ratio (SINR) calculation formula is:
[0136] ;
[0137] Where, For drones In the channel The signal-to-noise ratio on is the communication noise factor, For drones At the moment The channel allocation coefficient, , For drones The power allocated to communication functions, For drones Power allocated to communication functions.
[0138] According to Shannon capacity formula, UAV The communication link data rate is:
[0139] ;
[0140] To ensure reliable data transmission, the minimum data rate of the drone must meet the following conditions:
[0141] ;
[0142] In order to ensure the detection accuracy, each target must be sensed multiple times. At the moment The effective detection is recorded as , and it satisfies:
[0143] ;
[0144] Where, Indicates the time Until now, the detection target has been successfully sensed and the number of drones transmitting information;
[0145] The conditions for effective detection of drones and time ,satisfy , , a metric to quantify the effectiveness of target detection is to time So far, the effective detection frequency of the target is expressed as:
[0146] ;
[0147] In order to deal with the potential geographical unfairness problem, that is, targets in dense areas may be detected multiple times, while sparse targets in edge areas may be ignored, this paper introduces the Jain fairness index, which measures the So far, the detection fairness among all targets is expressed as:
[0148] ;
[0149] Where, To the time Jain Fairness Index to date, For the The detected target at time The average detection frequency of
[0150] The goal of this multi-hop ISAC system is to maximize the effective detection frequency and geographical fairness of all targets while ensuring To ensure the communication and detection quality within a time slot, the system is limited by four conditions: drone safety distance, regional range, minimum signal-to-noise ratio, and minimum communication power. It is necessary to optimize the drone's flight direction, flight distance, power allocation, and channel allocation. The corresponding optimization problem is expressed as:
[0151] ;
[0152] Where, is the number of time slots, To the time Jain Fairness Index to date, To detect the target At the moment The number of effective detections, is the number of detected targets, For drones With drones The distance between is the minimum safe distance between two drones, is the horizontal coordinate boundary of the detection area, For drones At the moment The horizontal coordinate of the position, is the vertical coordinate boundary of the detection area, For drones At the moment The vertical coordinate of the position, For drones At the moment The signal-to-noise ratio, For drones Reach the minimum signal-to-noise ratio threshold for effective detection, For drones The communication link data rate, For drones Reaching the minimum communication link data rate threshold for effective detection, For drones At the moment The moving length, , For drones The maximum moving length, For drones At the moment The moving angle, For drones At the moment The power distribution coefficient, is the number of the communication channel, The set of channels allocated to communication by the fusion center, For drones At the moment The channel allocation coefficient of
[0153] The constraints in (19e) of the optimization problem ensure that each UAV is assigned exactly one communication channel at each moment, preventing interference and optimizing bandwidth utilization. The constraints in equations (19a), (19b), (19c), (19d), and (19e), respectively, define the UAV's flight range, direction, power allocation, and channel assignment to ensure optimal system performance.
[0154] In this embodiment, the optimization problem is solved by combining the SAC algorithm with the MDP;
[0155] Among them, in order to solve the proposed problem model, the system is constructed as a Markov decision process (MDP), which is a widely used framework in reinforcement learning. The radar sensing and communication environment described above is the learning environment. The fusion center acts as a decision maker and is responsible for trajectory planning, power allocation, and channel allocation. It transmits decisions to each UAV through the network. This structure allows the decision process to be formulated, enabling UAVs to make optimal decisions based on their observations and past actions. The MDP consists of the tuple Definition, where The details are as follows:
[0156] Action Space :The optimization problem includes determining the flight direction, flight distance, power allocation coefficient and channel allocation of each UAV. At the moment Actions It is defined as a vector of its flight parameters and resource allocation decisions, expressed as:
[0157] ;
[0158] The action space includes the actions of all drones in the system, so the time The system state is the set of all drone states .
[0159] State Space :Each drone makes a local observation of the environment and collects information about its current state and previous actions, which affect the current state. At the moment The observation space of is defined as follows:
[0160] ;
[0161] in, and Indicates drone At the moment The remaining variables reflect the actions taken in the previous time slot;
[0162] The state space includes the observation data of all drones in the system, so the time The system state is the set of all drone states .
[0163] Reward Function : The reward function aims to improve the efficiency of target detection while encouraging regional fairness. It consists of two main parts: detection reward and incentive reward;
[0164] The detection reward reflects the drone At the moment The reward is the contribution to the overall detection perception task. The reward is to aggregate the detection information of all UAVs through the fusion center to ensure collaborative performance. The calculation formula is:
[0165] ;
[0166] The incentive reward is to further improve the system performance and ensure fairness. This reward will discourage drones from repeatedly detecting the same target and encourage them to explore targets that are rarely detected or not detected. At the moment Internal perception of the goal , then the incentive reward of the drone is expressed as:
[0167] ;
[0168] Where, is a scaling factor used to adjust the reward magnitude; is a discount factor that reduces the reward for frequently detected targets, is the decay coefficient of the detection reward; Is a binary variable representing the drone Whether the detection target is effectively detected If the drone At the moment Target detected ,but ,otherwise .
[0169] To summarize, the total reward calculation formula is:
[0170] ;
[0171] This reward structure can both motivate UAVs to contribute to the overall system performance and ensure fair detection of targets throughout the entire sensing area.
[0172] In reinforcement learning, Soft Actor-Critic (SAC) is a policy gradient-based algorithm that adopts a maximum entropy-based reinforcement learning framework. It aims to simultaneously optimize the policy and value function to achieve efficient exploration and more stable learning. Compared with traditional policy gradient methods, SAC introduces an entropy term, combining the Actor-Critic method with the maximum entropy principle of continuous action space, maximizing the expected cumulative reward and policy entropy, thereby encouraging exploration and preventing convergence to suboptimal solutions. It is particularly suitable for solving the proposed formulated problem and encourages the intelligent agent to engage in more exploratory behavior, thereby improving the robustness of learning and the speed of convergence.
[0173] The SAC algorithm includes an Actor network, a first Critic network, a second Critic network, a first target Critic network, and a second target Critic network, specifically:
[0174] (1) Initialize Actor network parameters , First Critic network parameters , Second Critic Network Parameters , the first target Critic network parameters and the second target Critic network parameters , and copy the two critic network parameters to the two target network parameters respectively, and at the same time, initialize the experience playback buffer and the number of iterations .
[0175] (2) Initialize the number of iterations and time slots , and the number of time slots in each round of training is 100, which means the number of detections is 100 times.
[0176] (3) In the data collection phase, the agent interacts with the environment to collect training data. Specifically, the agent obtains the current state from the environment. , and use it as the input of the policy network (Actor), the Actor network outputs the current action based on this state , and then the drone cluster performs the action to update the environment and returns to the state of the next moment and the current moment's reward During the training process, the quadruple generated by each round of interaction Stored in the experience pool Experience Pool After collecting 1024 pieces of data, they serve as a buffer to facilitate random sampling during subsequent training, thereby promoting the model to learn more generalized strategies.
[0177] (4) When updating the two critic networks, Randomly draw a batch of experience samples from the current state of each sample and actions As the input of the two Critic networks, the expected returns of the corresponding actions are calculated, namely Value, denoted as 、 ; In order to train the Critic network, we need to use the target Critic network to obtain the target reward value. First, the next moment state Input into the Actor network to obtain the optimal action at the next moment , then, and Input them into the two target Critic networks respectively to get the target reward value, recorded as 、 , select 、 The minimum value in the next state value, based on which the target is calculated value, Time slot target The value is calculated as:
[0178] ;
[0179] Where, is the discount factor, is the entropy temperature parameter, For the Actor network, Output the logarithmic form of the probability density of the action for the Actor network;
[0180] Then use the gradient descent method to update the first critic network parameters and the second critic network parameters Specifically, we evaluate the network predictions by comparing Values and Goals value , the mean square error can be calculated, and then the back propagation algorithm can be used to update the parameters of the two critic networks and , to minimize the loss value, that is:
[0181] ;
[0182] Where, For the The loss value of the Critic network, To represent the experience replay buffer All experience samples obtained from Calculate the expected value of the loss function;
[0183] (5) Update Critic network parameters :
[0184] ;
[0185] Where, is the updated Critic network parameter, is the learning rate, is the loss function For the first Critic network parameters gradient;
[0186] (6) Calculate the loss function of the Actor network and the effect of the loss function on the Actor network parameters gradient;
[0187] Among them, the loss function of the Actor network is:
[0188] ;
[0189] Where, Indicates status and actions The expectation of the joint distribution of
[0190] The loss function is used to determine the Actor network parameters. The gradient of is:
[0191] ;
[0192] Where, is the gradient of the policy distribution, It's the action The gradient, It's action Parameters gradient.
[0193] (7) Then update the parameters of the Actor network according to the gradient descent method:
[0194] ;
[0195] Where, is the updated Actor network parameter, is the learning rate, is the gradient of the policy function to the policy network parameters;
[0196] (8) In order to automatically adjust the entropy of the strategy so that the strategy reaches a dynamic balance between exploration and utilization, the entropy temperature parameter needs to be updated:
[0197] ;
[0198] Where, is the updated entropy temperature parameter, is the target entropy;
[0199] (9) Critic network for the first target and the second target Critic network Adopt a soft update strategy, namely:
[0200] ;
[0201] Where, After the update Target Critic network parameters, is the scaling factor for soft updates;
[0202] (10) Time slot Increase by 1 to determine whether the number of training time slots per round reaches 100; if not, return to step 3 to continue training; if reached, determine whether all algorithm iterations have been completed ; If not completed, increase the number of iterations by 1 and reinitialize the time slot , return to step 3 to continue training; if all iterations are completed, the algorithm ends.
[0203] By continuously iterating the above steps, the SAC algorithm can continuously optimize the policy network and value network during the training process, and ultimately obtain an effective decision-making strategy.
[0204] In this embodiment, the soft update scaling factor is 0.005, and the number of iterations is 15,000.
[0205] As mentioned above, the specific process of the SAC algorithm is shown in Table 1:
[0206] Table 1 SAC algorithm flow
[0207]
[0208] Step 4: Send the fusion center's instructions to each drone node according to the number of hops specified in the multi-hop network;
[0209] Step 5: Each hop UAV receives and executes the instructions assigned by the fusion center, i.e., moves according to the planned trajectory and performs target detection according to the allocated power;
[0210] Step 6: Each hop UAV forwards the current sensing information to the previous hop node along the assigned channel through the multi-hop network until the fusion center receives all the detection information;
[0211] Step 7: Repeat the above detection process until the specified number of detections is reached, ending this round of communication perception.
[0212] In this embodiment, there are 1 fusion center, 3 drones and 35 targets. The number of available channels is 7. The movement range of all drones is limited to the area Meters, flight altitude The maximum flight distance of each drone is 100 meters, the minimum safe distance between drones is 50 meters, and the height of the fusion center is For radar and communication functions, the maximum transmission power of each drone is 5 meters. 0.1W, the bandwidth of each channel The operating frequency is 1MHz The frequency is 3GHz; the sensor-related parameter settings are as follows: , , , , The communication-related parameter settings are as follows: , , , , The thresholds for sensing and communication are set to and .
[0213] The present invention uses an advanced reinforcement learning algorithm, namely the SAC algorithm, which is more efficient than traditional convex optimization and swarm intelligence algorithms and has better training results. Figure 3 and Figure 4 As shown in the figure, the superiority of the SAC algorithm is reflected in the system's reward value and the number of effective detections. During the training process, the average reward and the average number of effective detections of the SAC algorithm are higher than those of other reinforcement learning algorithms, ensuring convergence while maintaining strategy diversity.
[0214] The present invention designs and develops a multi-hop UAV trajectory planning and resource allocation method with integrated synesthesia. Different from the traditional ISAC and UAV combination system, it increases the UAV communication distance, expands the system's detection range, improves the system's spectrum utilization, and better leverages the advantages of synesthesia integration. It has strong portability, requires only one fusion center and multiple UAVs, is easy to deploy, and saves system overhead, making it more suitable for large-scale, energy-constrained collaborative detection environments.
[0215] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.
Claims
1. A multi-hop UAV trajectory planning and resource allocation method based on synaesthesia, characterized in that: The steps include: Step 1: Randomly deploy drone clusters within the detection area; Step 2: Based on the Cartesian coordinate system, the fusion center obtains the current position information; The current location information includes the location coordinates of the fusion center, the UAV cluster, and all detected targets; Step 3: The fusion center plans the next trajectory of the UAV cluster based on the current location information, constructs an optimization problem, and converts the optimization problem into a Markov decision process. The optimization problem is solved using the SAC algorithm to obtain the flight direction, flight distance, power, and channel of each UAV. The optimization problem is: ; Where, is the number of time slots, To the time Jain Fairness Index to date, To detect the target At the moment The number of effective detections, is the number of detected targets, For drones With drones The distance between is the minimum safe distance between two drones, is the horizontal coordinate boundary of the detection area, For drones At the moment The horizontal coordinate of the position, is the vertical coordinate boundary of the detection area, For drones At the moment The vertical coordinate of the position, For drones At the moment The signal-to-noise ratio, For drones Reach the minimum signal-to-noise ratio threshold for effective detection, For drones The communication link data rate, For drones Reaching the minimum communication link data rate threshold for effective detection, For drones At the moment The moving length, , For drones The maximum moving length, For drones At the moment The moving angle, For drones At the moment The power distribution coefficient, is the number of the communication channel, The set of channels allocated to communication by the fusion center, drones At the moment The channel allocation coefficient of The drone The communication link data rate satisfies: ; Where, For drones In the channel The signal-to-noise ratio on The drone In the channel The signal-to-noise ratio on satisfies: ; Where, is the communication noise factor, For drones At the moment The channel allocation coefficient, For drones With the previous hop device The average channel power gain between , For drones With the previous hop device The average channel power gain between For drones The power allocated to the communication function, For drones The power allocated to the communication function, The bandwidth for the drone; Step 4: Send the fusion center's instructions to each drone node according to the number of hops specified in the multi-hop network; Step 5: Each hop UAV receives and integrates the instructions assigned by the center, moves according to the planned trajectory, and detects the target according to the allocated power; Step 6: Each hop UAV forwards the current sensing information to the previous hop node along the assigned channel through the multi-hop network until the fusion center receives all the detection information; Step 7: Repeat the above detection process until the specified number of detections is reached, ending this round of communication perception.
2. The synesthesia-integrated multi-hop UAV trajectory planning and resource allocation method according to claim 1, characterized in that: The detection target At the moment The effective number of detections meets the following requirements: ; Where, Indicates the time Until now, the detection target has been successfully sensed and the number of drones transmitting information.
3. The synesthesia-integrated multi-hop UAV trajectory planning and resource allocation method according to claim 2, characterized in that: The time The Jain fairness index so far satisfies: ; Where, For the The detected target at time The average detection frequency.
4. The synesthesia-integrated multi-hop UAV trajectory planning and resource allocation method according to claim 3, characterized in that: The drone At the moment The position meets: ; Where, For drones At the moment The horizontal coordinate of the position, For drones At the moment The vertical coordinate of the position.
5. The synesthesia-integrated multi-hop UAV trajectory planning and resource allocation method according to claim 4, characterized in that: The drone At the moment The signal-to-noise ratio satisfies: ; Where, For drones The power allocated to the detection function, is the bandwidth of the drone, is the transmitting antenna gain, is the receiving antenna gain, is the working wavelength, Target radar cross section, is the Boltzmann constant, It's a drone and detection targets The distance between is the effective noise temperature, is the radar noise figure, To detect loss.
6. The synesthesia-integrated multi-hop UAV trajectory planning and resource allocation method according to claim 5, characterized in that: The drone With the previous hop device The average channel power gain between satisfies: ; Where, are the parameters of the traditional radar model, For the moment drones The device on the previous hop The distance between is the attenuation factor of the line-of-sight link, is the attenuation factor of the non-line-of-sight link, For the moment drones The device on the previous hop The probability of sight distance between For the moment drones The device on the previous hop The probability of non-line-of-sight between At the moment drones The device on the previous hop The probability of the line of sight between them satisfies: ; Where, is the first constant, is the second constant, For drones The device on the previous hop The elevation angle between At the moment drones The device on the previous hop The non-line-of-sight probability between them satisfies: ; The drone The device on the previous hop The elevation angle between them satisfies: ; Where, For drones The device on the previous hop The distance between For drones Previous hop device height, For drones At the moment The height of the location.
7. The synesthesia-integrated multi-hop UAV trajectory planning and resource allocation method according to claim 6, characterized in that: Converting the optimization problem into a Markov decision process specifically includes: Action Space: , ; State Space: , ; Reward function: ; in, For drones At the moment The action of time, For drones At the moment The observation space, For detection rewards, To motivate rewards, is a collection of detection targets; The detection reward satisfies: ; The incentive reward meets the following requirements: ; Where, is the scaling factor, is the decay coefficient of the detection reward, For drones Whether the detection target is effectively detected binary variable.
8. The synesthesia-integrated multi-hop UAV trajectory planning and resource allocation method according to claim 7, characterized in that: The SAC algorithm specifically includes the following steps: Step 1. Initialize Actor network parameters , First Critic network parameters , Second Critic Network Parameters , the first target Critic network parameters and the second target Critic network parameters , and copy the two critic network parameters to the two target critic network parameters, and initialize the experience replay buffer and the number of iterations; Step 2: Initialize the number of iterations and time slots, and the number of training time slots in each iteration is 100; Step 3: Get from the environment Status of time slot Enter the Actor network and get Time slot action ; Step 4: Drone swarm performs actions , get the state of the next time slot and the reward for the current time slot ; Step 5: Stored in the experience replay buffer; Step 6: If the number of experiences stored in the experience replay buffer is greater than the preset minimum sampling batch size, A set of experience samples are randomly selected from the set, and for each sample, Time slot status and Time slot action Input two Critic networks respectively, obtain the Q value, and the state of the next moment Input Actor network to generate the action at the next moment , and then 、 The target return value is calculated through the first target critic network and the second target critic network respectively, and the target Q value is obtained accordingly: ; Where, for Time slot target value, is the target return value output by the first target Critic network, is the target return value output by the second target Critic network, is the discount factor, is the entropy temperature parameter, For the Actor network, Output the logarithmic form of the probability density of the action for the Actor network; Step 7. Update the parameters of the two critic networks by gradient descent to minimize the loss values of the two critic networks: ; Where, For the The loss value of the Critic network, To represent the experience replay buffer All experience samples obtained from Calculate the expected value of the loss function; Step 8. Update Critic network parameters : ; Where, is the updated Critic network parameter, is the learning rate, is the loss function For the first critic network parameters gradient; Step 9. Calculate the loss function of the Actor network and the effect of the loss function on the Actor network parameters gradient; Among them, the loss function of the Actor network is: ; Where, Indicates status and actions The expectation of the joint distribution of The loss function is used to determine the Actor network parameters. The gradient of is: ; Where, is the gradient of the policy distribution, It's the action The gradient, It's action Parameters gradient; Step 10. Update the Actor network parameters, entropy temperature parameters, and two target Critic network parameters: ; ; ; Where, is the updated Actor network parameter, is the learning rate, is the gradient of the policy function to the policy network parameters, is the updated entropy temperature parameter, is the target entropy, After the update Target Critic network parameters, is the scaling factor for soft updates; Step 11: If , then jump directly to step 3; like and , then jump directly to step 2; like and , the algorithm ends.
Citation Information
Patent Citations
Unmanned aerial vehicle intelligent trajectory planning and communication resource allocation method based on reinforcement learning
CN116704823A
Radar and communication integrated unmanned aerial vehicle cooperative multi-target detection method
CN114679729A