Energy Consumption Optimization Method for Integrated Sensing and Communication of UAV Swarms Based on Reinforcement Learning
Through the method based on reinforcement learning, the action selection of the drone cluster is optimized, and the problem of high energy consumption of the drone cluster ISAC network is solved, efficient communication and perception in complex environments are achieved, and the service life of the drone cluster is extended.
Patent Information
- Application Number
- CN202310843486.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-07-11
AI Technical Summary
The ISAC network of the existing drone group consumes a lot of energy and cannot effectively meet the communication and perception needs in complex environments, resulting in insufficient network performance.
Using reinforcement learning-based method, the Q-value function network is initialized using the positioning and perception performance of the drone cluster, and combined with the energy consumption of communication tasks, the reward function is designed, and the action selection of the drone is optimized through the ε-greedy strategy to achieve energy consumption optimization.
On the premise of ensuring communication and perceptual performance, reduce the energy consumption of drone networks, extend the service life of drone clusters, and improve the network service capabilities of resource-constrained drone systems.
Smart Images

Figure CN116896777B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of integrated communication and sensing of unmanned aerial vehicles, and particularly relates to a method for optimizing the energy consumption of integrated communication and sensing of an unmanned aerial vehicle swarm based on reinforcement learning. Background Art
[0002] The next-generation wireless network (B5G / 6G) promotes the continuous innovation and development of wireless technologies, and at the same time provides key support for many emerging applications, such as connected intelligence, connected vehicles, and smart cities. These applications require high-quality wireless communication connections and high-precision sensing capabilities. Therefore, it can be foreseen that in the B5G / 6G network, the capabilities of both communication and sensing are required to improve the utilization rate of spectrum resources. Among them, the integrated communication and sensing (ISAC) technology is widely regarded as one of the effective solutions to achieve this goal. To meet the actual demand of users for the full-domain collaborative coverage of communication and sensing services, future wireless networks need to be more deployed in complex terrains and electromagnetic environments such as dense cities and mountainous areas. However, in the above complex environments, existing network technologies represented by cellular networks and global navigation satellite systems (GNSS) all have certain defects, making their network performance unable to meet the high-quality communication and sensing requirements of users.
[0003] Especially in harsh environments, such as areas with high-rise buildings or remote mountainous areas, traditional satellite positioning services may lead to poor service quality of ground user communication and sensing due to problems such as network marginalization. In this case, the unmanned aerial vehicle swarm, with its advantages of high mobility and flexible deployment, is expected to make up for this deficiency by combining corresponding communication and sensing technologies. Therefore, in recent years, the unmanned aerial vehicle cluster-assisted technology has received extensive attention from the academic and industrial circles. However, the current power supply method for unmanned aerial vehicles is usually based on battery charging. Even if a few models of unmanned aerial vehicles can supplement energy through solar energy or other solutions, their energy supply is relatively limited. Therefore, unmanned aerial vehicles themselves have the deficiency of limited resources. In this case, how to reasonably allocate resources such as power and served users to reduce the energy consumption of the airborne platform becomes increasingly important. Therefore, how to minimize the energy consumption of unmanned aerial vehicles by optimizing the power allocation and service strategy of unmanned aerial vehicles on the basis of completing communication and sensing tasks is the core research problem of the present invention. The research on this problem is expected to effectively reduce the energy consumption of the airborne cluster system on the premise of ensuring the communication and sensing performance of the system, thereby improving the service life of the integrated communication and sensing network of the unmanned aerial vehicle swarm. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for optimizing the energy consumption of integrated communication and sensing of an unmanned aerial vehicle swarm based on reinforcement learning, so as to solve the problem of large energy consumption in the ISAC network based on the unmanned aerial vehicle swarm in the prior art, and at the same time ensure the real-time service performance of the communication and positioning of terminal ground users, and effectively improve the network service life.
[0005] An energy consumption optimization method for integrated communication and sensing of an unmanned aerial vehicle (UAV) swarm based on reinforcement learning. The established system environment is as follows: In a cellular communication and sensing network covered by a single base station, the coordinates of the base station are (x0, y0, z0), and the set of UAVs is The set of ground users is A complete decision-making process of the UAV as an intelligent agent is called a decision-making cycle. In the present invention, a complete process of the UAV to complete positioning, sensing, and communication task offloading is a decision-making cycle. Assuming that each moment t is a decision-making cycle, the set of states of all decision-making cycles is called the state space set S, which is expressed as: S = {s1, s2,......, s t ,......}. The set of actions of all decision-making cycles is called the action space set A, which is expressed as: A = {a1, a2,......, a t ,......}. At moment t, the position coordinates of UAV-m and the user-l to be located are u m (t) = (x m (t), y m (t)) T and v l = (x l , y l ) T , and it is assumed that the height of UAV platform-m is a fixed value H m . All UAVs in the cluster fly on a pre-set flight trajectory that meets a certain threshold range, and at the same time provide communication and position sensing services for ground users, and each UAV can be recognized by the ground user terminals to be served.
[0006] First step, utilize the positioning and sensing performance of the UAV cluster to assign an initial value to the Q-value function network reflecting the relationship between UAV states and actions.
[0007] Second step, judge the current state of the UAV from the environment. A certain working point can be selected as the initial state of the UAV.
[0008] Third step, based on the current state of the UAV, select the current action according to the ε-greedy strategy.
[0009] Fourth step, introduce the UAV sensing performance and the energy consumption of communication tasks into the design of the Reward reward function, and obtain the actual environmental reward value of the previous step action, as well as the next state of the UAV.
[0010] Fifth step, use the actual environmental reward value of the previous step action to update the Q-value function network.
[0011] In the sixth step, set the new state as the current state, and repeat the third step to the sixth step until the value in the Q-value function network converges.
[0012] In the above steps, the following key technical points are mainly involved:
[0013] (1) Initialize the Q-value function network using the positioning and sensing performance of the UAV cluster
[0014] The geometric configuration between the UAV and the ground users to be served is the premise affecting its positioning and sensing performance, and the position (three-dimensional) dilution of precision PDOP can be used to characterize this performance. When the UAV is in different states, different actions are selected, so the geometric configuration presented between the UAV and the users is different, and its ability to provide sensing services for the users to be served is also different. Assume that at time t, the subset of UAV base stations providing positioning and sensing services for user - l is S k (t), and assume that the number of UAVs in the set is M0. Then calculate the position (three-dimensional) dilution of precision value, which can be expressed as:
[0015]
[0016] In the above formula, is the Jacobian matrix of the positioning and sensing observation equation of the UAV base station subset S k (t), and it can be further expressed as the following formula:
[0017]
[0018] In the above formula, u1(t) = (x1(t), y1(t)) T , (1 ∈ s k (t)), respectively represent the coordinates of UAV - 1, UAV - m0, and UAV - M0 in the UAV base station subset S k (t); similarly, H1, respectively are the fixed altitude values of UAV - 1, UAV - m0, and UAV - M0 in the UAV base station subset S k (t), and v l = (x l , y l ) T is the position coordinate of the user - l to be located.
[0019] The method for initializing the Q-value function network using the positioning and sensing performance of the UAV cluster is as follows: when the UAV is in the normal working state, at time t, when the UAV base station subset S k (t) provides communication and sensing services for users, the value in its corresponding Q-table grid is The remaining grid positions are assigned zero. Utilize the positioning and sensing performance of the UAV swarm to assign initial values to the Q-value function network, thereby providing prior information for the reinforcement learning network of the UAV agent and further facilitating the learning of the agent.
[0020] (2) Energy consumption for uploading communication tasks
[0021] At time t, the LoS channel power gain from the m-th UAV to the l-th ground user can be expressed as:
[0022]
[0023] where α is the path loss coefficient of the channel, β0 is the channel gain per unit (per meter), and d m,l (t) is the distance from the m-th UAV to the l-th ground user, and there is Furthermore, the signal-to-noise ratio SNR m,l (t) at this moment for this link can be expressed as:
[0024]
[0025] where P is the constant transmission power of the space-based platform, and σ 2 is the noise power, indicates whether the space-based platform provides communication and sensing services for the ground user. Specifically, the meaning is: when , it does not provide communication and sensing services, and when , it provides communication and sensing services. is the interference from other UAV platforms, that is, co-channel interference. Among them, P u (t) is the transmission power of other UAV platforms at time t; g u,l (t) is the LoS channel power gain of other UAV platforms at time t; represents the set of UAVs. Then the data transmission rate R m,l (t) from the m-th UAV to the l-th ground user for this link can be expressed as:
[0026] R m,l (t) = B · log2(1 + SNR m,l (t)) (5)
[0027] where B is the signal bandwidth. Furthermore, the energy consumption E m,l (t) for this link can be expressed as:
[0028]
[0029] where P is the constant transmission power of the space-based platform, Indicates whether the empty base platform provides communication and sensing services for ground users. Specifically, it means that when it does not provide communication and sensing services, and when it provides communication and sensing services. P m (t) ∈ [0, 1, …, 5] represents the number of users served by a single UAV. We assume that the data upload task of each ground user can be sent to at most one UAV.
[0030] (3) Action selection strategy
[0031] As a value-based reinforcement learning algorithm, the Q-learning algorithm uses the iterative update of the Q-value function to find the optimal policy π of the UAV (agent). * During the operation of the algorithm, the agent selects actions to execute according to the ε-greedy strategy. That is, the probability that the agent randomly selects an action is ε, and the probability of selecting the action corresponding to the maximum value in the Q-value function network is 1 - ε.
[0032] When the algorithm starts to execute, first initialize the Q-table using the method in (1) above, and then select the current state s t , for each action a in this state, there is a corresponding "state-action" value, denoted as Q(s t , a). In this case, select the action in this state according to the ε-greedy strategy, that is, select the action corresponding to the maximum value in the Q-value function network, as shown in the following formula:
[0033]
[0034] After selecting the action, the agent starts to execute this action, and then enters the next state s t+1 and obtains the reward value r(t) within the current action selection decision period. At the same time, update the value at the corresponding position in the corresponding Q-network:
[0035]
[0036] where γ is the discount factor and γ ∈ [0, 1]. Regarding the reward value function r(t), the specific description is given in (4).
[0037] (4) Design of the reward value function r(t)
[0038] To comprehensively ensure the communication and sensing performance of the UAV swarm network, the reward function r(t) at time t is designed as the following formula:
[0039]
[0040] where SNR thris a pre-set known system parameter signal-to-noise ratio threshold, aiming to ensure the communication performance of the agent; PDOP thr is the threshold of the known parameter three-dimensional precision factor. The purpose is to ensure the perception performance of the agent. Based on this function, the communication and sensing characteristics of the UAV swarm network in the process of energy consumption optimization action selection are ensured.
[0041] The method for optimizing the energy consumption of the UAV swarm with integrated communication and sensing based on reinforcement learning in the present invention, compared with the existing methods, its advantages and beneficial effects are as follows:
[0042] (a) The present invention comprehensively considers the communication and sensing performance of the UAV cluster and the network energy consumption problem, and proposes an energy efficiency optimal strategy based on the UAV cluster integrated communication and sensing network, which can effectively reduce the UAV network energy consumption on the premise of ensuring the communication and sensing performance of the UAV system, thereby improving the network service life of the UAV cluster with limited resources itself.
[0043] (b) The present invention designs an intelligent decision-making algorithm based on reinforcement learning, enabling the UAV to adaptively select the number of served users and power according to the dynamically changing environment. On the premise of ensuring the communication and sensing performance of the system, it offloads the most task upload traffic with the minimum network energy consumption, avoiding the rigid mode of the traditional centralized network control and overcoming the problems brought by the environmental dynamics to the strategy formulation. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical principle and specific process scheme of the proposed invention, the following will briefly describe and introduce the relevant drawings involved in the embodiments. Obviously, the attached Figures 1 to 3 drawings described below are only for the description and illustration of the embodiments, and for those ordinary related technical researchers in the field, such other drawings can be obtained without creative labor.
[0045] Figure 1 is a schematic flow chart of the method for optimizing the energy consumption of the UAV swarm with integrated communication and sensing based on reinforcement learning proposed by the present invention.
[0046] Figure 2 is a schematic learning flow chart of the Q-value function network proposed by the present invention.
[0047] Figure 3 is a schematic diagram of the integrated communication and sensing network scenario based on the UAV swarm targeted by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The following will further describe the features and principles of the present invention in combination with the drawings and embodiments. The role of the listed embodiments is only to explain the present invention, and is not used to limit the application scope of the present invention.
[0049] Reference Figure 1 As shown, considering a cellular network covered by a single base station, in the scenario where the network radius is 500 m, the present invention proposes a method for optimizing the energy consumption of integrated communication and sensing for a swarm of drones based on reinforcement learning. The following will further explain and introduce in detail the specific implementation manner of the overall inventive method by taking the parameters indicating whether the corresponding selected user is served; P m (t) ∈ [0, 1, …, 5] represents the number of users served by a drone; it is assumed that the exploration rate ε of the reinforcement learning network is set to 0.8; the discount factor γ is 0.9; l max = 1000; the communication threshold SNR thr = 2 dB; the positioning and sensing threshold PDOP thr = 1.5; the number of drones for providing positioning and sensing services to each ground user is 4; and taking the example that the data upload task of each ground user can be sent to at most 1 drone to further explain and introduce in detail the specific implementation manner of the overall inventive method.
[0050] Step 1: During the operation of the system, first establish a Q-table grid. Then, initialize the Q-value function network by using the positioning and sensing performance of the drone cluster according to the proposed scheme, that is, the three-dimensional dilution of precision. Specifically, when the drone is in a normal working state, at time t, when selecting a subset S k (t) of the drone base station to provide communication and sensing services to users, the value in the corresponding Q-table grid is -PDOP sk(t) , and the values of the remaining grid positions are assigned zero.
[0051] Step 2: Select a certain state as the initial state of the drone agent according to the current network environment.
[0052] Step 3: Based on the ε-greedy policy, select the action of the drone service state and the number of served users selected in the current state s t selected in Step 2, that is, determine and P m (t) according to formula (7) and ε = 0.8. Specifically, with a probability of 1 - ε = 0.2, select the action corresponding to the maximum value in the Q-value function network, that is, randomly select an action with a probability of ε = 0.8;
[0053] Step 4: After the action decision is completed, the drone obtains the energy consumption of the communication upload task during this action decision period, that is, obtain E m,l (t) through formula (6), and and the values of P m (t), as well as the communication threshold SNR thr = 2 dB and the positioning and sensing threshold PDOP thrSubstitute ε = 1.5 and γ = 1.5 into formula (9) to calculate the reward value r(t), and simultaneously transfer to the next state s t+1 ;
[0054] Step 5: Substitute the exploration rate ε = 0.8 and the discount factor γ = 0.9 of the reinforcement learning network into formula (8) to obtain Q(s t , a t ), thereby updating the value of the Q-value function network;
[0055] Step 6: Set the new state s t+1 as the current state, and repeat Steps 3 - 6 until the value in the Q-value function network converges.
[0056] When the value in the Q-value function network finally reaches the convergence state through continuous updates, the Q-table can be used to guide the UAV to make the best decision in the corresponding state, that is, to select the optimal number of served users and the transmission power in the corresponding state, and obtain the optimal user communication task traffic offloading strategy, that is, the optimal energy efficiency of the UAV. The following gives the full process of the algorithm:
[0057] Q-Learning algorithm: Obtain the optimal energy efficiency strategy of the integrated communication and sensing network of UAV clusters
[0058] Initialize for any s ∈ S, a ∈ A(s),
[0059] Use the key technology (1) to assign an initial value to the Q-table
[0060] Initialize t = 1, ε = 0.8, γ = 0.9,
[0061] Repeat:
[0062] Initialize the state s according to the current environmental information
[0063] Repeat in each action decision period t:
[0064] Select the action a under the state s according to the ε-greedy strategy
[0065] Execute the action a, obtain the reward function value using formula (9) and enter the next state s'
[0066] Update the corresponding value of the Q-table using formula (8)
[0067] Let t = t + 1, s' = s
[0068] Repeat the above steps until the maximum number of iterations l max = 1000.
[0069] The overall detailed flowchart of the features and principles of the present invention involved in the above-listed embodiments is as follows Figure 2 shown, and the relevant scenario diagram targeted by the present invention is as follows Figure 3 shown.
[0070] In summary, a method for optimizing the energy consumption of an integrated communication and sensing drone swarm based on reinforcement learning proposed by the present invention provides prior information for the reinforcement learning network by utilizing the sensing performance of the drones, enabling the drones to reach the optimal target state more quickly and efficiently. On the other hand, on the premise of ensuring the communication and sensing performance of the drone swarm, the system task energy consumption is reduced, thereby effectively improving the service life of the drone network, which is of great significance for the drone system with limited resources itself.
Claims
1. An energy consumption optimization method for integrated communication and sensing of UAV swarms based on reinforcement learning, characterized in that: It includes the following steps: In the first step, utilize the positioning and sensing performance of the UAV swarm to initialize the Q-value function network that reflects the relationship between the UAV state and actions; In the second step, judge the current state of the UAV from the environment; select a certain working point as the initial state of the UAV; In the third step, based on the current state of the UAV, select the current action according to the ε-greedy strategy; In the fourth step, introduce the UAV sensing performance and the communication task energy consumption into the design of the Reward function, and obtain the actual environmental reward value of the previous action, as well as the next state of the UAV; In the fifth step, use the actual environmental reward value of the previous action to update the Q-value function network; In the sixth step, set the new state as the current state, and repeat the third step to the sixth step until the values in the Q-value function network converge; Among them, the Q-value function network is initialized as follows: At time t, the subset of UAV base stations providing positioning perception services for user -l is S k (t), and let the number of UAVs in the set be M0; then calculate the dilution of precision of this UAV subset value, expressed as: In the above formula, is the Jacobian matrix of the positioning and sensing observation equation of the UAV base station subset S k (t); Among them, selecting the current action based on the ε-greedy strategy is: the probability that the agent randomly selects an action is ε, and the probability of selecting the action corresponding to the maximum value in the Q-value function network is 1 - ε; When starting to execute, first initialize the Q-table using the method in formula (1), and then select the current state s t , for each action a in this state, there is a corresponding "state-action" value, which is denoted as Q(s t , a); in this case, select the action in this state according to the ε-greedy strategy, that is, select the action corresponding to the maximum value in the Q-value function network, as shown in the following formula:
2. The method for optimizing the energy consumption of an integrated communication and sensing drone swarm based on reinforcement learning according to claim 1, wherein: It is further expressed as the following formula: where, u1(t) = (x1(t), y1(t)) T , (1 ∈ s k (t)), respectively represent the coordinates of UAV-1, UAV-m0, and UAV-M0 in the UAV base station subset S k (t); similarly, H1, respectively are the fixed altitude values of UAV-1, UAV-m0, and UAV-M0 in the UAV base station subset S k (t), and v l = (x l , y l ) T is the position coordinate of the user-l to be located.
3. The energy consumption optimization method for integrated communication and sensing of an unmanned aerial vehicle swarm based on reinforcement learning according to claim 1, characterized in that: When the drone is in the normal working state, select a subset S of the drone base stations at time t k (t) When providing communication and sensing services for users, the value in the corresponding Q-table grid is The values of the remaining grid positions are assigned zero; utilize the positioning and sensing performance of the drone swarm to initialize the Q-value function network.
4. The energy consumption optimization method for integrated communication and sensing of an unmanned aerial vehicle swarm based on reinforcement learning according to claim 1, wherein: The communication task energy consumption is that at time t, the LoS channel power gain from the m-th UAV to the l-th ground user is expressed as: where α is the path loss coefficient of the channel, β0 is the channel gain per unit (per meter), and d m,l (t) is the distance from the m-th UAV to the l-th ground user, and there is 5. The energy consumption optimization method for integrated communication and sensing of an unmanned aerial vehicle swarm based on reinforcement learning according to claim 4, wherein: The signal-to-noise ratio SNR of the link at this moment m,l (t) is expressed as: where P is the constant transmission power of the empty base platform, and σ 2 is the noise power, indicates whether the empty base platform provides remote sensing services for ground users. Specifically, when it does not provide remote sensing services, and when it provides remote sensing services; is the interference from other UAV platforms, that is, co-channel interference. Among them, P u (t) is the transmission power of other UAV platforms at time t; g u,l (t) is the LoS channel power gain of other UAV platforms at time t; represents the set of UAVs; then the data transmission rate R m,l (t) of the link from the m-th UAV to the l-th ground user is expressed as: R m,l R(t) = B·log2(1 + SNR m,l (t)) (5) Among them, B is the signal bandwidth.
6. The integrated communication and sensing energy consumption optimization method for UAV swarms based on reinforcement learning according to claim 5, characterized in that: The energy consumption E of the link m,l (t) is expressed as: Among them, P is the constant transmission power of the space-based platform, indicating whether the space-based platform provides remote sensing services for ground users. Specifically, it means that when it does not provide remote sensing services, and when it provides remote sensing services; P m (t) ∈ [0, 1, …, 5] represents the number of users served by a single UAV; the data upload task of each ground user can be sent to at most one UAV.
7. A method for optimizing the energy consumption of an integrated communication and sensing UAV swarm based on reinforcement learning according to claim 1, characterized in that: After the selected action, the agent starts to execute this action and then enters the next state s t+1 and obtains the reward value r(t) within the current action selection decision cycle. At the same time, the value at the corresponding position in the corresponding Q-network is updated: Among them, γ is the discount factor, and γ ∈ [0, 1].
8. A method for optimizing the energy consumption of an integrated communication and sensing drone swarm based on reinforcement learning according to claim 7, characterized in that: Design of the reward value r(t): Design that at time t, the reward function r(t) is expressed as the following formula: Among them, SNR thr is a pre-set known system parameter signal-to-noise ratio threshold, aiming to ensure the communication performance of the agent; PDOP thr is the threshold of the three-dimensional dilution of precision for known parameters; the purpose is to ensure the perception performance of the agent; based on this function, the communication and sensing characteristics of the UAV swarm network in the process of energy consumption optimization action selection are ensured.