Reflector enhanced drone assisted deployment data optimization method and related devices
By constructing objective functions and constraints to optimize the deployment data of UAVs and intelligent reflectors, the problems of excessive communication quality and energy consumption in the research of combining UAVs and RIS were solved. A balance between secure communication rate and energy consumption in cell-free networks was achieved, improving the rationality and efficiency of network deployment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2025-04-21
- Publication Date
- 2026-08-04
AI Technical Summary
In the research on the combination of UAV and RIS, the improper deployment of UAVs and smart reflective surfaces in the existing technology leads to poor coverage and communication quality of cell-free networks, as well as excessive communication energy consumption, making it difficult to achieve a balance between total secure communication rate and total communication energy consumption.
By constructing an objective function to maximize the ratio of total secure communication rate to total communication energy consumption, and combining constraints, the deployment data of UAVs and intelligent reflectors are optimized. The optimized deployment data is obtained by solving the problem using a deep deterministic policy gradient algorithm and a convex optimization algorithm.
It achieves a dynamic balance between total secure communication rate and total communication energy consumption in decellularized networks, provides a reasonable data foundation for network deployment, and improves communication security and energy efficiency.
Smart Images

Figure CN120499675B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method and related equipment for optimizing deployment data of unmanned aerial vehicles (UAVs) with enhanced reflectivity. Background Technology
[0002] In recent years, thanks to their flexibility and deployment capabilities, unmanned aerial vehicles (UAVs) have been widely used in remote monitoring, photography, traffic control, and cargo transportation. In particular, in cell-free networks, using UAVs as aerial platforms, leveraging their high mobility and ease of deployment, can effectively improve the coverage and communication quality of cell-free communication systems. Meanwhile, reconfigurable intelligent surfaces (RIS) have shown great potential in improving the wireless propagation environment. Deploying RIS in cell-free networks can effectively reduce the base station's transmission power while achieving higher service quality.
[0003] In current research combining UAVs and RIS, the proper deployment of UAVs and RIS is crucial, affecting network communication quality and energy consumption. Inappropriate network deployment can lead to poor coverage and communication quality in cell-free networks, resulting in low communication security. Furthermore, efforts to improve communication quality may consume excessive energy, leading to excessive communication power consumption. Summary of the Invention
[0004] In view of this, the purpose of this application is to propose a method and related equipment for optimizing deployment data of UAVs with enhanced reflective surfaces, so as to overcome all or part of the shortcomings of the prior art.
[0005] To achieve the above objectives, this application provides a method for optimizing deployment data of UAVs with enhanced reflectivity, comprising: acquiring deployment data corresponding to UAVs within a predetermined area; constructing an objective function for optimizing the deployment data based on the deployment data, with the goal of maximizing the ratio of total secure communication rate to total communication energy consumption, and constructing constraints corresponding to the objective function; and solving the objective function based on the constraints to obtain the maximized ratio and optimized deployment data.
[0006] Optionally, the objective function includes a first sub-objective function and a second sub-objective function. The deployment data includes at least the beam matrix at the base station within the predetermined area, the phase matrix at the reflector, the position of the UAV, the value of artificial noise, the signal-to-noise ratio for each user, the signal-to-noise ratio for each eavesdropper, the weight for each user, the precoding vector for each user, the predetermined phase matrix, and the total number of users. Based on the deployment data, constructing an objective function to optimize the deployment data with the objective of maximizing the ratio of total secure communication rate to total communication energy consumption includes: determining the first sub-objective function using the following formula: in, R represents the maximum ratio of total secure communication rate to total communication power consumption. total P represents the total secure communication rate. total η represents the total communication energy consumption. k Let γ be the weight of the k-th user. k Let S be the signal-to-noise ratio of the k-th user. Let W be the signal-to-noise ratio (SNR) of the eavesdropper corresponding to the k-th user, where the SNR of each user is jointly determined by the beam matrix at the base station, the phase matrix at the reflector, the position of the drone, and the value of the artificial noise. W is the beam matrix at the base station, μ is the phase matrix at the reflector, L is the position of the drone, q is the artificial noise, K is the total number of users, B is the total number of base stations, and w... b,j For the predetermined beam matrix from the b-th base station to the j-th user, q AN The artificial noise emitted by the UAV is used; the second sub-objective function is determined by the following formula: Where, η k Let K be the weight of the k-th user, and K be the total number of users. This represents a pre-constructed second formula. This represents a pre-constructed third formula. This represents the first formula that has been pre-constructed. This represents the pre-constructed fourth formula. W obtained in the t-th training round k The value of W, where U represents the predetermined phase matrix. k Let t be the pre-encoded vector of the k-th user, and t be the training round.
[0007] Optionally, the constraints include: a security rate threshold constraint, a transmission rate constraint, and a unity modulus constraint. The deployment data includes at least the beam matrix at the base station within the predetermined area, the phase matrix at the reflector, the location of the UAV, the value of artificial noise, the signal-to-noise ratio for each user, the signal-to-noise ratio for each eavesdropper, the weight for each user, the precoding vector for each user, the security rate threshold, the maximum transmission power, the total number of base stations within the predetermined area, the unity modulus of the reflector, and the total number of reflectors. Based on the deployment data, the constraints corresponding to the objective function are constructed, including: determining the security rate threshold constraint using the following formula: Where, η k Let γ be the weight of the k-th user. k Let be the signal-to-noise ratio (SNR) of the k-th user, where the SNR of each user is jointly determined by the beam matrix at the base station, the phase matrix at the reflector, the position of the UAV, and the value of the artificial noise. R represents the signal-to-noise ratio of the eavesdropper corresponding to the k-th user, where K is the total number of users, and R is the signal-to-noise ratio of the eavesdropper. s The security rate threshold is defined; the transmission rate constraint is determined using the following formula: Among them, w b,j Let P be the precoding vector for communication between the b-th base station and the j-th user. b Where B is the maximum transmission power threshold, and B is the total number of base stations within the predetermined area; the unit modulus constraint is determined by the following formula: Where, |μ n | represents the unit modulus of the reflective surface, n represents the number of rows in the matrix, and N represents the total number of reflective surfaces.
[0008] Optionally, the objective function includes a first sub-objective function and a second sub-objective function; the step of solving the objective function based on the constraints to obtain the maximized ratio and the optimized deployment data includes: performing multiple rounds of alternating solution operations on the first sub-objective function and the second sub-objective function based on the constraints to obtain the maximized ratio and the optimized deployment data.
[0009] Optionally, the ratio includes a first target ratio and a second target ratio, and the optimized deployment data includes first target deployment data corresponding to the first target ratio and second target deployment data corresponding to the second target ratio; the step of performing multiple rounds of alternating solution operations on the first sub-objective function and the second sub-objective function based on the constraints to obtain the maximized ratio and the optimized deployment data includes: each round of alternating solution operation is performed as follows: based on the constraints, the first sub-objective function is iteratively solved multiple times to obtain the first ratio and the first deployment data corresponding to the current round of alternating solution operation, and based on the constraints, the ratio is further iterated. The second sub-objective function undergoes multiple rounds of individual solution operations to obtain the second ratio and the second deployment data corresponding to the current round of alternating solution operations; in response to determining that the number of alternating solutions executed is less than the predetermined number of alternating solutions, the next round of alternating solution operations is executed; in response to determining that the number of alternating solutions executed is greater than or equal to the predetermined number of alternating solutions, the first ratio corresponding to the current round of alternating solution operations is determined as the first target ratio, the first deployment data is determined as the first target deployment data, the second ratio is determined as the second target ratio, the second deployment data is determined as the second target deployment data, and the multi-round alternating solution operations are exited.
[0010] Optionally, the step of solving the first sub-objective function iteratively multiple times based on the constraints includes: solving the first sub-objective function iteratively multiple times using a deep deterministic policy gradient algorithm based on the constraints.
[0011] Optionally, the step of performing multiple rounds of individual solution operations on the second sub-objective function based on the constraints to obtain the second ratio and second deployment data corresponding to the current round of alternating solution operations includes: each round of individual solution operations is performed as follows: based on the constraints, the second sub-objective function is solved using a convex optimization algorithm to obtain the second ratio corresponding to the current round of individual solution operations; in response to determining that the difference between the second ratio corresponding to the current round of individual solution operations and the second ratio corresponding to the previous round of individual solution operations is greater than a predetermined rate difference, the next round of individual solution operations is executed; in response to determining that the difference between the second ratio corresponding to the previous round of individual solution operations is less than or equal to the predetermined rate difference, the second ratio corresponding to the current round of individual solution operations is determined as the second ratio corresponding to the current round of alternating solution operations, and the second deployment data corresponding to the current round of individual solution operations is determined as the second deployment data corresponding to the current round of alternating solution operations, and the multiple rounds of individual solution operations are exited.
[0012] Based on the same inventive concept, this application also provides a drone-assisted deployment data optimization device with enhanced reflectivity, comprising: an acquisition module configured to acquire deployment data corresponding to drones within a predetermined area; a construction module configured to construct an objective function for optimizing the deployment data based on the deployment data, with the objective of maximizing the ratio of total secure communication rate to total communication energy consumption, and construct constraints corresponding to the objective function; and a solution module configured to solve the objective function based on the constraints to obtain the maximized ratio and the optimized deployment data.
[0013] Based on the same inventive concept, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.
[0014] Based on the same inventive concept, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing a computer to perform the method described above.
[0015] As described above, the UAV-assisted deployment data optimization method and related equipment with enhanced reflective surface provided in this application include: acquiring deployment data of UAVs within a predetermined area; constructing an objective function to optimize the deployment data with the goal of maximizing the ratio of total secure communication rate to total communication energy consumption, and constructing constraints corresponding to the objective function to achieve the purpose of optimizing the deployment data, making the optimized deployment data reasonable while meeting user needs; and solving the objective function based on the constraints to obtain the maximized ratio and optimized deployment data, achieving a dynamic balance between total secure communication rate and total communication energy consumption in decellularized networks, thus providing a data foundation for subsequent network deployment. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the drone-assisted deployment data optimization method with enhanced reflective surface according to an embodiment of this application.
[0018] Figure 2This is a flowchart illustrating another embodiment of the UAV-assisted deployment data optimization method with enhanced reflective surface according to this application;
[0019] Figure 3 This is a schematic diagram illustrating how the safety and energy efficiency of another embodiment of this application changes with training rounds;
[0020] Figure 4 This is a schematic diagram comparing the average safety and energy efficiency of the present solution and the comparative solution in another embodiment of this application;
[0021] Figure 5 This is a schematic diagram illustrating the variation of average safety energy efficiency with the number of RIS in yet another embodiment of this application;
[0022] Figure 6 A schematic diagram of the structure of the drone-assisted deployment data optimization device with enhanced reflective surface according to an embodiment of this application;
[0023] Figure 7 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0025] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0026] As described in the background section, in recent years, thanks to their flexible mobility and deployment capabilities, drones have been widely used in fields such as remote monitoring, photography, traffic control, and cargo transportation. Compared with traditional terrestrial communication networks, drones can obtain higher-quality line-of-sight channels through reasonable flight trajectory design, thereby greatly improving communication performance. Utilizing drones as aerial mobile base stations or relay nodes provides flexible and efficient wireless coverage and communication services to ground users. Their core advantages lie in their high mobility, rapid deployment, and good line-of-sight transmission conditions. Drones can carry communication equipment and improve spectrum efficiency and reduce signal obstruction by dynamically adjusting their flight trajectories and optimizing network performance. Furthermore, drones can work collaboratively with satellites, terrestrial base stations, and smart reflectors. In particular, in cell-free network cells, using drones as aerial platforms, leveraging their high mobility and ease of deployment, can effectively improve the coverage and communication quality of cell-free communication systems. Meanwhile, in recent years, smart reflectors have shown great potential in improving the wireless propagation environment. Deploying smart reflectors in cell-free networks can effectively reduce the base station's transmission power while achieving high service quality. Smart reflectors are a novel wireless communication technology composed of programmable metasurfaces. By dynamically adjusting the phase, amplitude, and reflection direction of electromagnetic waves, they optimize the wireless channel environment, enhancing signal coverage, suppressing interference, and improving energy efficiency. Smart reflectors require no active radio frequency components; they adjust reflection parameters solely through intelligent algorithms (such as AI or optimized control). In traditional research on smart reflectors-assisted cell-free networks, they are typically placed within terrestrial networks. Due to the inherent characteristics of smart reflectors, reflected signals can only cover one side of the reflector, potentially failing to cover all users.
[0027] However, in cell-free networks, due to the broadcast nature of wireless communication, legitimate users' signals are easily leaked to eavesdroppers, and drones' line-of-sight channels are more vulnerable to malicious attacks from the ground. Smart reflectors can enhance the signal received by legitimate users while suppressing the signal strength received by eavesdroppers, effectively improving network security. Furthermore, smart reflectors do not amplify noise or cause self-interference, significantly reducing energy consumption. Therefore, using multiple smart reflectors to enhance the secure transmission performance of drone-assisted cell-free networks is a promising solution.
[0028] However, in existing technologies, RIS (Radio Signal Enhancement) combined with UAVs is typically not deployed in terrestrial networks, relying solely on UAVs for signal enhancement. Proper deployment of UAVs and RIS is crucial, impacting network communication quality and energy consumption. Inappropriate network deployment can lead to poor coverage and communication quality in cell-free networks, resulting in low communication security. Furthermore, improving communication quality may require more energy, necessitating more efficient energy allocation strategies to avoid interference while ensuring communication quality, thus leading to excessive energy consumption.
[0029] In view of this, embodiments of this application propose a drone-assisted deployment data optimization method with enhanced reflective surface, referring to... Figure 1 This includes the following steps:
[0030] Step 101: Obtain deployment data for drones within the designated area.
[0031] In this step, existing technologies utilize cellular networks to provide communication for a predetermined area, dividing the area into multiple adjacent hexagonal "cells," each with signal service provided by a base station. However, cellular networks have relatively high energy consumption, leading to the development of decellularized communication methods. This application utilizes drones and smart reflectors to deploy decellularized networks. The combination of drones and smart reflectors can act as mobile relay nodes, reflecting incident signals from ground base stations to target users. The drones can also act as airborne base stations, emitting artificial noise towards target users and potential eavesdroppers to improve network security. Specifically, the smart reflectors are deployed on the ground within the predetermined area and mounted on drones. The predetermined area is the region where the network is to be deployed; for example, it can be a decellularized cell within the network to be deployed. Within the predetermined area, there are multiple deployment schemes for drones and smart reflectors. First, deployment data corresponding to drones within the predetermined area is acquired. This deployment data is the data used for network deployment using drones and smart reflectors, and it represents one of the multiple deployment schemes. Subsequent optimization of the deployment data is needed to find the optimal deployment data that meets user needs.
[0032] Step 102: Based on the deployment data, construct an objective function to optimize the deployment data with the goal of maximizing the ratio of total secure communication rate to total communication energy consumption, and construct the corresponding constraints for the objective function.
[0033] In this step, after acquiring the deployment data, it's important to note that the current deployment data may cause a conflict between the total secure communication rate and the total communication energy consumption. This conflict manifests as either increasing the total secure communication rate at the expense of increasing total communication energy consumption, or decreasing total communication energy consumption at the expense of relatively low communication security. Therefore, to maximize the security and energy efficiency of communication, an objective function for optimizing the deployment data is constructed, with the goal of maximizing the ratio of the total secure communication rate to the total communication energy consumption. To solve the objective function, corresponding constraints also need to be constructed. By constructing the objective function and constraints, the deployment data is optimized, ensuring that the optimized deployment data is reasonable while meeting user needs.
[0034] Step 103: Based on the constraints, solve the objective function to obtain the maximized ratio and optimized deployment data.
[0035] In this step, the objective function is solved using constraints to obtain the maximized ratio and optimized deployment data. The optimized deployment data maximizes the ratio of total secure communication rate to total communication energy consumption. Given the constraints, the objective function and constraints are solved jointly using the constraints themselves, ensuring that the solution satisfies both the objective function and the constraints. Subsequently, the optimized deployment data is used to deploy the network within a predetermined area. For example, the deployment data includes the 3D coordinates of the UAV and the intelligent reflector. Using all the 3D coordinates, the UAV and intelligent reflector are deployed to ensure that subsequent communication maximizes both the total secure communication rate and minimizes the total communication energy consumption. Solving the objective function using constraints resolves the conflict between total secure communication rate and total communication energy consumption, achieving a dynamic balance between these two factors in decellularized networks, and providing a data foundation for subsequent network deployment.
[0036] The above scheme acquires deployment data for drones within a predetermined area. Based on this deployment data, an objective function is constructed to optimize the deployment data, aiming to maximize the ratio of total secure communication rate to total communication energy consumption. Constraints on this objective function are also constructed to optimize the deployment data, ensuring that the optimized data is reasonable while meeting user needs. Based on these constraints, the objective function is solved to obtain the maximized ratio and the optimized deployment data, achieving a dynamic balance between total secure communication rate and total communication energy consumption in decellularized networks. This provides a data foundation for subsequent network deployment.
[0037] In some embodiments, the objective function includes a first sub-objective function and a second sub-objective function. The deployment data includes at least the beam matrix at the base station within the predetermined area, the phase matrix at the reflector, the location of the UAV, the value of artificial noise, the signal-to-noise ratio for each user, the signal-to-noise ratio for each eavesdropper, the weight for each user, the precoding vector for each user, the predetermined phase matrix, and the total number of users. Based on the deployment data, constructing an objective function to optimize the deployment data with the objective of maximizing the ratio of total secure communication rate to total communication energy consumption includes: determining the first sub-objective function using the following formula: in, R represents the maximum ratio of total secure communication rate to total communication power consumption. total P represents the total secure communication rate. total η represents the total communication energy consumption. k Let γ be the weight of the k-th user. k Let S be the signal-to-noise ratio of the k-th user. Let W be the signal-to-noise ratio (SNR) of the eavesdropper corresponding to the k-th user, where the SNR of each user is jointly determined by the beam matrix at the base station, the phase matrix at the reflector, the position of the drone, and the value of the artificial noise. W is the beam matrix at the base station, μ is the phase matrix at the reflector, L is the position of the drone, q is the artificial noise, K is the total number of users, B is the total number of base stations, and w... b,j For the predetermined beam matrix from the b-th base station to the j-th user, q AN The artificial noise emitted by the UAV is used; the second sub-objective function is determined by the following formula: Where, η k Let K be the weight of the k-th user, and K be the total number of users. This represents a pre-constructed second formula. This represents a pre-constructed third formula. This represents the first formula that has been pre-constructed. This represents the pre-constructed fourth formula. W obtained in the t-th training round k The value of W, where U represents the predetermined phase matrix. k Let t be the pre-encoded vector of the k-th user, and t be the training round.
[0038] In this embodiment, the objective function is constructed with the goal of maximizing the ratio of total secure communication rate to total communication energy consumption. The formula for the total secure communication rate in the objective function is as follows: The formula for total communication energy consumption in the objective function is: In this system, the artificial noise is emitted by the drone and can be manually set; the reflector is an intelligent reflector. However, due to the coupling of multiple variables in the objective function, which is a non-convex function, it is difficult to directly solve the objective function. Furthermore, ignoring some terms unrelated to the optimization variables does not affect the final simulation results. The original problem can be rewritten as a first subproblem and an intermediate subproblem. The intermediate subproblem is then used to derive a second subproblem. The first subproblem includes a first objective function and constraints, and the second subproblem includes a second objective function and constraints. The second subproblem is a convex function, facilitating the solution of both the first and second subproblems. Solving the first and second subproblems yields the beam matrix at the base station, the drone's position, the value of the artificial noise, and the phase matrix at the reflector in the deployment data. The signal-to-noise ratio (SNR) for each user and the SNR for each user's corresponding eavesdropper are jointly determined by the beam matrix at the base station, the phase matrix at the reflector, the drone's position, and the value of the artificial noise. The phase matrix at the intelligent reflector is represented as follows: Where θ r,n The nth reflection unit of the r-th RIS can be represented as:
[0039] The formula for determining the signal-to-noise ratio for each user is as follows: Where, μ H This is a pre-constructed matrix related to the phase matrix at the smart reflector. For line-of-sight path channel components, For non-line-of-sight path channel components, w k Let w be the beam matrix at the base station corresponding to the k-th user. j Let σ be the beam matrix at the base station corresponding to the j-th user. k Let q be the variance of the thermal noise at the k-th user. AN Artificial noise emitted by drones. For the channel gain from the drone to the k-th user, The product of the channel-dependent channel gains constructed for the smart reflector. For the channel gain from the r-th RIS to the k-th user, G b,r Let diag(x) be the channel gain from the b-th base station to the r-th RIS, and let diag(x) represent an N*N diagonal matrix constructed using x. Let b be the channel gain from the b-th base station to the k-th user. These are the diagonal elements of the RIS phase matrix. Where β0 represents the path loss when the reference distance is 1m, d rThe distance between the signal transmitter and receiver is represented by α, the path loss exponent is represented by λ, the wavelength of the transmitted signal is represented by K, and the Ricean factor is represented by l. M Let M be the distance between the base station and the Mth element at the smart reflective surface.
[0040] The formula for determining the signal-to-noise ratio of each user's corresponding eavesdropper is as follows: Among them, h e For the pre-constructed formula related to the eavesdropper's channel gain, σ e The variance of thermal noise for the eavesdropper. Channel gain from drones to eavesdroppers.
[0041] The first formula pre-constructed in the second sub-objective function is: The pre-constructed second formula is determined by the following formula:
[0042]
[0043] The pre-constructed third formula is determined by the following formula:
[0044]
[0045] The pre-constructed fourth formula is Where Tr represents the trace of the matrix, and U = μμ H , for h k The conjugate transpose of . for h e The conjugate transpose of .
[0046] The derivation process of the signal-to-noise ratio for each user and the signal-to-noise ratio for the corresponding eavesdropper is as follows:
[0047] Assume the network deploys B base stations, K users, R RIS (Radio Routers), one eavesdropper, and one UAV (User Avatar). The direct channel gains from the b-th base station to the k-th user and the eavesdropper are respectively... The direct channel gain from the UAV to the k-th user and the eavesdropper is respectively... Let represent the channel gain from the b-th base station to the RIS, from the RIS to the k-th user, and from the RIS to the eavesdropper, respectively.
[0048] The signal transmitted at base station b is:
[0049] Among them, w b,k It is the precoding vector from the b-th base station to the k-th user, s kIt is the signal sent by the base station to the k-th user, and for this typical scenario, E{|s k | 2} = 1.
[0050] Since the reflection channel constructed by RIS can bypass obstacles and establish a Loss of Position (LoS) link, the Ricean channel model is used for modeling:
[0051] in, Indicates the non-line-of-sight path channel components. It can be represented as:
[0052]
[0053] The artificial noise emitted by a UAV can be expressed as q AN Therefore, the received signal at the k-th user can be expressed as:
[0054]
[0055] Where, n k Let be the thermal noise at the k-th user.
[0056] Similarly, the signal that the eavesdropper is listening to from the k-th user is:
[0057]
[0058] according to
[0059]
[0060] Where (a) is defined by θ r =[θ r,1 ,…,θ r,N ] T This can be achieved by defining the following equations: (b) and (c)
[0061]
[0062] Similarly, we can conclude that:
[0063]
[0064] According to the above equation, the SINR of the k-th user and the eavesdropper eavesdropping on the k-th user's information can be expressed as:
[0065]
[0066] The process of solving the first sub-objective function using the DDPG algorithm is as follows: First, in order to apply the reinforcement learning algorithm, the first subproblem P is... 0This is expressed as a Markov Decision Process (MDP). The MDP framework consists of four key components: state space S, action space A, transition probabilities, and reward function R(s). t ,a t In an MDP (Multi-Active Programming Principle), the agent interacts with the environment and obtains optimization results. The state space, action space, and reward function of the MDP will be introduced below.
[0067] State space: at step t, state s t Includes channel status information and total network power consumption {P} total The state is defined as follows:
[0068] s t =({h(t)},{P total}) Formula (20),
[0069] Where h(t) represents the set of channel gains for each link, i.e., h(t) = {(F r,k (t)) real ,(F r,k (t)) imag ,(H a,k (t)) real ,(H a,k (t)) imag ,(G b,r ) real ,(G b,r ) imag ,(H a,e ) real ,(H a,e ) imag},(X(t)) real 、(X(t)) imag Let represent the real and imaginary parts of the complex matrix, respectively. This comprehensive state representation provides all the necessary information for the decision at step t, including the current channel conditions and energy consumption.
[0070] Action space: In step t, the agent's action a t The location of the UAV, the beamforming matrix of the base station, and the artificial noise vector are represented as follows:
[0071]
[0072] Where L(t) represents the position of the UAV at step t. q represents the beamforming matrix of the base station. AN (t) represents the artificial noise transmitted by the UAV. This action determines how the network's security efficiency changes within step t.
[0073] Reward Function: The immediate reward function aims to reflect the optimization objective while incorporating relevant constraints. To encourage agents to optimize system performance and comply with constraints, the reward function is defined as follows:
[0074]
[0075] Among them, U(s) t ,a t λ represents the objective function value given the state and actions, indicating the system's performance in terms of safety and energy efficiency. i Let Penalty represent the penalty term coefficient for constraint i, and Penalty... i (s t ,a t The penalty for violating constraint i at step t is quantified and can be expressed as:
[0076]
[0077] The above equation guarantees that when constraint i is satisfied, f i (s t ,a t The value of f is 0 when the constraint is violated. i (s t ,a t The value of ) is positive. By subtracting the penalty term in the reward function from the objective function value, any action that violates the constraints will reduce the immediate reward, thereby preventing the agent from taking actions that violate the constraints. This scheme has two penalty terms: one is a safe rate constraint, during training the agent must ensure that the system's safe rate is greater than a minimum threshold; the other is a power constraint, the agent must ensure that the power consumed by the base station is less than the base station's maximum transmission power. Therefore, the reward function encourages the agent to minimize constraint violations while maximizing the objective function U(s). t ,a t This effectively balances the relationship between performance optimization and constraint satisfaction.
[0078] The first subproblem is derived by rewriting the original problem, which is:
[0079]
[0080] Next, optimize the phase matrix at RIS. It's important to note that ignoring some terms unrelated to the optimization variable μ in the original problem will not affect the final simulation results. Therefore, the original problem can be rewritten as an intermediate subproblem:
[0081]
[0082] P 1 This is an intermediate subproblem.
[0083] To solve the aforementioned intermediate sub-problem, we first define... and U = uu H Then, ignoring the rank-one constraint, the problem can be represented as:
[0084]
[0085] The formula pre-constructed for each round is as follows:
[0086]
[0087] By fixing W k and According to Taylor expansion, F2 and F3 can be expressed as:
[0088]
[0089]
[0090] Therefore, the optimization problem of the phase matrix U can be rewritten as:
[0091]
[0092] Second subproblem P 2 This is a standard convex optimization problem, which can be solved using existing convex optimization algorithms.
[0093] In some embodiments, the constraints include: a security rate threshold constraint, a transmission rate constraint, and a unity modulus constraint. The deployment data includes at least the beam matrix at the base station within the predetermined area, the phase matrix at the reflector, the location of the UAV, the value of artificial noise, the signal-to-noise ratio (SNR) for each user, the SNR for each eavesdropper, the weight for each user, the precoding vector for each user, the security rate threshold, the maximum transmission power, the total number of base stations within the predetermined area, the unity modulus of the reflector, and the total number of reflectors. Based on the deployment data, the constraints corresponding to the objective function are constructed, including: determining the security rate threshold constraint using the following formula: Where, η k Let γ be the weight of the k-th user. k Let be the signal-to-noise ratio (SNR) of the k-th user, where the SNR of each user is jointly determined by the beam matrix at the base station, the phase matrix at the reflector, the position of the UAV, and the value of the artificial noise. R represents the signal-to-noise ratio of the eavesdropper corresponding to the k-th user, where K is the total number of users, and R is the signal-to-noise ratio of the eavesdropper. s The security rate threshold is defined; the transmission rate constraint is determined using the following formula: Among them, wb,j Let P be the precoding vector for communication between the b-th base station and the j-th user. b Where B is the maximum transmission power threshold, and B is the total number of base stations within the predetermined area; the unit modulus constraint is determined by the following formula: Where, |μ n | represents the unit modulus of the reflective surface, n represents the number of rows in the matrix, and N represents the total number of reflective surfaces.
[0094] In this embodiment, since the objective function aims to maximize the ratio of total secure communication rate to total communication energy consumption, the total secure communication rate should be as large as possible, and the total communication energy consumption should be as small as possible. A secure communication rate constraint is constructed to ensure that the total secure communication rate is greater than or equal to a secure communication rate threshold, which is determined based on historical experience. A transmission rate constraint is also constructed to ensure that the transmission power of all user communications is less than or equal to a maximum transmission power threshold, which is also determined based on historical experience. The precoding vector optimizes signal transmission performance by adjusting the signal's phase and amplitude, while the transmission power directly affects the signal's energy and coverage. In practical systems, the precoding vector and transmission power need to be jointly designed to meet the goals of power constraints and performance optimization. Through reasonable precoding and power allocation, the system's spectral efficiency, power efficiency, and reliability can be significantly improved. The unit modulus of the phase matrix of the intelligent reflector (i.e., the magnitude of the reflection coefficient is 1) is used as a constraint for three main reasons: First, most practical RIS systems employ passive designs, where the reflection unit can only adjust the phase and cannot actively amplify the signal amplitude; the unit modulus constraint aligns with the physical characteristics of passive reflection. Second, this constraint ensures that no additional energy loss is introduced during signal reflection, conforming to the principle of energy conservation and maintaining system energy efficiency. Finally, the unit modulus simplifies hardware implementation (e.g., requiring only a phase-adjustable metasurface structure) and reduces the complexity of the optimization problem (reducing degrees of freedom). Through the safety rate threshold constraint, transmission rate constraint, and unit modulus constraint, the solution to the objective function is restricted to a range that meets the requirements, ensuring the feasibility and practicality of the solution.
[0095] In some embodiments, the objective function includes a first sub-objective function and a second sub-objective function; the step of solving the objective function based on the constraints to obtain the maximized ratio and the optimized deployment data includes: performing multiple rounds of alternating solution operations on the first sub-objective function and the second sub-objective function based on the constraints to obtain the maximized ratio and the optimized deployment data.
[0096] In this embodiment, since the first and second sub-problems are rewritten from the original problem, the constraints of the first and second sub-objective functions are the same. Because the optimized deployment data requires continuous iteration, multiple rounds of alternating solution operations are performed on the first and second sub-objective functions based on the constraints to obtain the maximized ratio and optimized deployment data. The first and second sub-objective functions and their constraints are solved jointly, ensuring that the solution satisfies both the objective function and the constraints. Through alternating solution operations, the first and second sub-objective functions achieve sufficient iteration, resulting in a more accurate maximum optimized ratio and optimized deployment data.
[0097] In some embodiments, the ratio includes a first target ratio and a second target ratio, and the optimized deployment data includes first target deployment data corresponding to the first target ratio and second target deployment data corresponding to the second target ratio; the step of performing multiple rounds of alternating solution operations on the first sub-objective function and the second sub-objective function based on the constraints to obtain the maximized ratio and the optimized deployment data includes: each round of alternating solution operation is performed as follows: based on the constraints, the first sub-objective function is iteratively solved multiple times to obtain the first ratio and the first deployment data corresponding to the current round of alternating solution operation, and based on the constraints... The second sub-objective function is solved individually in multiple rounds to obtain the second ratio and the second deployment data corresponding to the current round of alternating solution operation. In response to determining that the number of alternating solutions executed is less than the predetermined number of alternating solutions, the next round of alternating solution operation is executed. In response to determining that the number of alternating solutions executed is greater than or equal to the predetermined number of alternating solutions, the first ratio corresponding to the current round of alternating solution operation is determined as the first target ratio, the first deployment data is determined as the first target deployment data, the second ratio is determined as the second target ratio, and the second deployment data is determined as the second target deployment data, and the multi-round alternating solution operation is exited.
[0098] In this embodiment, each round of alternating operations is performed as follows: Based on the constraints, the first sub-objective function is solved iteratively multiple times to obtain the first ratio and first deployment data corresponding to the current round of alternating operations. When the number of iterations reaches a predetermined number, based on the constraints, the second sub-objective function is solved separately for multiple rounds to obtain the second ratio and second deployment data corresponding to the current round of alternating operations. Although only one round of alternating operations has already iterated the first and second sub-objective functions extensively, the iteration is not sufficient; therefore, multiple rounds of alternating operations are required. The predetermined number of iterations is determined based on historical experience.
[0099] When the number of alternating solutions performed is greater than or equal to the predetermined number of alternating solutions, it indicates that the first and second sub-objective functions have reached sufficient iteration. The first ratio corresponding to the current round of solution is determined as the first target ratio, the first deployment data is determined as the first target deployment data, the second ratio is determined as the second target ratio, and the second deployment data is determined as the second target deployment data. Here, the first target ratio and the second target ratio are the maximized ratios, and the first target deployment data and the second target deployment data are the optimized deployment data. The predetermined number of alternating solutions is determined based on historical experience. By performing multiple rounds of iterative solutions on the first sub-objective function and multiple rounds of separate solutions on the second sub-objective function, the first and second sub-objective functions undergo deep iteration, resulting in more accurate first target ratios, second target ratios, first target deployment data, and second target deployment data.
[0100] In some embodiments, the step of solving the first sub-objective function iteratively based on the constraints includes: solving the first sub-objective function iteratively using a deep deterministic policy gradient algorithm based on the constraints.
[0101] In this embodiment, multiple variables are coupled in the numerator and denominator of the first sub-objective function, making the first sub-problem non-convex. The first sub-problem includes the first sub-objective function and constraints. Traditional convex optimization techniques are difficult to use to solve this problem; therefore, the Deep Deterministic Policy Gradient (DDPG) algorithm is used to iteratively solve the first sub-objective function. The DDPG algorithm efficiently solves the objective function of continuous control problems because it combines the representational power of deep learning with the optimization direction of deterministic policy gradients. The DDPG algorithm implements policy iteration through an Actor-Critic architecture: the Actor network directly outputs deterministic actions, avoiding the variance problem of stochastic policies; the Critic network evaluates the value of actions and provides gradient feedback, guiding the policy update direction. An experience replay mechanism breaks data correlation, soft updates of the target network stabilize the training process, and the introduction of action noise ensures thorough exploration. This design allows DDPG to utilize deep networks to handle high-dimensional state spaces while ensuring convergence through the policy gradient theorem, making it particularly suitable for tasks requiring precise continuous actions. By using the deep deterministic policy gradient algorithm to solve the first sub-objective function multiple times, it is ensured that the solution of the first sub-objective function gradually improves towards the optimal direction while satisfying the constraints.
[0102] In some embodiments, the step of performing multiple rounds of individual solution operations on the second sub-objective function based on the constraints to obtain the second ratio and second deployment data corresponding to the current round of alternating solution operations includes: each round of individual solution operations is performed as follows: based on the constraints, the second sub-objective function is solved using a convex optimization algorithm to obtain the second ratio corresponding to the current round of individual solution operations; in response to determining that the difference between the second ratio corresponding to the current round of individual solution operations and the second ratio corresponding to the previous round of individual solution operations is greater than a predetermined rate difference, the next round of individual solution operations is executed; in response to determining that the difference between the second ratio corresponding to the previous round of individual solution operations is less than or equal to the predetermined rate difference, the second ratio corresponding to the current round of individual solution operations is determined as the second ratio corresponding to the current round of alternating solution operations, and the second deployment data corresponding to the current round of individual solution operations is determined as the second deployment data corresponding to the current round of alternating solution operations, and the multiple rounds of individual solution operations are exited.
[0103] In this embodiment, each round of individual solution operation is performed as follows: A convex optimization algorithm is used to solve the second sub-objective function, obtaining the second ratio and second deployment data corresponding to the current round of individual solution operation. The convex optimization algorithm can be a combination of a semi-definite relaxation (SDR) algorithm and a Gaussian random algorithm. If the difference between the second ratio corresponding to the current round of individual solution operation and the second ratio corresponding to the previous round of individual solution operation is greater than a predetermined rate difference, it indicates that the solution of the second sub-objective function corresponding to the current round of individual solution operation has not yet reached sufficient iteration, and the next round of individual solution operation is executed. If the difference between the second ratio corresponding to the current round of individual solution operation and the second ratio corresponding to the previous round of individual solution operation is less than or equal to the predetermined rate difference, it indicates that the solution of the second sub-objective function corresponding to the current round of individual solution operation has reached sufficient iteration, the second ratio corresponding to the current round of individual solution operation is determined as the second ratio corresponding to the current round of alternating solution operation, the second deployment data corresponding to the current round of individual solution operation is determined as the second deployment data corresponding to the current round of alternating operation, and the multi-round individual solution operation is terminated. Through multiple rounds of individual solution operations, the solution of the second sub-objective function corresponding to each round of individual solution operations is fully iterated, and under the condition of satisfying the constraints, the solution of the second sub-objective function is gradually improved towards the optimal direction.
[0104] In another embodiment provided in this application, such as Figure 2As shown, the initialization information includes setting the channel state information, number of RIS, number of users, etc., through the initialization function, and initializing the optimization variables UAV location, base station beamforming, and RIS phase matrix. To apply the reinforcement learning algorithm, the reinforcement learning algorithm parameters and state space are set simultaneously with the above initialization.
[0105] The algorithm checks if the number of alternating solutions executed so far is less than the predetermined number of alternating solutions. If the number of alternating solutions executed so far is less than the predetermined number, the environment information and optimization variables are reset. The environment information and the optimization variables trained by reinforcement learning are reset first to ensure the accuracy of the iteration process. For solving the first sub-objective function, the algorithm checks if the number of iterations is less than the set maximum number of iterations. If the number of iterations is less than the set maximum number of iterations, an action is selected from the initial space, the reward after the selected action is calculated, and the environment information is updated. The environment information and deployment data for this round are stored. Reinforcement learning algorithms are applied to optimize UAV location, base station beamforming, and the RIS phase matrix. During the iteration process, the action to be executed in this round is first selected based on the changes in the state space, then the corresponding reward is calculated based on the selected action, and the state space is updated and relevant information is stored. When the number of iterations exceeds the set maximum number of iterations, the reward function value is saved.
[0106] If the number of alternating solutions performed is greater than or equal to a predetermined number, for the second sub-objective function: the SDR algorithm and Gaussian stochastic optimization algorithm are used to solve the problem, and the deployment data is updated. It is then determined whether the difference between the second ratio in the updated deployment data and the second ratio corresponding to the previous round of individual solution operations is greater than a predetermined rate difference. If it is greater than the predetermined rate difference, the next round of individual solution operations is executed. Convex optimization theory is used to solve the second sub-objective function, such as SDR and Gaussian randomization. The second sub-objective function is transformed into a standard positive semidefinite programming problem using SDR and solved using the CVX toolbox. To satisfy the rank-one constraint, Gaussian randomization is used to obtain a feasible solution to the original problem. During convex optimization, a small convergence threshold is set to prevent the algorithm from stopping before reaching the global optimum, making the optimization result closer to the global optimum. After the convex optimization process ends, the RIS phase matrix is saved, and the above process is repeated until the number of training rounds is greater than the set maximum number of training rounds. If the difference is less than or equal to the predetermined rate difference, it is further determined whether the number of alternating solutions performed is less than the predetermined number of alternating solutions. If the number of alternating solutions performed is greater than or equal to the predetermined number of alternating solutions, output the optimized safety and energy efficiency and the optimized deployment data.
[0107] In another embodiment provided in this application, the simulation effect of this application can be further illustrated through simulation. In the simulation example, a total of 3 ground base stations are deployed in the predetermined area. The location of the b-th base station is set to (10(b-1), 0, 3)m, and each base station is equipped with two antennas. The 3 users are distributed on a circle with a center of (120, 20)m and a radius of 3m, at a height of 1.5m. The two ground RIS are deployed at (100, 50, 6)m and (120, 50, 6)m respectively, with the number of components N = 30. The eavesdropper is deployed at (90, 30, 1.5)m, and the initial position of the UAV is deployed at (60, 40, 10)m. The noise power is set to σ. k =σ e = -70dB. A simulation example introduces a comparative scheme to highlight the superiority of the proposed scheme. This scheme is a multi-RIS enhanced UAV-cell-free network, while the comparative scheme is a multi-RIS assisted decellularized network.
[0108] like Figure 3 As shown, Figure 3 This is a schematic diagram illustrating the change in safety and energy efficiency with training rounds according to another embodiment of this application. Figure 3 The horizontal axis represents the number of training rounds, which reflects the number of iterations in the alternating solution operation. The vertical axis represents the average security energy efficiency, which is determined by the ratio of the total secure communication rate to the total communication energy consumption. Figure 3 The results demonstrate that the average security efficiency gradually converges with increasing training epochs. Simulation results show that the proposed algorithm can effectively solve the aforementioned non-convex problem, and when the number of training epochs exceeds 300, the average security efficiency of the network approaches the convergence value, proving that the proposed algorithm can achieve convergence within a relatively short number of training epochs. These results indicate that the proposed scheme can effectively improve the security efficiency of the network and achieve a good balance between security performance and energy consumption in decellularized networks.
[0109] like Figure 4 As shown, Figure 4 This is a schematic diagram comparing the average safety and energy efficiency of the proposed solution and a comparative solution in another embodiment of this application. The proposed solution is a multi-RIS enhanced UAV-cell-free network. Figure 4 The blue line in the diagram represents the contrast solution, which is a multi-RIS-assisted decellularized network. Figure 4 The orange line in the middle. Figure 4The horizontal axis represents the number of training rounds, reflecting the number of iterations in the alternating solution operation, while the vertical axis represents the average security efficiency. In this comparative scheme, the multi-RIS-assisted cell-free network and this scheme have the same base station transmit power, number of RIS, and number of RIS units. Despite having the same number of RIS, this scheme achieves a higher security rate compared to the ground-based multi-RIS-assisted cell-free network. This is because this scheme can utilize the high mobility of UAVs to bring the RIS closer to network users, providing a higher transmission rate. Simultaneously, the artificial noise emitted by the UAVs can effectively suppress eavesdroppers from intercepting legitimate user information, thereby improving the security of information transmission.
[0110] like Figure 5 As shown, Figure 5 This is a schematic diagram illustrating the change in average security efficiency with the number of RIS components according to another embodiment of this application, showing the relationship between average security efficiency and the number of RIS components. Figure 5 The horizontal axis represents the number of training rounds, which reflects the number of iterations in the alternating solution operation. The vertical axis represents the average safety energy efficiency. Figure 5 The red line in the diagram indicates that the number of RIS is 5. Figure 5 The green line in the image indicates that the number of RIS is 10. Figure 5 The orange line in the image indicates that the number of RIS is 15. Figure 5 The blue line in the diagram represents a RIS count of 20. Experimental results show that the average security efficiency of the system significantly improves with the increase in the number of RIS. These results indicate that as the number of RIS increases, the RIS can better regulate the wireless channel, allowing legitimate users to receive more favorable signals. This improves the strength of the received signal while suppressing the transmitted signal power intercepted by eavesdroppers, effectively enhancing system security. Furthermore, introducing RIS effectively improves the network's system energy efficiency. This is because RIS consists of passive components and controls the reflected phase and amplitude of the incident signal via a control circuit board. Since the energy consumed by the control circuit is negligible, compared to traditional repeater circuits, RIS-assisted communication systems can effectively reduce energy consumption, thus extending the lifespan of devices such as UAVs. Moreover, the air platform formed by combining RIS and UAV can be adjusted according to the user's location, keeping the RIS closer to the user and allowing the user to receive a stronger signal. Therefore, this solution effectively improves network security while reducing energy consumption, achieving green communication.
[0111] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.
[0112] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0113] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides a drone-assisted deployment data optimization device with enhanced reflective surface.
[0114] refer to Figure 6 The reflective surface-enhanced UAV-assisted deployment data optimization device includes:
[0115] The acquisition module 10 is configured to acquire deployment data of drones within a predetermined area.
[0116] The construction module 20 is configured to construct an objective function to optimize the deployment data based on the deployment data, with the goal of maximizing the ratio of total secure communication rate to total communication energy consumption, and to construct the constraints corresponding to the objective function.
[0117] The solver module 30 is configured to solve the objective function based on the constraints to obtain the maximized ratio and optimized deployment data.
[0118] The aforementioned device acquires deployment data for drones within a predetermined area. Based on this deployment data, an objective function is constructed to optimize the deployment data, aiming to maximize the ratio of total secure communication rate to total communication energy consumption. Constraints corresponding to this objective function are also constructed to optimize the deployment data, ensuring that the optimized data is reasonable while meeting user needs. Based on the constraints, the objective function is solved to obtain the maximized ratio and the optimized deployment data, achieving a dynamic balance between total secure communication rate and total communication energy consumption in decellularized networks. This provides a data foundation for subsequent network deployment.
[0119] In some embodiments, the construction module 20 is further configured such that the objective function includes a first sub-objective function and a second sub-objective function, and the deployment data includes at least the beam matrix at the base station within the predetermined area, the phase matrix at the reflector, the position of the UAV, the value of artificial noise, the signal-to-noise ratio for each user, the signal-to-noise ratio for each eavesdropper, the weight for each user, the precoding vector for each user, the predetermined phase matrix, and the total number of users; the first sub-objective function is determined by the following formula: in, R represents the maximum ratio of total secure communication rate to total communication power consumption. total P represents the total secure communication rate. total η represents the total communication energy consumption. k Let γ be the weight of the k-th user. k Let S be the signal-to-noise ratio of the k-th user. Let W be the signal-to-noise ratio (SNR) of the eavesdropper corresponding to the k-th user, where the SNR of each user is jointly determined by the beam matrix at the base station, the phase matrix at the reflector, the position of the drone, and the value of the artificial noise. W is the beam matrix at the base station, μ is the phase matrix at the reflector, L is the position of the drone, q is the artificial noise, K is the total number of users, B is the total number of base stations, and w... b,j For the predetermined beam matrix from the b-th base station to the j-th user, q AN The artificial noise emitted by the UAV is used; the second sub-objective function is determined by the following formula: Where, η k Let K be the weight of the k-th user, and K be the total number of users. This represents a pre-constructed second formula. This represents a pre-constructed third formula. This represents the first formula that has been pre-constructed. This represents the pre-constructed fourth formula. W obtained in the t-th training round k The value of W, where U represents the predetermined phase matrix. k Let t be the pre-encoded vector of the k-th user, and t be the training round.
[0120] In some embodiments, the construction module 20 is further configured such that the constraints include: a security rate threshold constraint, a transmission rate constraint, and a unity modulus constraint. The deployment data includes at least the beam matrix at the base station within the predetermined area, the phase matrix at the reflector, the location of the UAV, the value of artificial noise, the signal-to-noise ratio for each user, the signal-to-noise ratio for each eavesdropper, the weight for each user, the precoding vector for each user, the security rate threshold, the maximum transmission power, the total number of base stations within the predetermined area, the unity modulus of the reflector, and the total number of reflectors. The security rate threshold constraint is determined by the following formula: Where, η k Let γ be the weight of the k-th user. k Let be the signal-to-noise ratio (SNR) of the k-th user, where the SNR of each user is jointly determined by the beam matrix at the base station, the phase matrix at the reflector, the position of the UAV, and the value of the artificial noise. R represents the signal-to-noise ratio of the eavesdropper corresponding to the k-th user, where K is the total number of users, and R is the signal-to-noise ratio of the eavesdropper. s The security rate threshold is defined; the transmission rate constraint is determined using the following formula: Among them, w b,j Let P be the precoding vector for communication between the b-th base station and the j-th user. b Where B is the maximum transmission power threshold, and B is the total number of base stations within the predetermined area; the unit modulus constraint is determined by the following formula: Where, |μ n | represents the unit modulus of the reflective surface, n represents the number of rows in the matrix, and N represents the total number of reflective surfaces.
[0121] In some embodiments, the solution module 30 is further configured such that the objective function includes a first sub-objective function and a second sub-objective function; based on the constraints, multiple rounds of alternating solution operations are performed on the first sub-objective function and the second sub-objective function to obtain the maximized ratio and the optimized deployment data.
[0122] In some embodiments, the solving module 30 is further configured such that the ratio includes a first target ratio and a second target ratio, and the optimized deployment data includes first target deployment data corresponding to the first target ratio and second target deployment data corresponding to the second target ratio; each round of alternating solution operation is performed as follows: based on the constraints, the first sub-objective function is iteratively solved multiple times to obtain the first ratio and first deployment data corresponding to the current round of alternating solution operation, and based on the constraints, the second sub-objective function is solved multiple times individually to obtain the second ratio and second deployment data corresponding to the current round of alternating solution operation; in response to determining that the number of alternating solutions executed is less than the predetermined number of alternating solutions, the next round of alternating solution operation is executed; in response to determining that the number of alternating solutions executed is greater than or equal to the predetermined number of alternating solutions, the first ratio corresponding to the current round of alternating solution operation is determined as the first target ratio, the first deployment data is determined as the first target deployment data, the second ratio is determined as the second target ratio, the second deployment data is determined as the second target deployment data, and the multi-round alternating solution operation is exited.
[0123] In some embodiments, the solution module 30 is further configured to perform multiple iterations of the first sub-objective function based on the constraints using a deep deterministic policy gradient algorithm.
[0124] In some embodiments, the solving module 30 is further configured to perform the following for each round of individual solving operations: based on the constraints, solve the second sub-objective function using a convex optimization algorithm to obtain the second ratio corresponding to the current round of individual solving operations; in response to determining that the difference between the second ratio corresponding to the current round of individual solving operations and the second ratio corresponding to the previous round of individual solving operations is greater than a predetermined rate difference, execute the next round of individual solving operations; in response to determining that the difference between the second ratio corresponding to the previous round of individual solving operations is less than or equal to the predetermined rate difference, determine the second ratio corresponding to the current round of individual solving operations as the second ratio corresponding to the current round of alternating solving operations, determine the second deployment data corresponding to the current round of individual solving operations as the second deployment data corresponding to the current round of alternating solving operations, and exit the multi-round individual solving operations.
[0125] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0126] The apparatus of the above embodiments is used to implement the corresponding reflective surface enhanced UAV-assisted deployment data optimization method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0127] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the UAV-assisted deployment data optimization method with enhanced reflective surface as described in any of the above embodiments.
[0128] Figure 7 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0129] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0130] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0131] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0132] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0133] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0134] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0135] The electronic devices described above are used to implement the corresponding reflective surface-enhanced UAV-assisted deployment data optimization method in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0136] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the UAV-assisted deployment data optimization method with enhanced reflectivity as described in any of the above embodiments.
[0137] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0138] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the UAV-assisted deployment data optimization method with enhanced reflectivity as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0139] Based on the same concept, corresponding to the methods of any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to execute the UAV-assisted deployment data optimization method with enhanced reflective surface as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0140] It should be noted that the embodiments of this application can also be further described in the following ways:
[0141] It is understood that before using the technical solutions of the various embodiments in this disclosure, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.
[0142] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations of this disclosed technical solution.
[0143] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0144] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0145] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application is limited to these examples; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in detail for the sake of brevity.
[0146] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0147] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0148] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A method for optimizing deployment data for unmanned aerial vehicles (UAVs) with enhanced reflective surfaces, characterized in that, include: Obtain deployment data for drones within a predetermined area; Based on the deployment data, with the goal of maximizing the ratio of total secure communication rate to total communication energy consumption, an objective function for optimizing the deployment data is constructed, as well as the corresponding constraints for the objective function; Based on the constraints, the objective function is solved to obtain the maximized ratio and the optimized deployment data; The objective function includes a first sub-objective function and a second sub-objective function. The deployment data includes at least the beam matrix at the base station in the predetermined area, the phase matrix at the reflector, the position of the UAV, the value of artificial noise, the signal-to-noise ratio for each user, the signal-to-noise ratio for each eavesdropper, the weight for each user, the precoding vector for each user, the predetermined phase matrix, and the total number of users. Based on the deployment data, and with the objective of maximizing the ratio of total secure communication rate to total communication energy consumption, an objective function is constructed to optimize the deployment data, including: The first sub-objective function is determined by the following formula: , in, This represents the maximum ratio of total secure communication rate to total communication power consumption. Indicates the total secure communication rate. Indicates total communication energy consumption. For the first The weight of each user For the first Signal-to-noise ratio per user For the first The signal-to-noise ratio (SNR) of each user for the eavesdropper is determined by the beam matrix at the base station, the phase matrix at the reflector, the position of the drone, and the value of the artificial noise. The beam matrix at the base station, Let be the phase matrix at the reflecting surface. The location of the drone. The artificial noise, The total number of users, The total number of base stations For the first The base station to the Pre-defined beam matrix for each user Artificial noise emitted by drones; The second sub-objective function is determined by the following formula: , in, For the first The weight of each user The total number of users, This represents a pre-constructed second formula. This represents a pre-constructed third formula. This represents the first formula that has been pre-constructed. This represents the pre-constructed fourth formula. The result obtained in the t-th training round The value, This represents the predetermined phase matrix. For the first Precoded vectors for each user For training rounds; The first formula pre-constructed is: ; The pre-constructed second formula is: The pre-constructed third formula is: The pre-constructed fourth formula is: ; in, This indicates taking the trace of the matrix. , , These are the diagonal elements of the RIS phase matrix. , Indicates the non-line-of-sight path channel components. , This refers to the path loss when the reference distance is 1m. Indicates the distance between the signal transmitter and the receiver. It is represented as the path loss index. Indicates the wavelength of the transmitted signal. Represents Rice factor, Let M be the distance between the base station and the Mth element at the smart reflective surface. for The conjugate transpose of . , Let the channel gain be from the r-th RIS to the k-th user. Let diag(x) represent the channel gain from the b-th base station to the r-th RIS. A diagonal matrix of N, Let b be the channel gain from the b-th base station to the k-th user. for The conjugate transpose of . For the pre-constructed formula related to the eavesdropper's channel gain, For the first Beam matrix at the base station corresponding to each user For the first The variance of thermal noise at each user location The variance of thermal noise for the eavesdropper.
2. The method according to claim 1, characterized in that, The constraints include: a security rate threshold constraint, a transmission rate constraint, and a unit modulus constraint. The deployment data includes at least the beam matrix at the base station in the predetermined area, the phase matrix at the reflector, the position of the UAV, the value of artificial noise, the signal-to-noise ratio for each user, the signal-to-noise ratio for each eavesdropper, the weight for each user, the precoding vector for each user, the security rate threshold, the maximum transmission power, the total number of base stations in the predetermined area, the unit modulus of the reflector, and the total number of reflectors. Based on the deployment data, the constraints corresponding to the objective function are constructed, including: The safe rate threshold constraint is determined using the following formula: , in, For the first The weight of each user For the first The signal-to-noise ratio (SNR) for each user is determined by the beam matrix at the base station, the phase matrix at the reflector, the position of the UAV, and the value of the artificial noise. For the first The signal-to-noise ratio of each user to the eavesdropper. For the total number of users, The security rate threshold; The transmission rate constraint is determined using the following formula: , in, For the first The base station and the first Precoded vectors for user communication The maximum transmission power threshold, The total number of base stations within the predetermined area; The unit modulus constraint condition is determined by the following formula: , in, The unit modulus of the reflecting surface. Indicates the number of rows in the matrix. This represents the total number of the reflecting surfaces.
3. The method according to claim 1, characterized in that, The objective function includes a first sub-objective function and a second sub-objective function; The step of solving the objective function based on the constraints to obtain the maximized ratio and optimized deployment data includes: Based on the constraints, multiple rounds of alternating solution operations are performed on the first sub-objective function and the second sub-objective function to obtain the maximized ratio and the optimized deployment data.
4. The method according to claim 3, characterized in that, The ratio includes a first target ratio and a second target ratio, and the optimized deployment data includes the first target deployment data corresponding to the first target ratio and the second target deployment data corresponding to the second target ratio; The step of performing multiple rounds of alternating solution operations on the first sub-objective function and the second sub-objective function based on the constraints to obtain the maximized ratio and the optimized deployment data includes: Each round of alternating solution operations is performed as follows: Based on the constraints, the first sub-objective function is solved iteratively multiple times to obtain the first ratio and the first deployment data corresponding to the current round of alternating solution operation. Based on the constraints, the second sub-objective function is solved individually multiple times to obtain the second ratio and the second deployment data corresponding to the current round of alternating solution operation. In response to the determination that the number of alternating solutions performed is less than the predetermined number of alternating solutions, the next round of alternating solution operations is executed; In response to determining that the number of alternating solutions performed is greater than or equal to the predetermined number of alternating solutions, the first ratio corresponding to the current round of alternating solution operation is determined as the first target ratio, the first deployment data is determined as the first target deployment data, the second ratio is determined as the second target ratio, the second deployment data is determined as the second target deployment data, and the multi-round alternating solution operation is exited.
5. The method according to claim 4, characterized in that, The step of iteratively solving the first sub-objective function based on the constraints includes: Based on the aforementioned constraints, the first sub-objective function is solved iteratively multiple times using a deep deterministic policy gradient algorithm.
6. The method according to claim 4, characterized in that, The step of performing multiple rounds of individual solution operations on the second sub-objective function based on the constraints to obtain the second ratio and second deployment data corresponding to the current round of alternating solution operations includes: Each round of individual solution operations is performed as follows: Based on the constraints, the second sub-objective function is solved using a convex optimization algorithm to obtain the second ratio corresponding to the individual solution operation in the current round; In response to determining that the difference between the second ratio corresponding to the current round of individual solution operation and the second ratio corresponding to the previous round of individual solution operation is greater than a predetermined rate difference, the next round of individual solution operation is executed. In response to determining that the difference between the second ratio and the second ratio corresponding to the previous round of individual solution operation is less than or equal to the predetermined rate difference, the second ratio corresponding to the current round of individual solution operation is determined as the second ratio corresponding to the current round of alternating solution operation, the second deployment data corresponding to the current round of individual solution operation is determined as the second deployment data corresponding to the current round of alternating solution operation, and the multi-round individual solution operation is exited.
7. A drone-assisted deployment data optimization device with enhanced reflective surface, characterized in that, include: The acquisition module is configured to acquire deployment data of drones within a predetermined area; The construction module is configured to construct an objective function to optimize the deployment data based on the deployment data, with the goal of maximizing the ratio of total secure communication rate to total communication energy consumption, and to construct the constraints corresponding to the objective function. The solution module is configured to solve the objective function based on the constraints to obtain the maximized ratio and optimized deployment data; The objective function includes a first sub-objective function and a second sub-objective function. The deployment data includes at least the beam matrix at the base station in the predetermined area, the phase matrix at the reflector, the position of the UAV, the value of artificial noise, the signal-to-noise ratio for each user, the signal-to-noise ratio for each eavesdropper, the weight for each user, the precoding vector for each user, the predetermined phase matrix, and the total number of users. The building module is also configured as follows: The first sub-objective function is determined by the following formula: , in, This represents the maximum ratio of total secure communication rate to total communication power consumption. Indicates the total secure communication rate. Indicates total communication energy consumption. For the first The weight of each user For the first Signal-to-noise ratio per user For the first The signal-to-noise ratio (SNR) of each user for the eavesdropper is determined by the beam matrix at the base station, the phase matrix at the reflector, the position of the drone, and the value of the artificial noise. The beam matrix at the base station, Let be the phase matrix at the reflecting surface. The location of the drone. The artificial noise, The total number of users, The total number of base stations For the first The base station to the Pre-defined beam matrix for each user Artificial noise emitted by drones; The second sub-objective function is determined by the following formula: , in, For the first The weight of each user The total number of users, This represents a pre-constructed second formula. This represents a pre-constructed third formula. This represents the first formula that has been pre-constructed. This represents the pre-constructed fourth formula. The result obtained in the t-th training round The value, This represents the predetermined phase matrix. For the first Precoded vectors for each user For training rounds; The first formula pre-constructed is: ; The pre-constructed second formula is: The pre-constructed third formula is: The pre-constructed fourth formula is: ; in, This indicates taking the trace of the matrix. , , These are the diagonal elements of the RIS phase matrix. , Indicates the non-line-of-sight path channel components. , This refers to the path loss when the reference distance is 1m. Indicates the distance between the signal transmitter and the receiver. It is represented as the path loss index. Indicates the wavelength of the transmitted signal. Represents Rice factor, Let M be the distance between the base station and the Mth element at the smart reflective surface. for The conjugate transpose of . , Let the channel gain be from the r-th RIS to the k-th user. Let diag(x) represent the channel gain from the b-th base station to the r-th RIS. A diagonal matrix of N, Let b be the channel gain from the b-th base station to the k-th user. for The conjugate transpose of . For the pre-constructed formula related to the eavesdropper's channel gain, For the first Beam matrix at the base station corresponding to each user For the first The variance of thermal noise at each user location The variance of thermal noise for the eavesdropper.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method described in any one of claims 1 to 6.