Air-ground cooperative efficient access and low latency short packet security communication method and system
Patent Information
- Application Number
- CN202610958241.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2046-06-30
AI Technical Summary
[0004]然而,将上述现有技术应用于存在信息泄露问题的可重构智能表面辅助多设备短包通信场景时,仍存在以下缺陷:
[0044](1)本发明将有限块长度理论引入可重构智能表面辅助的RSMA短包安全通信场景,建立了包含信道色散项和非零译码错误概率的短包安全速率模型,能够准确表征短包通信中可达速率与块错误概率、信干噪比之间的非线性耦合关系。该模型克服了传统香农理论无法精确刻画短包通信性能的局限,为联合优化提供了精确的理论基础。本发明通过将RSMA与有限块长度理论融合,使系统能够在满足可靠性约束的同时最大化安全传输速率,显著提升了短包通信的安全可靠性。
Smart Images

Figure CN122476342B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, specifically to a method and system for efficient air-ground cooperative access and low-latency short-packet secure communication. Background Technology
[0002] Short packet communication, as a key technology for realizing ultra-reliable and low-latency communications (URLLC), faces inherent performance bottlenecks under the finite block length mechanism. At the same time, the massive number of devices accessing the network exacerbates the scarcity of spectrum resources, and the openness of wireless channels makes their signal transmission vulnerable to unauthorized interception.
[0003] To address the aforementioned issues, existing technologies primarily improve upon these technologies at the multiple access level and the wireless environment reconfiguration level. At the multiple access level, Rate-Splitting Multiple Access (RSMA) is employed to segment user information into public and private flows, combined with serial interference cancellation. This approach offers more flexible and effective management of inter-user interference compared to traditional Space-Division Multiple Access (SDMA) and Non-Orthogonal Multiple Access (NOMA). At the wireless environment reconfiguration level, Reconfigurable Intelligent Surfaces (RIS), through their numerous passive reflective elements, can intelligently modulate the electromagnetic wave phase with low power consumption to reconfigure the wireless propagation environment. This dynamically establishes reliable virtual line-of-sight links for devices in communication dead zones, effectively overcoming signal obstruction and attenuation problems.
[0004] However, when applying the aforementioned existing technologies to reconfigurable smart surface-assisted multi-device short packet communication scenarios where information leakage issues exist, the following drawbacks still exist:
[0005] First, most existing solutions for reconfigurable smart surface communication are based on Shannon's theory, which assumes infinite block length and negligible decoding error probability. This theory cannot accurately characterize the performance characteristics of short packet communication under finite block length mechanisms, such as the coupled impact of non-zero decoding error probability on system security and reliability. Therefore, when optimization methods based on Shannon's theory are directly applied to short packet communication scenarios, the security performance of short packet communication cannot be fully guaranteed.
[0006] Secondly, in multi-user scenarios, the risks of interference between users and information leakage are intertwined, making resource allocation and interference management exceptionally complex. Traditional RSMA resource allocation methods struggle to effectively suppress unauthorized users from intercepting information while satisfying the reliability constraints of all legitimate users.
[0007] Therefore, how to balance the risks of finite block length mechanism, multiple access interference and information leakage in a multi-user short packet communication system assisted by reconfigurable smart surfaces, and jointly optimize multi-dimensional resources such as the reflectivity of reconfigurable smart surfaces to improve communication security performance is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0008] To address the shortcomings of existing technologies, this invention provides a method and system for efficient air-to-ground collaborative access and low-latency short-packet secure communication. It constructs a short-packet secure communication system model consisting of a multi-antenna transmitter, an airborne platform equipped with a reconfigurable smart surface, multiple authorized receiving devices, and an unauthorized receiving device. The system employs RSMA and serial interference cancellation, and jointly optimizes the reflection coefficient of the reconfigurable smart surface and the flight position of the airborne platform. This effectively manages interference between users in a multi-user environment, balancing the high reliability and secure transmission requirements of short-packet communication.
[0009] To achieve the above objectives, the present invention adopts the following technical solution:
[0010] This invention provides a method for efficient air-to-ground collaborative access and low-latency short-packet secure communication, comprising the following steps:
[0011] S1. Construct a short packet secure communication system model, which includes a multi-antenna transmitter, an air platform, and a receiver. The receiver includes... One authorized receiving device and one unauthorized receiving device; the air platform is equipped with a reconfigurable smart surface; the short packet secure communication system model adopts RSMA; channel modeling is performed on each communication link to determine the signal-to-interference-plus-noise ratio expression of the receiving end;
[0012] S2. Based on the short packet secure communication system model, a joint optimization problem is constructed with the goal of maximizing the total achievable secure rate of the system; wherein, the objective function of the joint optimization problem couples the position coordinates of the airborne platform with the reflection coefficient of the reconfigurable smart surface;
[0013] S3. The joint optimization problem is solved iteratively until the convergence condition is met. In each iteration, the current position coordinates of the air platform are fixed first, and the reflection coefficient of the reconfigurable smart surface is optimized using a successive convex approximation algorithm. Then, the current reflection coefficient is fixed, the optimization problem of the air platform position is modeled as a Markov decision process, and the position coordinates of the air platform are optimized using the DSAC-T (Distributive Soft Actor-Critic with Three Refinements) deep reinforcement learning algorithm.
[0014] S4. Output the final optimized reconfigurable smart surface reflection coefficient and the aerial platform position coordinates to achieve secure transmission of short packet communication.
[0015] Furthermore, in S1, channel modeling for each communication link includes: modeling the link between the multi-antenna transmitter and the legitimate receiving device as a Rayleigh fading model; and modeling the link between the air platform and the multi-antenna transmitter, as well as the link between the air platform and the receiver, as a Rice fading model.
[0016] Furthermore, in S2, the total reachable safe rate of the system is determined based on the finite block length, and its expression is:
[0017]
[0018]
[0019]
[0020] In the formula, For the first The total achievable safe rate of the system across all time slots In the first The first time slot, the first The maximum achievable decoding rate of a legitimate receiving device for a private information stream; In the first In the first time slot, the unauthorized receiving device... The achievable decoding rate of a legitimate receiving device's private information stream; To perform the positive part operation; In the first The first time slot, the first The signal-to-interference-plus-noise ratio of a private information stream from a legitimate receiving device; In the first In the first time slot, the unauthorized receiving device acquires the first... Signal-to-interference-to-noise ratio (SIR) when a legitimate receiving device receives a private information stream; To be assigned to the The number of symbols in the private information stream of a legitimate receiving device. For the channel dispersion function, For Gauss The inverse function of a function For the first Block error probability constraint value for a valid receiving device For the first The information leakage probability constraint value of a legitimate receiving device.
[0021] Furthermore, in S2, the joint optimization problem is expressed as:
[0022]
[0023] Constraints:
[0024]
[0025]
[0026]
[0027]
[0028] In the formula, For the first The total achievable safe rate of the system across all time slots In the first The first time slot, the first The maximum achievable decoding rate of a legitimate receiving device for a private information stream; In the first In the first time slot, the unauthorized receiving device... The achievable decoding rate of a legitimate receiving device's private information stream. The preset safe rate threshold, For the first Phase offset vector of each time slot, For the first in reconfigurable smart surface Phase shift of each reflective element, The total number of reflective elements, In the first Each time slot contains the location coordinates of the aerial platform. , , They are respectively axis, axis, Axis coordinates , , They are respectively , , The range of values for .
[0029] Furthermore, in S3, the reflection coefficient of the reconfigurable smart surface is optimized using a successive convex approximation algorithm, specifically including:
[0030] By introducing auxiliary variables, the fractional constraints in the joint optimization problem are transformed into linear constraints;
[0031] A first-order Taylor expansion is performed on the channel dispersion term introduced by the finite block length to construct its linear upper bound, thus approximating the non-convex objective function and constraints as convex functions.
[0032] The approximated convex function is transformed into a second-order cone programming problem;
[0033] The second-order cone programming problem is solved iteratively until the convergence condition is met, and the optimized reflection coefficient is obtained.
[0034] Furthermore, in S3, the Markov decision process includes a state space, an action space, and a reward function; the state space is defined by the aerial platform in the first... The position coordinates of each time slot are constituted; the action space includes the flight speed, elevation angle and azimuth angle of the air platform; the reward function uses the total reachable safe rate of the system within the allowed flight area as a positive reward and the total reachable safe rate of the system outside the allowed flight area as a negative reward.
[0035] Furthermore, in S3, the DSAC-T deep reinforcement learning algorithm includes the following training mechanism:
[0036] An adaptive clipping boundary is adopted, which dynamically adjusts the clipping range according to the current standard deviation of the value distribution; a gradient scaling weight based on the variance of the value distribution is used to adaptively adjust the gradient update magnitude; and a dual evaluation network structure is adopted, which estimates the target value through two independent evaluation networks and selects the smaller estimate for policy update.
[0037] Furthermore, in S3, the convergence condition is: the change in the total achievable safe rate of the system is less than a preset threshold, or the preset number of iterations is reached.
[0038] This invention also provides an air-to-ground cooperative high-efficiency access and low-latency short packet secure communication system for performing the above method, including:
[0039] A building module is used to construct the short packet secure communication system model.
[0040] The optimization module is used to construct the joint optimization problem based on the short packet secure communication system model and by using RSMA to manage inter-user interference.
[0041] The first solution module is used to decompose the joint optimization problem into a first sub-problem for optimizing the reflectivity of the reconfigurable smart surface, and to use a successive convex approximation algorithm to transform the first sub-problem into a convex problem for solution, so as to obtain the optimized reflectivity of the reconfigurable smart surface.
[0042] The second solution module is used to decompose the joint optimization problem into a second sub-problem for optimizing the position of the air platform, and to model the second sub-problem as a Markov decision process. The DSAC-T deep reinforcement learning algorithm is used to solve the Markov decision process to obtain the optimized position coordinates of the air platform.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] (1) This invention introduces finite block length theory into the reconfigurable smart surface-assisted RSMA short packet secure communication scenario, establishing a short packet secure rate model that includes channel dispersion terms and non-zero decoding error probabilities. This model accurately characterizes the nonlinear coupling relationship between achievable rate, block error probability, and signal-to-interference-plus-noise ratio in short packet communication. This model overcomes the limitation of traditional Shannon theory in accurately characterizing the performance of short packet communication, providing a precise theoretical basis for joint optimization. By integrating RSMA with finite block length theory, this invention enables the system to maximize secure transmission rate while satisfying reliability constraints, significantly improving the security and reliability of short packet communication.
[0045] (2) To address the complex non-convex characteristics of the objective function under the finite block length mechanism, this invention adopts a successive convex approximation solution strategy: introducing auxiliary variables to transform fractional constraints into linear constraints, constructing a linear upper bound by performing a first-order Taylor expansion on the channel dispersion term, and transforming the original non-convex subproblem into a solvable second-order cone programming problem, which has the advantages of fast convergence speed and high solution accuracy. At the same time, for the joint optimization problem of the strongly coupled intelligent surface reflection coefficient and the position of the airborne platform, this invention proposes an alternating iterative solution framework that combines successive convex approximation and deep reinforcement learning. By decomposing the original problem into two subproblems and optimizing them alternately, the solutions of the two subproblems reinforce each other and gradually approach the optimum, effectively solving the problem of collaborative optimization of strongly coupled non-convex problems.
[0046] (3) This invention employs the DSAC-T deep reinforcement learning algorithm in the aerial platform position optimization stage. This algorithm includes three training stabilization mechanisms: adaptive clipping boundary, which dynamically adjusts the clipping range based on the standard deviation of the value distribution to reduce gradient oscillations; gradient scaling weights based on the variance of the value distribution, which adaptively adjusts the gradient update magnitude to improve robustness to changes in reward scale; and a dual-value distribution learning structure, which estimates the target value through two independent evaluation networks and selects the smaller value for policy updates, effectively mitigating the overestimation problem of the action value function. The synergistic effect of these three training stabilization mechanisms gives the aerial platform position optimization higher training stability, sample efficiency, and convergence performance.
[0047] (4) This invention deploys reconfigurable smart surfaces on an aerial platform, creating a mobile aerial smart reflective platform. This design utilizes the high mobility of the aerial platform to establish optimal line-of-sight links. Simultaneously, it leverages the passive beamforming capability of the reconfigurable smart surface to reconfigure wireless channels with low energy consumption, dynamically establishing reliable coverage for communication blind spots and effectively overcoming signal blocking and attenuation problems. Compared with traditional fixed reconfigurable smart surface deployments, this invention significantly enhances the coverage, deployment flexibility, and robustness of the communication network, providing wide coverage and high reliability technical guarantees for short packet communication. Attached Figure Description
[0048] Figure 1 This is a flowchart of the air-ground cooperative high-efficiency access and low-latency short packet secure communication method provided in Embodiment 1 of the present invention;
[0049] Figure 2 This is a comparison chart of the total achievable safe rate of the system under different transmission powers for Embodiment 1, Comparative Example 1, and Comparative Example 2 of the present invention;
[0050] Figure 3 This is a comparison chart of the total achievable safe rate of the system under different numbers of reflective elements in Embodiment 1, Comparative Example 1, and Comparative Example 2 of the present invention;
[0051] Figure 4 This is a comparison chart showing the change in average reward during the training process of Embodiment 1, Comparative Example 3, and Comparative Example 4 of the present invention as a function of the number of training rounds. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] Example 1
[0054] refer to Figure 1 This embodiment provides a method for efficient air-to-ground collaborative access and low-latency short packet secure communication, which is performed according to the following steps:
[0055] S1. Construct a short packet secure communication system model, which includes a multi-antenna transmitter, an air platform, and a receiver. The receiver includes... One legitimate receiving device and one unauthorized receiving device;
[0056] The aerial platform is equipped with a reconfigurable smart surface; the reconfigurable smart surface is composed of... It consists of 10 or more reflective elements, preferably ≤ 10 ≤25, each reflective element can independently modulate the phase of the incident signal. The phase response of the reconfigurable smart surface is represented by a phase shift vector:
[0057]
[0058] in For the first Phase offset vector of each time slot, For the first in reconfigurable smart surface Phase shift of each reflective element, .
[0059] The short packet secure communication system model employs RSMA, dividing the information stream of each legitimate receiving device into a common information stream and a private information stream. The common information streams of all legitimate receiving devices are merged and encoded into a global common information stream, while the private information streams of each legitimate receiving device are encoded independently, ensuring that each private information stream can only be decoded by its corresponding legitimate receiving device. A linear precoding matrix is used to beamform the global common information stream and each private information stream to form the transmitted signal. :
[0060]
[0061] in, The beamforming precoding vector represents the global common information flow. Represents the global common information flow. Indicates the first Beamforming precoding vectors of the private information stream of a legitimate receiving device Indicates the first A private information stream of a legitimate receiving device.
[0062] Channel modeling is performed on each communication link in the short packet secure communication system model to determine the signal-to-interference-plus-noise ratio (SINR) expression at the receiver. The specific modeling is as follows:
[0063] The non-line-of-sight link between the multi-antenna transmitter and the legitimate receiver is modeled as a Rayleigh fading model.
[0064] The links between the air platform and the multi-antenna transmitter, the links between the air platform and the authorized receiving equipment, and the links between the air platform and the unauthorized receiving equipment are all modeled as Ricean fading models.
[0065] The signal-to-interference-plus-noise ratio (SINR) expressions for each receiver are as follows:
[0066] In the The first time slot, the first Signal-to-interference-plus-noise ratio of public information streams from a legitimate receiving device for:
[0067]
[0068] In the formula, For the first Each time slot, from the multi-antenna transmitter, is reflected via the air platform to the first... The equivalent channel vector of a legitimate receiving device The variance is the additive white Gaussian noise.
[0069] In the The first time slot, the first Signal-to-interference-plus-noise ratio of a private information stream from a legitimate receiving device for:
[0070]
[0071] In the formula, Beamforming precoding vectors for private information streams from other legitimate receiving devices.
[0072] In the In the first time slot, the unauthorized receiving device acquires the first... Signal-to-interference-plus-noise ratio of private information streams received by a legitimate receiving device for:
[0073]
[0074] In the formula, For the first Each time slot is the equivalent channel vector reflected from the multi-antenna transmitter via the air platform to the unlicensed receiving device.
[0075] In this embodiment, the multi-antenna transmitter is a central controller (CC) with 8 antennas, the authorized receiving devices and the unauthorized receiving devices are each equipped with a single antenna, the aerial platform is a drone, and the number of authorized receiving devices is 3.
[0076] S2. Based on the short packet secure communication system model, the total reachable security rate of the system is determined by using a finite block length, and a joint optimization problem is constructed with the goal of maximizing the total reachable security rate of the system.
[0077] In the The first time slot, the first The maximum achievable decoding rate of a legitimate receiving device for a private information stream for:
[0078]
[0079] In the formula, To be assigned to the The number of symbols in the private information stream of a legitimate receiving device. For Gauss The inverse function of a function For the first In this embodiment, the block error probability constraint value for each legitimate receiving device is... , For the channel dispersion function, .
[0080] In the In the first time slot, the unauthorized receiving device... The achievable decoding rate of a private information stream for a legitimate receiving device for:
[0081]
[0082] In the formula, , For the first In this embodiment, the probability constraint value for information leakage of a legitimate receiving device is defined as follows: .
[0083] No. The total achievable safe rate of the system in each time slot for:
[0084]
[0085] In the formula, To perform the positive part operation, i.e. When the time comes, take its original value; When the time is right, take 0.
[0086] As can be seen from the signal-to-interference-plus-noise ratio (SIRR) formula, the allocation strategy of public and private information flows in RSMA directly affects the SIRR of each link. In the finite block length rate model, the achievable rate is determined by both the principal term and the penalty term: the principal term monotonically increases with increasing SIRR; in the penalty term, channel dispersion is a nonlinear function of SIRR. When SIRR is small, the channel dispersion value is close to 0, and the absolute value of the penalty term is small; when SIRR increases, the channel dispersion value approaches 1, and the absolute value of the penalty term increases. This means that increasing SIRR increases the principal term, but also increases the magnitude of the penalty term, with the principal and penalty terms showing opposite trends.
[0087] RSMA's traffic splitting strategy manages multi-user interference by adjusting the signal-to-interference-plus-noise ratio (SIR) distribution across links. Therefore, RSMA's changes to the SIR simultaneously affect both the main term and the penalty term, and the directions of influence may be opposite.
[0088] The principal term is a concave function whose growth rate decreases with increasing signal-to-interference-plus-noise ratio (SNR), while the penalty term is also a concave function of SNR, exhibiting drastic changes in the low SNR region and tending to saturate in the high SNR region. The combined effect of these two terms results in a complex non-convex objective function that cannot be directly solved using conventional convex optimization methods. Therefore, this invention constructs a joint optimization problem and subsequently introduces alternating optimization.
[0089] The joint optimization problem is expressed as:
[0090]
[0091] Constraints:
[0092]
[0093]
[0094]
[0095]
[0096] In the formula, In this embodiment, a preset safe rate threshold is used. =0.5bps / Hz, In the first Each time slot contains the location coordinates of the aerial platform. , , They are respectively axis, axis, Axis coordinates , , They are respectively , , In this embodiment, a three-dimensional coordinate system is constructed with the position of the multi-antenna transmitter as the origin.
[0097] The objective function of the joint optimization problem couples the position coordinates of the airborne platform with the reflection coefficient of the reconfigurable smart surface.
[0098] S3. Solve the joint optimization problem iteratively until the convergence condition is met. The convergence condition is: the change in the total achievable safe rate of the system is less than a preset threshold, or the preset number of iterations is reached.
[0099] In each iteration, the following sub-steps are performed:
[0100] S301. Given the current aerial platform position coordinates, decompose the joint optimization problem into a first subproblem for optimizing the reconfigurable smart surface phase offset vector:
[0101]
[0102] Constraints:
[0103]
[0104]
[0105]
[0106] Because the objective function of the first subproblem contains a logarithmic function and a nonlinear dispersion term related to the signal-to-interference-plus-noise ratio (SINR), it exhibits non-convexity. This embodiment utilizes a successive convex approximation algorithm to transform it into a solvable convex problem, optimizing the reflection coefficient of the reconfigurable smart surface. The specific processing is as follows:
[0107] Introducing auxiliary variables and The fractional constraints in the joint optimization problem are transformed into linear constraints. In the first The first time slot, the first The lower bound of the signal-to-interference-plus-noise ratio (SIR) for a legitimate receiving device's private information stream. In the first In the first time slot, the unauthorized receiving device acquires the first... Upper bound of signal-to-interference-plus-noise ratio for private information streams from a legitimate receiving device;
[0108] The channel dispersion term introduced by the finite block length is expanded by a first-order Taylor expansion at the expansion point of this iteration to construct its linear upper bound, thus approximating the non-convex objective function and constraints as convex functions.
[0109] The approximated convex function is transformed into a second-order cone programming problem:
[0110] The second-order cone programming problem is solved iteratively until the reflection coefficient between two adjacent iterations is less than the preset convergence tolerance. The optimized reflection coefficient is obtained when the number of iterations reaches a preset value. In this embodiment, .
[0111] S302. Given the reflection coefficients obtained from the current optimization, decompose the joint optimization problem into a second sub-problem for optimizing the location of the airborne platform:
[0112]
[0113] Constraints:
[0114]
[0115]
[0116] Due to the location coordinates of the aerial platform The second subproblem appears in the channel model and is time-coupled, meaning the position of the current time slot affects the position of subsequent time slots. Therefore, it cannot be directly solved using conventional convex optimization methods. In this embodiment, the second subproblem is modeled as a Markov decision process and solved using the DSAC-T deep reinforcement learning algorithm to obtain the optimized airborne platform position coordinates.
[0117] The Markov decision process includes: a state space, an action space, and a reward function. The state space is defined by the aerial platform in the [missing information - likely a specific step or event]. Position coordinates of each time slot Composition; the action space includes the aerial platform in the first... Flight speed per time slot Angle of elevation and azimuth ,in The maximum flight speed of the airborne platform; the reward function uses the total reachable safe rate of the system within the allowed flight area as a positive reward, and the total reachable safe rate of the system exceeding the allowed flight area as a negative reward. The positive reward is... Negative reward ,in As a penalty factor, in this embodiment .
[0118] The DSAC-T deep reinforcement learning algorithm introduces three training stabilization mechanisms based on the traditional SAC algorithm:
[0119] An adaptive clipping boundary is adopted, which dynamically adjusts the clipping range according to the current standard deviation of the value distribution to reduce gradient oscillations and improve training stability;
[0120] Gradient scaling weights based on the variance of value distribution are used to adaptively adjust the gradient update magnitude, reduce the impact of reward scale changes on network training, and improve the robustness of the algorithm to different reward scales.
[0121] A dual-evaluation network structure is adopted, in which two independent evaluation networks jointly estimate the target value and select the smaller estimate for policy update, which effectively alleviates the overestimation problem of action value function and improves training stability.
[0122] Repeat steps S301 and S302 until the convergence condition is met. In this embodiment, the convergence condition is:
[0123]
[0124] For the first The iteration yielded the... The system can achieve a total safe rate across all time slots.
[0125] Example 2
[0126] This embodiment provides an efficient air-to-ground cooperative access and low-latency short-packet secure communication system for executing the method provided in Embodiment 1. The system includes:
[0127] A building module is used to construct a short-packet secure communication system model, which includes a multi-antenna transmitter, an airborne platform equipped with a reconfigurable smart surface, and... One legitimate receiving device and one unauthorized receiving device.
[0128] An optimization module is used to construct a joint optimization problem based on the short packet secure communication system model, using RSMA to manage inter-user interference, with the goal of maximizing the total achievable secure rate.
[0129] The first solution module is used to decompose the joint optimization problem into a first subproblem for optimizing the reflectivity of the reconfigurable smart surface, and to use a successive convex approximation algorithm to transform the first subproblem into a convex problem for solution, so as to obtain the optimized reflectivity of the reconfigurable smart surface.
[0130] The second solution module is used to decompose the joint optimization problem into a second sub-problem for optimizing the position of the air platform, and to model the second sub-problem as a Markov decision process, and to solve it using the DSAC-T deep reinforcement learning algorithm to obtain the optimized position coordinates of the air platform.
[0131] In this embodiment, the modules work together and through alternating iterative optimization, output the final optimized reconfigurable smart surface reflection coefficient and the position coordinates of the aerial platform.
[0132] Comparative Example 1
[0133] The difference between this comparative example and Example 1 is that RSMA in Example 1 is replaced with NOMA, while the rest is the same as in Example 1.
[0134] Comparative Example 2
[0135] The difference between this comparative example and Example 1 is that RSMA in Example 1 is replaced with SDMA in this comparative example, while the rest is the same as in Example 1.
[0136] Comparative Example 3
[0137] The difference between this comparative example and Example 1 is that the DSAC-T deep reinforcement learning algorithm in Example 1 is replaced with TD3 (Twin Delayed Deep Deterministic Policy Gradient), while the rest is the same as in Example 1.
[0138] Comparative Example 4
[0139] The difference between this comparative example and Example 1 is that this comparative example replaces the DSAC-T deep reinforcement learning algorithm in Example 1 with DDPG (Deep Deterministic Policy Gradient), while the rest is the same as Example 1.
[0140] To better illustrate the beneficial effects of the present invention, simulation verification was performed on Examples 1, 1, 2, 3, and 4 under the same conditions. In the simulation experiment, a spatial rectangular coordinate system was established, with the position of the multi-antenna transmitter as the origin (0,0,0). There were three legal receiving devices, located at (80,70,0) m, (100,80,0) m, and (90,50,0) m, respectively. The unlicensed receiving device was located at (120,60,0) m. The path loss factor was 2.3. The flight area of the airborne platform was 200m × 200m × 200m, and the maximum flight speed of the airborne platform was 10m / s. The channel noise power was -70dBm. The Rice factor was 6.
[0141] Figure 2 The graphs show a comparison of the overall achievable safe rate of the system as a function of the transmit power of the multi-antenna transmitter for different schemes. Figure 2As can be seen, Example 1 achieves a higher overall achievable security rate across all power levels, significantly outperforming Comparative Example 1 and Comparative Example 2. This demonstrates that by employing the RSMA multiple access scheme, the present invention can more flexibly manage inter-user interference, converting some interference signals into useful information, thereby achieving superior secure transmission performance under different transmit power conditions.
[0142] Figure 3 The diagram shows a comparison of the overall achievable safe rate of the system as a function of the number of reconfigurable smart surface reflective elements for different schemes. Figure 3 It can be seen that the overall achievable safe rate of the system in Example 1 is significantly higher than that in Comparative Examples 1 and 2. This indicates that there is a significant synergistic gain effect between RSMA and the reconfigurable smart surface in this invention, and the flexible interference management capability of RSMA can fully utilize the channel control freedom provided by the reconfigurable smart surface.
[0143] Figure 4 The convergence curves of different schemes are shown, comparing the changes in average reward over training epochs. From Figure 4 It can be seen that the DSAC-T algorithm used in Example 1 can improve the average reward value more quickly and reach a stable convergence state after fewer training rounds. In contrast, the TD3 algorithm used in Comparative Example 3 and the DDPG algorithm used in Comparative Example 4 have slower convergence speeds and their final stable average reward values are lower than those of the DSAC-T algorithm. This shows that the present invention can enhance the adaptability to random fading of wireless channels, interference from unlicensed receiving devices, and fluctuations in short packet communication rewards through value distribution learning, and reduce the policy bias caused by value overestimation through a dual value estimation mechanism, thereby improving the stability of the training process and the reliability of the optimization results, and thus improving the total achievable safe rate and overall secure transmission performance of the system.
[0144] The specific embodiments of the present invention are provided to enable those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention.
[0145] It should be understood that the present invention is not limited to the content already described above, and various modifications and changes can be made without departing from its scope. The scope of the present invention is limited only by the appended claims.
Claims
1. A method for efficient air-to-ground collaborative access and low-latency short-packet secure communication, characterized in that, Includes the following steps: S1. Construct a short packet secure communication system model, which includes a multi-antenna transmitter, an air platform, and a receiver. The receiver includes... One legitimate receiving device and one unauthorized receiving device; The aerial platform is equipped with a reconfigurable smart surface; the short packet secure communication system model adopts RSMA; channel modeling is performed on each communication link to determine the signal-to-interference-plus-noise ratio expression at the receiver; S2. Based on the short packet secure communication system model, a joint optimization problem is constructed with the goal of maximizing the total achievable secure rate of the system; wherein, the objective function of the joint optimization problem couples the position coordinates of the airborne platform with the reflection coefficient of the reconfigurable smart surface; The total reachable safety rate of the system is determined based on the finite block length, and its expression is: In the formula, For the first The total achievable safe rate of the system across all time slots In the first The first time slot, the first The maximum achievable decoding rate of a legitimate receiving device for a private information stream; In the first In the first time slot, the unauthorized receiving device... The achievable decoding rate of a legitimate receiving device's private information stream; To perform the positive part operation; In the first The first time slot, the first The signal-to-interference-plus-noise ratio of a private information stream from a legitimate receiving device; In the first In the first time slot, the unauthorized receiving device acquires the first... Signal-to-interference-to-noise ratio (SIR) when a legitimate receiving device receives a private information stream; To be assigned to the The number of symbols in the private information stream of a legitimate receiving device. For the channel dispersion function, For Gauss The inverse function of a function For the first Block error probability constraint value for a valid receiving device For the first The probability constraint value for information leakage of a legitimate receiving device; The joint optimization problem is expressed as: Constraints: In the formula, The preset safe rate threshold, For the first Phase offset vector of each time slot, For the first in reconfigurable smart surface Phase shift of each reflective element The total number of reflective elements, In the first Each time slot contains the location coordinates of the aerial platform. , , They are respectively axis, axis, Axis coordinates , , They are respectively , , The range of values for ; S3. The joint optimization problem is solved iteratively until the convergence condition is met. In each iteration, the position coordinates of the current air platform are fixed first, and the reflection coefficient of the reconfigurable smart surface is optimized using the successive convex approximation algorithm. Then, the current reflection coefficient is fixed, the optimization problem of the air platform position is modeled as a Markov decision process, and the position coordinates of the air platform are optimized using the DSAC-T deep reinforcement learning algorithm. The optimization of the reflection coefficient of the reconfigurable smart surface using a successive convex approximation algorithm specifically includes: By introducing auxiliary variables, the fractional constraints in the joint optimization problem are transformed into linear constraints; A first-order Taylor expansion is performed on the channel dispersion term introduced by the finite block length to construct its linear upper bound, thus approximating the non-convex objective function and constraints as convex functions. The approximated convex function is transformed into a second-order cone programming problem; The second-order cone programming problem is solved iteratively until the convergence condition is met, and the optimized reflection coefficient is obtained. The Markov decision process includes a state space, an action space, and a reward function; the state space is defined by the aerial platform in the [missing information - likely a specific step or event]. The position coordinates of each time slot are constituted; the action space includes the flight speed, elevation angle and azimuth angle of the air platform; the reward function uses the total reachable safe rate of the system within the allowed flight area as a positive reward and the total reachable safe rate of the system beyond the allowed flight area as a negative reward. The DSAC-T deep reinforcement learning algorithm includes the following training mechanism: An adaptive clipping boundary is adopted, which dynamically adjusts the clipping range according to the current standard deviation of the value distribution; a gradient scaling weight based on the variance of the value distribution is used to adaptively adjust the gradient update magnitude; a dual evaluation network structure is adopted, which estimates the target value through two independent evaluation networks and selects the smaller estimate for policy update. S4. Output the final optimized reconfigurable smart surface reflection coefficient and the aerial platform position coordinates to achieve secure transmission of short packet communication.
2. The air-to-ground cooperative high-efficiency access and low-latency short packet secure communication method according to claim 1, characterized in that, In S1, channel modeling for each communication link includes: modeling the link between the multi-antenna transmitter and the legitimate receiving device as a Rayleigh fading model; and modeling the link between the air platform and the multi-antenna transmitter, as well as the link between the air platform and the receiver, as a Ricean fading model.
3. The air-to-ground cooperative high-efficiency access and low-latency short packet secure communication method according to claim 1, characterized in that, In S3, the convergence condition is: the change in the total achievable safe rate of the system is less than a preset threshold, or the preset number of iterations is reached.
4. A high-efficiency air-to-ground cooperative access and low-latency short-packet secure communication system, used to execute the method of any one of claims 1 to 3, characterized in that, include: A building module is used to construct the short packet secure communication system model. The optimization module is used to construct the joint optimization problem based on the short packet secure communication system model and by using RSMA to manage inter-user interference. The first solution module is used to decompose the joint optimization problem into a first sub-problem for optimizing the reflectivity of the reconfigurable smart surface, and to use a successive convex approximation algorithm to transform the first sub-problem into a convex problem for solution, so as to obtain the optimized reflectivity of the reconfigurable smart surface. The second solution module is used to decompose the joint optimization problem into a second sub-problem for optimizing the position of the air platform, and to model the second sub-problem as a Markov decision process. The DSAC-T deep reinforcement learning algorithm is used to solve the Markov decision process to obtain the optimized position coordinates of the air platform.
Citation Information
Patent Citations
Performance evaluation method for short packet coding communication
CN120321695A
AmBC-NOMA communication resource joint optimization method for high mobility V2X scene
CN120416940A