A dense heterogeneous network secure communication resource allocation method, device and equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIDIAN UNIV
- Filing Date
- 2026-04-14
- Publication Date
- 2026-06-23
Smart Images

Figure CN122269478A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication security resource allocation technology, specifically to a method, apparatus, and device for allocating dense heterogeneous network security communication resources. Background Technology
[0002] With the development of 5G and post-5G communication systems, dense heterogeneous networks, through the dense deployment of multiple layers of small base stations under the coverage of macro base stations, have become a key architecture to meet the demands of high data density and high user concurrency. However, the broadcast nature of wireless channels makes these networks highly vulnerable to eavesdropping attacks, posing a serious challenge to communication security. Physical layer security technologies, by leveraging the inherent randomness of channels to enhance security, have become a research focus in this field. In the dynamic and complex environment of real-world dense heterogeneous networks, it is necessary to jointly optimize multi-dimensional resources such as power and spectrum to combat the threat of multiple eavesdroppers while meeting user quality of service and interference management constraints. This places higher demands on the performance and real-time capabilities of resource allocation methods.
[0003] To address the aforementioned issues, existing research has evolved primarily in two directions. On one hand, research focuses on improving physical layer security performance in heterogeneous network scenarios, such as using deep reinforcement learning to jointly optimize beamforming and power allocation to maximize security rate or security efficiency. On the other hand, to address the scalability issues arising from large-scale networking, research has further shifted towards distributed and collaborative optimization mechanisms. For example, distributed deep reinforcement learning is used to achieve autonomous decision-making among multiple nodes, reducing reliance on global information. Furthermore, to protect data privacy and promote collaboration, frameworks such as federated learning and hierarchical federated learning have been introduced, combining with reinforcement learning to form a distributed training paradigm. This aims to achieve efficient model aggregation and cross-node knowledge sharing to optimize the allocation efficiency of resources such as spectrum and power.
[0004] Nevertheless, existing technical solutions still have several limitations. First, most studies on security resource allocation simplify the modeling of eavesdropping environments, failing to fully characterize the complex threats posed by multiple eavesdroppers, thus affecting the effectiveness of strategies in real-world environments. Second, traditional centralized optimization methods are computationally complex, have high signaling overhead, and lack scalability. Third, while distributed reinforcement learning-based methods improve environmental adaptability, each node learns independently based on local observations, lacking efficient experience sharing and global collaboration mechanisms, which can easily lead to policy instability and weak generalization ability. Finally, existing federated learning frameworks mostly focus on the model aggregation process itself, failing to deeply coordinate with physical layer security objectives (such as confidentiality rate) and strict wireless resource constraints, resulting in limited overall optimization performance. Summary of the Invention
[0005] To address the aforementioned problems in the existing technology, the present invention provides a method, apparatus, and device for allocating dense heterogeneous network security communication resources.
[0006] The technical problem to be solved by this invention is achieved through the following technical solution: In a first aspect, the present invention provides a method for allocating dense heterogeneous network security communication resources, including: Based on the current dense heterogeneous network, a dense heterogeneous network system model is constructed. The dense heterogeneous network system model includes: macro base stations, multiple micro base stations, and the set of legitimate users served by each micro base station. Based on the dense heterogeneous network system model and Shannon's capacity formula, the user link communication rate of all legitimate users and the eavesdropping link communication rate of all eavesdroppers in the current dense heterogeneous network are calculated respectively. A confidentiality rate model is constructed based on the user link communication rate and the eavesdropping link communication rate. The confidentiality rate model is an optimization problem with the goal of maximizing the total confidentiality rate of a dense heterogeneous network system model, using the transmission power allocated to the corresponding legitimate users by multiple micro base stations as the optimization variable. The confidentiality rate model is modeled as a multi-agent Markov decision process, and each micro base station is regarded as an agent in the multi-agent Markov decision process. The policy parameters of the agents in the multi-agent Markov decision-making process are updated based on the policy gradient method to obtain the local policy model parameters of each micro base station. Based on the federated learning framework and macro base stations, the parameters of the local policy model are collaboratively optimized until the convergence condition is reached. The transmit power allocation strategy is determined based on the local policy model parameters corresponding to the convergence condition.
[0007] Secondly, the present invention provides a dense heterogeneous network security communication resource allocation device, which includes: a model building unit, a calculation unit, an update unit, a collaborative optimization unit, and an output unit; The model building unit is used to construct a dense heterogeneous network system model based on the current dense heterogeneous network. The dense heterogeneous network system model includes: macro base stations, multiple micro base stations, and the set of legitimate users served by each micro base station. The computing unit is used to calculate, based on the dense heterogeneous network system model and Shannon's capacity formula, the user link communication rate of all legitimate users and the eavesdropping link communication rate of all eavesdroppers in the current dense heterogeneous network. The model building unit is also used to build a secure rate model based on the user link communication rate and the eavesdropping link communication rate. The secure rate model is an optimization problem with the transmit power allocated to the corresponding legitimate users by multiple micro base stations as the optimization variable and the goal of maximizing the total secure rate of the dense heterogeneous network system model. The secure rate model is modeled as a multi-agent Markov decision process, and each micro base station is regarded as an agent in the multi-agent Markov decision process. The update unit is used to update the policy parameters of the agents in the multi-agent Markov decision-making process based on the policy gradient method, so as to obtain the local policy model parameters of each micro base station. The collaborative optimization unit is used to perform collaborative optimization of local policy model parameters based on the federated learning framework and macro base stations until the convergence condition is reached. The output unit is used to determine the transmit power allocation strategy based on the local policy model parameters corresponding to the convergence condition.
[0008] Thirdly, the present invention provides a dense heterogeneous network security communication resource allocation device, comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the dense heterogeneous network security communication resource allocation device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the dense heterogeneous network security communication resource allocation method of the first aspect described above.
[0009] This invention provides a method, apparatus, and device for allocating dense heterogeneous network security communication resources. One method for allocating network communication resources in dense heterogeneous networks includes: constructing a dense heterogeneous network system model based on the current dense heterogeneous network; the dense heterogeneous network system model includes: macro base stations, multiple micro base stations, and a set of legitimate users served by each micro base station; calculating the user link communication rate of all legitimate users and the eavesdropping link communication rate of all eavesdroppers in the current dense heterogeneous network based on the dense heterogeneous network system model and the eavesdropping link communication rate; constructing a security rate model based on the user link communication rate and the eavesdropping link communication rate; the security rate model is an optimization problem with the transmit power allocated by multiple micro base stations to corresponding legitimate users as the optimization variable and the objective of maximizing the total security rate of the dense heterogeneous network system model; modeling the security rate model as a multi-agent Markov decision process, and treating each micro base station as an agent in the multi-agent Markov decision process; updating the policy parameters corresponding to the agents in the multi-agent Markov decision process based on the policy gradient method to obtain the local policy model parameters corresponding to each micro base station; performing collaborative optimization processing on the local policy model parameters based on a federated learning framework and macro base stations until the convergence condition is reached; and determining the transmit power allocation strategy based on the local policy model parameters corresponding to the convergence condition. In this invention, firstly, by constructing a security rate model and aiming to maximize the overall security rate, the physical layer security performance is directly optimized. This addresses the problems of existing research's simplistic modeling of eavesdropping environments and its inability to fully characterize the collaborative threats of multiple eavesdroppers, thereby improving the effectiveness of the strategy in real-world complex environments. Secondly, the security rate model is modeled as a multi-agent Markov decision process, employing a distributed learning framework based on policy gradient methods. This allows each micro-base station to make independent decisions as an agent, solving the problems of computational complexity, high signaling overhead, and insufficient scalability inherent in traditional centralized optimization methods. Next, a federated learning framework is introduced, enabling macro-base stations to collaboratively optimize local policy model parameters. This achieves global experience sharing and a collaborative mechanism, addressing the policy instability and weak generalization capabilities of distributed reinforcement learning methods due to a lack of efficient sharing. Finally, this invention deeply integrates federated learning with the physical layer security objective (security rate) and wireless resource constraints (transmit power), directly integrating security and resource factors during the optimization process. This solves the problems of existing federated learning frameworks' inability to deeply collaborate with security objectives and their limited optimization performance. In summary, this invention achieves efficient, stable, and scalable secure communication resource allocation in dense heterogeneous networks by deeply integrating distributed learning, federated collaboration, and security objectives.
[0010] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0011] Figure 1A flowchart illustrating a method for allocating dense heterogeneous network security communication resources according to an embodiment of the present invention; Figure 2 A schematic diagram of a dense heterogeneous network system model is shown as an example; Figure 3 An exemplary simulation diagram illustrates how the system security rate of a centralized resource allocation scheme varies with the number of training rounds under different eavesdropping modes. Figure 4 An exemplary diagram illustrates the performance of the dense heterogeneous network security communication resource allocation method proposed in this invention under the same scenario; Figure 5 An exemplary diagram illustrates how the system security rate varies with training rounds in a "dual eavesdropper" scenario. Figure 6 An exemplary diagram illustrates how the system's security rate varies with training rounds in a "four eavesdroppers" scenario. Figure 7 This is a schematic diagram of the structure of a dense heterogeneous network security communication resource allocation device provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of a dense heterogeneous network security communication resource allocation device provided in an embodiment of the present invention. Detailed Implementation
[0012] This invention aims to address the key technical challenges of secure communication resource allocation in dense, heterogeneous networks. In complex wireless environments with multiple users, base stations, and eavesdroppers, especially under conditions where eavesdroppers may exhibit both cooperative and non-cooperative behavior, and the system is affected by intra-cell and inter-cell interference as well as dynamic time-varying channel conditions, how can we achieve effective improvement and stable optimization of secure rates without the need for globally accurate channel state information and without sharing raw observation data? Furthermore, traditional centralized optimization methods suffer from high computational complexity and poor scalability in such non-convex problems, while single distributed learning methods struggle to achieve cross-base station collaboration and global performance improvement. Therefore, a resource allocation mechanism that balances distributed training, privacy protection, and global optimization capabilities is urgently needed. To address this, this invention proposes a secure communication resource allocation method based on Federated Deep Reinforcement Learning (FDRL). First, a system model is constructed, including legitimate links, eavesdropping links, and their interference relationships. Secure rate representations are then characterized for both cooperative and non-cooperative scenarios involving multiple eavesdroppers. Based on this, the non-convex secure rate maximization problem, which is inherently difficult to solve directly, is transformed into a multi-agent reinforcement learning problem. A Markov decision process is used to model the system state, actions, and rewards, enabling each micro-base station to learn policies based on local observation information as an independent agent. Furthermore, a federated learning mechanism is introduced, with the macro-base station acting as the central control node to periodically aggregate the local model parameters uploaded by each micro-base station. This achieves cross-base station collaborative optimization and knowledge sharing without leaking the original data, ultimately enabling the system to learn efficient and secure resource allocation strategies in dynamic and complex environments.
[0013] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0014] To achieve efficient, stable, and scalable secure communication resource allocation in dense heterogeneous networks, this invention provides a method for secure communication resource allocation in dense heterogeneous networks. Figure 1 This is a flowchart illustrating a method for allocating dense heterogeneous network security communication resources according to an embodiment of the present invention, as shown below. Figure 1 As shown, it includes: S101. Based on the current dense heterogeneous network, construct a dense heterogeneous network system model.
[0015] The dense heterogeneous network system model includes: macro base stations, multiple micro base stations, and a set of legitimate users served by each micro base station.
[0016] Optionally, the dense heterogeneous network system model also includes: the set of eavesdroppers within the coverage area of each micro base station, and the set of neighboring micro base stations of each micro base station.
[0017] It should be noted that dense heterogeneous networks are a complex network architecture that improves network capacity and coverage by densely deploying various types of wireless access nodes (such as macro base stations and micro base stations). Its core characteristics lie in the spatial multiplexing gain brought about by the multi-layered structure, as well as the resulting dense inter-cell interference and security challenges.
[0018] Figure 2 A schematic diagram of a dense heterogeneous network system model is shown as an example. For example... Figure 2 As shown, the system consists of a central macro base station and multiple micro base stations distributed around it (index 1 to 1). This constitutes a layered coverage. Each micro base station serves a set of multiple legitimate users, while potential eavesdroppers exist within its coverage area (as shown by the red human-shaped icon in the figure). The figure clearly marks four key links: communication links representing legitimate communication (blue solid arrows), intra-cell interference from other users of the same base station (orange lightning arrows), eavesdropping links where eavesdroppers attempt to intercept information (red dashed arrows), and inter-cell interference from adjacent base stations (yellow lightning arrows). This schematic diagram visually presents the system entities, communication relationships, and interference topology, providing a foundation for subsequent modeling and analysis.
[0019] S102. Based on the dense heterogeneous network system model and Shannon's capacity formula, calculate the user link communication rate of all legitimate users and the eavesdropping link communication rate of all eavesdroppers in the current dense heterogeneous network.
[0020] Optionally, S102 includes: Based on a dense heterogeneous network system model, the user-side signal-to-interference-plus-noise ratio (SIN-NMR) of all legitimate users and the eavesdropper-side SIN-NMR of all eavesdroppers are calculated. Based on the user-side signal-to-interference-plus-noise ratio, the user link communication rate is calculated using the Shannon capacity formula. Based on the eavesdropper's side signal-to-interference-plus-noise ratio (SINNR), the communication rate of the eavesdropping link is calculated using the Shannon capacity formula.
[0021] First, for micro base stations PBS With its service users Between, in discrete time slots The channel gain is defined as: in: Indicates the first The first micro base station service A legitimate user in a discrete time slot Small-scale fading coefficient; Indicates the first The first micro base station service A legitimate user in a discrete time slot Large-scale fading coefficient. Satisfying a first-order Gaussian-Markov process: in, This refers to the channel time correlation coefficient; Indicates the first The first micro base station service A legitimate user in a discrete time slot Independent and identically distributed complex Gaussian random variables, Indicates the first The first micro base station service A legitimate user in a discrete time slot The small-scale fading coefficient. For PBS With eavesdroppers The channel between them adopts the same composite fading model as the user side, and its channel gain is expressed as: Small-scale fading also satisfies: .
[0022] Optionally, the user-side signal-to-interference-plus-noise ratio is expressed as: ; The eavesdropper's side signal-to-interference-plus-noise ratio is expressed as: ; Indicates the first The first micro base station service A legitimate user in a discrete time slot User-side signal-to-interference-plus-noise ratio, Indicates the first The first micro base station service A legitimate user in a discrete time slot The transmission power, Indicates the first The first micro base station service A legitimate user in a discrete time slot Channel gain, Indicates the first Index variables of legitimate users of micro base station services , Indicates the first The total number of legitimate users served by each micro base station Indicates the first The first micro base station service A legitimate user in a discrete time slot The transmission power, Indicates the first The first micro base station service A legitimate user in a discrete time slot Channel gain, Indicates the relationship with the first The adjacent micro base station Adjacent micro base stations, Indicates the first The total number of adjacent micro base stations corresponding to each micro base station Indicates adjacent micro base stations The first service A legitimate user, Indicates adjacent micro base stations The total number of all legitimate users served. Indicates the time slot in discrete time No. The neighboring micro base stations are assigned to the first The transmit power of legitimate users, Indicates the first The service of the first adjacent micro base station A legitimate user in a discrete time slot Channel gain, Indicates the first The power of additive white Gaussian noise received by a legitimate user. Indicates the time slot in discrete time No. The eavesdropper attempted to eavesdrop on the first... The micro base station sends to the first The eavesdropper side signal-to-interference-plus-noise ratio during communication between legitimate users Indicates the time slot in discrete time No. From micro base stations to eavesdroppers Channel gain, Indicates the time slot in discrete time No. From adjacent micro base stations to eavesdroppers Channel gain, Indicates eavesdropper The power of the received additive white Gaussian noise.
[0023] Based on Shannon's capacity formula, the signal-to-interference-plus-noise ratio (SIR / NDR) is mapped to the communication rate to obtain the user link communication rate. Communication rate with the eavesdropping link They are respectively: ; ; in This refers to the system bandwidth.
[0024] S103. Construct a confidentiality rate model based on the user link communication rate and the eavesdropping link communication rate.
[0025] The security rate model is an optimization problem that uses the transmit power allocated to corresponding legitimate users by multiple micro base stations as optimization variables, with the goal of maximizing the total security rate of a dense heterogeneous network system model.
[0026] Optionally, the security rate model includes: a first security rate model and a second security rate model; The first confidentiality rate model is a model for non-collusion eavesdropping scenarios, and is equal to the difference between the current legitimate user's user link communication rate and the largest eavesdropping link communication rate among all eavesdroppers; The second confidentiality rate model is a model for collusion eavesdropping scenarios, and is equal to the difference between the current legitimate user's user link communication rate and the sum of the eavesdropping link communication rates of all eavesdroppers.
[0027] After obtaining the user link communication rate and the eavesdropping link communication rate based on the above process, this step constructs a secure rate model and establishes a secure rate maximization problem with power allocation variables as the optimization object. According to physical layer security theory, legitimate users... Security rate Defined as the legal link communication rate Communication rate with the eavesdropping link The non-negative part of the difference is expressed as: in, This indicates that the difference is 0 when it is negative, which is used to ensure the non-negativity of the confidentiality rate. In a multi-eavesdropper environment, consider both non-collusion and collusion eavesdropping scenarios: Non-collusive eavesdropping scenarios: when eavesdroppers are independent of each other, their equivalent eavesdropping capability Determined by the strongest eavesdropping link: The corresponding first confidentiality rate model is: ; Collusive eavesdropping scenario: When eavesdroppers coordinate information, their equivalent eavesdropping capability... For signal-to-interference-plus-noise ratio superposition: The corresponding second security rate model is: ; The optimization problem is constructed with the goal of maximizing the overall security rate of the system. The constructed optimization problem (overall security rate) is expressed as: The constraints are: ,in: For time slots Within PBS Assigned to user The transmit power, as an optimization variable, affects the signal-to-interference-plus-noise ratio. and This further affects communication speed and security rate; Indicates PBS The maximum transmit power is preset by the base station hardware capabilities or system configuration parameters. This represents the minimum communication rate threshold for users, which is preset based on service quality requirements. S104. The confidentiality rate model is modeled as a multi-agent Markov decision process, and each micro base station is regarded as an agent in the multi-agent Markov decision process.
[0028] To address the non-convex optimization problem constructed in the above process, this step models the power allocation optimization process as a multi-agent Markov decision process (MDP) to achieve policy learning based on interactive data. Each PBS is abstracted as an intelligence, and each intelligence operates in discrete time slots. It makes independent decisions and optimizes its strategies through interaction with the environment.
[0029] Optionally, the state of the micro base station in the multi-agent Markov decision-making process is represented as follows: ; in, Indicates the time slot in discrete time No. The environmental conditions that a micro base station bases use to make decisions. Indicates the first The first micro base station service A legitimate user in a discrete time slot Channel gain, Indicates the time slot in discrete time No. From micro base stations to eavesdroppers Channel gain, Indicates the first The first micro base station service A legitimate user in a discrete time slot The transmission power at that time Indicates the first A legitimate user in a discrete time slot Total security rate at that time; The actions of the micro base station in the multi-agent Markov decision process are represented as follows: ; in, Indicates the time slot in discrete time No. The actions of a micro base station Indicates the first The first micro base station service A legitimate user in a discrete time slot The transmission power at that time Indicates the first The first micro base station service A legitimate user, Indicates the first The total number of legitimate users served by each micro base station; The reward function of the micro base station in the multi-agent Markov decision process is expressed as:
[0030] Indicates the time slot in discrete time No. Each micro base station performs actions The corresponding reward at that time Indicates the first weight. Indicates the second weight. Indicates the third weight. Expressing the fourth weight, Indicates the first Each micro base station in discrete time slots Legitimate users serving it Provided user link communication rate, This represents the minimum communication rate threshold. Indicates the first Maximum transmit power budget for a micro base station Indicates the first The total number of neighboring micro base stations corresponding to each micro base station.
[0031] S105. Update the policy parameters of the agents in the multi-agent Markov decision-making process based on the policy gradient method to obtain the local policy model parameters of each micro base station.
[0032] In this step, the decision-making strategy of each micro base station (agent) is first modeled as a parameterized policy function: ;in, This represents the parameter vector of the policy network. The agent's optimization objective is to maximize the expected cumulative reward it obtains during interaction with the environment. This objective function is defined as: ;in, Indicates micro base station The cumulative returns obtained Indicating targeting The expected value is obtained from the distribution. To achieve this goal, the policy gradient method is used for optimization. Objective function Regarding parameters The gradient can be expressed as: Subsequently, the policy parameters are updated based on this gradient, as follows: , This represents the learning rate.
[0033] S106. Based on the federated learning framework and macro base stations, the parameters of the local policy model are collaboratively optimized until the convergence condition is reached.
[0034] Optionally, S106 includes: S201. The macro base station generates and distributes the initial global policy model parameters to all micro base stations.
[0035] S202. Each micro base station updates its local policy model parameters based on the received current global policy model parameters, and obtains the locally updated policy model parameters.
[0036] S203. Each micro base station uploads the locally updated strategy model parameters to the macro base station.
[0037] S204. The macro base station performs aggregation calculations on all uploaded locally updated policy model parameters to generate new global policy model parameters.
[0038] S205. Determine whether the convergence condition is met; if not, distribute the new global policy model parameters to each micro base station as the current global policy model parameters in S202 of the next iteration.
[0039] S206. Repeat steps S202-S205 until the convergence condition is met.
[0040] Optionally, the convergence condition is that the norm of change of the global policy model parameters, which are composed of local policy model parameters, is less than a preset threshold in multiple consecutive collaborative optimization processes.
[0041] In the aforementioned S106 process, firstly, the macro base station, acting as the central node, generates an initial set of global policy model parameters and distributes them to all micro base stations in the network. Next, each micro base station, upon receiving these parameters, uses them as a basis to train its own local policy model parameters in conjunction with local network environment data. Subsequently, each micro base station uploads its updated local parameters back to the macro base station. The macro base station aggregates all the local parameters uploaded by the micro base stations and performs a fusion calculation using a preset aggregation algorithm (such as weighted average) to generate a new generation of global policy model parameters with superior performance. Afterward, the system determines whether the newly generated global parameters meet preset convergence conditions. If the convergence conditions are not met, the macro base station distributes these new parameters back to each micro base station as the starting point for local training in the next iteration. The process then repeats from the step of "each micro base station updating its local parameters based on the received current global parameters," forming a closed-loop iteration of "local training → parameter upload → global aggregation → judgment." This cycle will continue until the parameters of the generated global strategy model finally meet the convergence condition. At this point, the collaborative optimization process ends, and each micro base station can use this converged model to make the final resource allocation decision.
[0042] It should be further noted that, without changing the overall technical chain constituted by the above steps, the present invention also allows for the replacement of some specific implementations: for example, the policy optimization method in the above process can adopt other deep reinforcement learning algorithms, and the action space can be expanded to the form of continuous power control.
[0043] S107. Determine the transmit power allocation strategy based on the local strategy model parameters corresponding to the convergence condition.
[0044] This invention provides a method for allocating network communication resources in a dense heterogeneous network, comprising: constructing a dense heterogeneous network system model based on the current dense heterogeneous network; the dense heterogeneous network system model includes: a macro base station, multiple micro base stations, and a set of legitimate users served by each micro base station; calculating the user link communication rate of all legitimate users and the eavesdropping link communication rate of all eavesdroppers in the current dense heterogeneous network based on the dense heterogeneous network system model and the eavesdropping link communication rate; constructing a security rate model based on the user link communication rate and the eavesdropping link communication rate; the security rate model is an optimization problem with the transmit power allocated by multiple micro base stations to corresponding legitimate users as optimization variables and the goal of maximizing the total security rate of the dense heterogeneous network system model; modeling the security rate model as a multi-agent Markov decision process, and treating each micro base station as an agent in the multi-agent Markov decision process; updating the policy parameters corresponding to the agents in the multi-agent Markov decision process based on the policy gradient method to obtain the local policy model parameters corresponding to each micro base station; performing collaborative optimization processing on the local policy model parameters based on a federated learning framework and a macro base station until the convergence condition is reached; and determining the transmit power allocation strategy based on the local policy model parameters corresponding to the convergence condition. In this invention, firstly, by constructing a security rate model and aiming to maximize the overall security rate, the physical layer security performance is directly optimized. This addresses the problems of existing research's simplistic modeling of eavesdropping environments and its inability to fully characterize the collaborative threats of multiple eavesdroppers, thereby improving the effectiveness of the strategy in real-world complex environments. Secondly, the security rate model is modeled as a multi-agent Markov decision process, employing a distributed learning framework based on policy gradient methods. This allows each micro-base station to make independent decisions as an agent, solving the problems of computational complexity, high signaling overhead, and insufficient scalability inherent in traditional centralized optimization methods. Next, a federated learning framework is introduced, enabling macro-base stations to collaboratively optimize local policy model parameters. This achieves global experience sharing and a collaborative mechanism, addressing the policy instability and weak generalization capabilities of distributed reinforcement learning methods due to a lack of efficient sharing. Finally, this invention deeply integrates federated learning with the physical layer security objective (security rate) and wireless resource constraints (transmit power), directly integrating security and resource factors during the optimization process. This solves the problems of existing federated learning frameworks' inability to deeply collaborate with security objectives and their limited optimization performance. In summary, this invention achieves efficient, stable, and scalable secure communication resource allocation in dense heterogeneous networks by deeply integrating distributed learning, federated collaboration, and security objectives.
[0045] To verify the effectiveness and superiority of the proposed dense heterogeneous network security communication resource allocation method, comparative experiments were conducted in a typical wireless communication simulation environment with centralized resource allocation methods and distributed methods under different parameter configurations to evaluate its performance under conditions of multiple eavesdroppers, multiple interferences, and dynamic channels. All experiments were implemented using a deep learning framework. By constructing simulation environments with multiple base stations, multiple users, and multiple eavesdroppers, the changing trends of system security rates and model convergence characteristics under different eavesdropping modes and different federated training mechanisms were analyzed, thereby comprehensively verifying the effectiveness of the proposed method.
[0046] Simulation Scenario and Parameter Settings: To comprehensively evaluate the applicability and robustness of this invention in complex wireless environments, a dense heterogeneous network scenario was constructed, comprising one macro base station and multiple micro base stations. Each micro base station is distributed in different hotspot areas and serves multiple users, while multiple eavesdroppers are deployed within the coverage area of the micro base stations. Typical parameter settings are shown in Table 1. Under the above parameter conditions, channel models for legitimate links and eavesdropping links were constructed, and the system security rate was defined based on the signal-to-interference-plus-noise ratio (SIR) of legitimate users and the SIR of eavesdroppers. The performance of the proposed method in different scenarios was verified while satisfying transmit power constraints and user quality of service constraints.
[0047] Model and Training Setup: During the resource allocation strategy learning process, each micro base station corresponds to an agent, and its policy network uses a deep neural network for function approximation, with local state information as the input. The output is the transmit power allocation policy for each user; the local model is updated using the policy gradient method, and the objective function is optimized using Monte Carlo reward estimation. The parameter iteration is completed. During the federated learning phase, each base station uploads its local model parameters to the macro base station according to the set aggregation period. The macro base station performs weighted aggregation to obtain the global model and then distributes it to each node for continued training, thereby achieving distributed collaborative optimization without sharing the original data.
[0048] Evaluation metrics: To objectively evaluate the performance of the method of this invention, the average security rate of the system is used as the core evaluation metric. At the same time, the model convergence speed, performance differences under different eavesdropping scenarios, and the impact of changes in federated aggregation strategies on system performance are comprehensively analyzed. The security rate is defined based on the difference between the capacity of the legitimate link and the capacity of the eavesdropping link, and is evaluated under two typical scenarios of non-cooperative eavesdropping and cooperative eavesdropping, so as to comprehensively reflect the adaptability and stability of the method under different security threat conditions.
[0049] Under the above-mentioned unified experimental setup, several specific embodiments of the present invention are given below to further illustrate the practical effects and application value of the technical solution of the present invention.
[0050] Table 1 Typical Simulation Parameter Settings
[0051] In this embodiment, the proposed dense heterogeneous network security communication resource allocation method based on federated deep reinforcement learning (FDRL) is compared and analyzed with an idealized centralized resource allocation scheme. Under the same network parameter configuration, the system security rate of the two schemes under different eavesdropping modes (cooperative and non-cooperative) is simulated and evaluated, and the results are as follows. Figure 3 and Figure 4 As shown.
[0052] Figure 3 An exemplary simulation diagram illustrates how the system security rate of a centralized resource allocation scheme varies with training rounds under different eavesdropping modes. For example... Figure 3 As shown, the system's security rate converges rapidly in both non-cooperative and cooperative eavesdropping scenarios. However, contrary to theoretical expectations, the results show that the final security rate levels under the two eavesdropping modes are extremely close, stabilizing at approximately 0.11 bps, and the convergence curves almost overlap. This phenomenon may stem from the fact that, under the ideal assumption of perfect global information, the resource allocation strategy formulated by centralized optimization has similar resistance capabilities to the two eavesdropping threats. This allows the performance loss caused by cooperative eavesdropping to be largely compensated for in the simulation, failing to intuitively reflect its additional threat.
[0053] Figure 4 An exemplary diagram illustrates the performance of the dense heterogeneous network security communication resource allocation method proposed in this invention under the same scenario. For example... Figure 4 As shown, under distributed, local observation conditions, the difference in the impact of the two eavesdropping modes on system security performance is significant: for the non-cooperative eavesdropping scenario, the system confidentiality rate starts from approximately 0.1 bps and eventually converges and stabilizes at a relatively high level of approximately 0.2 bps; while for the cooperative eavesdropping scenario, the confidentiality rate starts from approximately 0.07 bps and eventually stabilizes at approximately 0.1 bps. The two curves are clearly separated, and the final performance in the cooperative eavesdropping scenario is approximately half that of the non-cooperative scenario. This clearly indicates that information sharing among multiple eavesdroppers significantly weakens system security performance, consistent with theoretical analysis.
[0054] Comprehensive comparison Figure 3 and Figure 4It is known that although centralized schemes may achieve higher or more balanced performance at theoretical limits, they rely on unrealistic assumptions about global information. In contrast, the FDRL distributed scheme proposed in this invention, although its absolute performance (~0.1 bps) in the collaborative eavesdropping scenario is slightly lower than the simulation results of the centralized scheme (~0.11 bps), can still effectively learn and converge under a fully distributed decision-making framework. It also reveals more realistically the differentiated impact of different eavesdropping threat patterns on the system, verifying the effectiveness of the proposed method in actual deployment and its ability to characterize complex threats.
[0055] Next, in this embodiment, based on the dense heterogeneous network security communication resource allocation method based on federated deep reinforcement learning proposed in this invention, the impact of the number of eavesdroppers and the federated aggregation strategy on the final security performance of the system is further analyzed. By setting different numbers of eavesdroppers and different model aggregation periods, the convergence of the system's security rate was simulated, and the results are as follows. Figure 5 and Figure 6 As shown.
[0056] Figure 5 An exemplary diagram illustrates how the system's security rate varies with training rounds in a "two-eavesdropper" scenario. Figure 5 As shown, when the number of eavesdroppers is relatively small (2 per cell), the overall system security rate converges to a relatively high level (approximately 0.25 bps). Two key phenomena can be observed: First, the final security rate in the non-cooperative eavesdropping scenario (orange and blue curves) is significantly higher than that in the cooperative eavesdropping scenario (green and purple curves), verifying that cooperation among eavesdroppers can seriously compromise system security. Second, under the same eavesdropping mode, the curves corresponding to shorter model aggregation cycles (Agg Per=100, i.e., more frequent global model updates) (orange and green) have better convergence speed and final performance than the curves (blue and purple) with longer aggregation cycles (Agg Per=1000). This indicates that in scenarios with relatively mild threats, more frequent exchange and integration of learning experiences from each node through a federated learning mechanism can effectively accelerate policy optimization and improve overall security performance.
[0057] Figure 6 An exemplary diagram illustrates how the system's security rate varies with training rounds in a "four-eavesdropper" scenario. Figure 6 As shown, when the eavesdropping threat increases (the number of eavesdroppers per cell increases to 4), the overall security performance of the system declines, eventually reducing the confidentiality rate to approximately 0.15 bps (non-cooperative) and 0.05 bps (cooperative). (Comparison) Figure 5 and Figure 6It can be observed that: First, the performance loss caused by collaborative eavesdropping is further amplified in a strong eavesdropping environment, and its confidentiality rate is much lower than in non-collaborative scenarios. Second, and more importantly, the performance gap between different aggregation cycle curves has significantly narrowed. Figure 6 In both non-cooperative (orange and blue) and cooperative (green and purple) scenarios, the converged values of the two curves with aggregation periods of 100 and 1000 are very close. This indicates that in a strong eavesdropping environment, the "bottleneck" of system performance is more determined by the eavesdropper's own powerful eavesdropping capabilities, while the performance improvement potential brought about by optimizing resource allocation strategies through federated learning becomes relatively limited.
[0058] In summary, the simulation results of this embodiment demonstrate that the dense heterogeneous network security communication resource allocation method proposed in this invention can operate stably under different eavesdropping scales and threat modes, exhibiting good adaptability. While federated aggregation mechanisms can effectively utilize distributed experience to improve performance, their effectiveness is limited by the actual security environment. When the eavesdropping threat is weak, frequent model aggregation can bring significant gains; however, when the eavesdropping threat is extremely strong, the limitations of this optimization method must be recognized, and it should be considered to combine it with other enhancement technologies such as physical layer security coding to build a more robust defense system. This reflects the robustness of the proposed solution to complex dynamic environments and also provides important design basis for parameter configuration (such as aggregation cycle) in its actual deployment.
[0059] The beneficial effects of this invention, which addresses the complex security threats and resource allocation challenges in dense, heterogeneous networks, can be systematically summarized as follows: First, accurate modeling and efficient optimization: By constructing the problem of maximizing the security rate as a multi-agent deep reinforcement learning problem, and comprehensively considering intra-cell / inter-cell interference, as well as cooperative and non-cooperative eavesdropping scenarios in the modeling, this invention achieves accurate characterization and dynamic adaptive optimization of complex security environments, fundamentally solving the problem that traditional methods simplify threat modeling and are difficult to cope with real complex environments.
[0060] Secondly, distributed decision-making and high scalability: By constructing a distributed multi-agent learning framework, each micro base station makes decisions based solely on local observations, significantly reducing reliance on global channel state information and centralized signaling overhead. This effectively overcomes the bottlenecks of computational complexity and poor scalability in traditional centralized optimization methods, significantly improving the deployability and real-time response capability of the method in practical networks.
[0061] Furthermore, collaborative learning and privacy protection: An innovative federated learning mechanism is introduced, enabling base stations to achieve efficient experience sharing and collaborative optimization across base stations through encrypted model parameter aggregation, without sharing raw sensitive data (such as local channel information and user location). This solves the problems of policy isolation and weak generalization ability in distributed reinforcement learning, while enhancing the system's privacy protection capabilities and the overall collaborative defense level of the network.
[0062] Ultimately, performance robustness and practical value: Experiments show that the proposed method can converge stably under various eavesdropping threats and dynamic environments, significantly improving the system's security rate and maintaining good performance in large-scale networks. This verifies its strong robustness and practicality, achieving an effective balance between security performance, resource overhead, and system complexity.
[0063] In summary, this invention creatively integrates distributed decision-making, secure collaborative optimization, and privacy protection through a federated deep reinforcement learning framework. It not only achieves adaptive and efficient resource allocation in complex security environments of dense heterogeneous networks, but also theoretically solves the limitations of traditional methods in terms of modeling, scalability, collaboration, and privacy. This provides an efficient, reliable, and scalable solution for the secure deployment of practical wireless communication systems.
[0064] The method provided in this embodiment of the invention can be applied to electronic devices. Specifically, the electronic device can be a desktop computer, a portable computer, a smart mobile terminal, a server, etc., and this embodiment of the invention does not limit the application to such devices.
[0065] Based on the same inventive concept, embodiments of the present invention also provide a dense heterogeneous network security communication resource allocation device. Figure 7 This is a schematic diagram of a dense heterogeneous network security communication resource allocation device provided in an embodiment of the present invention, as shown below. Figure 7 As shown, it includes: a model building unit 701, a calculation unit 702, an update unit 703, a collaborative optimization unit 704, and an output unit 705; Model building unit 701 is used to build a dense heterogeneous network system model based on the current dense heterogeneous network; the dense heterogeneous network system model includes: macro base stations, multiple micro base stations, and a set of legitimate users served by each micro base station; The computing unit 702 is used to calculate the user link communication rate of all legitimate users and the eavesdropping link communication rate of all eavesdroppers in the current dense heterogeneous network based on the dense heterogeneous network system model and Shannon's capacity formula. The model building unit 701 is also used to build a confidentiality rate model based on the user link communication rate and the eavesdropping link communication rate. The confidentiality rate model is an optimization problem with the transmit power allocated to the corresponding legitimate users by multiple micro base stations as optimization variables and the goal of maximizing the total confidentiality rate of the dense heterogeneous network system model. The confidentiality rate model is modeled as a multi-agent Markov decision process, and each micro base station is regarded as an agent in the multi-agent Markov decision process. Update unit 703 is used to update the policy parameters of agents in the multi-agent Markov decision-making process based on the policy gradient method, so as to obtain the local policy model parameters of each micro base station. The collaborative optimization unit 704 is used to perform collaborative optimization processing on the parameters of the local policy model based on the federated learning framework and macro base station until the convergence condition is reached. Output unit 705 is used to determine the transmit power allocation strategy based on the local policy model parameters corresponding to the convergence condition.
[0066] Figure 8 This is a schematic diagram of a dense heterogeneous network security communication resource allocation device provided in an embodiment of the present invention. It includes a processor 710, a storage medium 720, and a bus 730. The storage medium 720 stores machine-readable instructions executable by the processor 710. When the dense heterogeneous network security communication resource allocation device is running, the processor 710 communicates with the storage medium 720 via the bus 730. The processor 710 executes the machine-readable instructions to perform the steps of the above-described method embodiment. Specific implementation methods and technical effects are similar and will not be described in detail here.
[0067] The storage medium may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the storage medium may also be at least one storage device located remotely from the aforementioned processor.
[0068] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0069] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings and the disclosure, will understand and implement other variations of the disclosed embodiments in carrying out the claimed invention. In this description, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.
[0070] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the inventive concept, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A method for allocating dense heterogeneous network security communication resources, characterized in that, include: Based on current dense heterogeneous networks, construct a dense heterogeneous network system model; The dense heterogeneous network system model includes: a macro base station, multiple micro base stations, and a set of legitimate users served by each micro base station; Based on the dense heterogeneous network system model and Shannon capacity formula, calculate the user link communication rate of all legitimate users and the eavesdropping link communication rate of all eavesdroppers in the current dense heterogeneous network. A confidentiality rate model is constructed based on the user link communication rate and the eavesdropping link communication rate. The confidentiality rate model is an optimization problem with the goal of maximizing the total confidentiality rate of the dense heterogeneous network system model, using the transmission power allocated to the corresponding legitimate users by multiple micro base stations as optimization variables. The confidentiality rate model is modeled as a multi-agent Markov decision process, and each of the micro base stations is regarded as an agent in the multi-agent Markov decision process. The policy parameters of each agent in the multi-agent Markov decision-making process are updated based on the policy gradient method to obtain the local policy model parameters of each micro base station. Based on the federated learning framework and the macro base station, the parameters of the local policy model are collaboratively optimized until the convergence condition is reached. The transmit power allocation strategy is determined based on the local policy model parameters corresponding to the convergence condition.
2. The method for allocating dense heterogeneous network security communication resources according to claim 1, characterized in that, The dense heterogeneous network system model also includes: a set of eavesdroppers within the coverage area of each micro base station, and a set of neighboring micro base stations of each micro base station.
3. The method for allocating dense heterogeneous network security communication resources according to claim 1, characterized in that, The calculation of the user link communication rate for all legitimate users and the eavesdropping link communication rate for all eavesdroppers in the current dense heterogeneous network, based on the dense heterogeneous network system model and Shannon capacity formula, includes: Based on the aforementioned dense heterogeneous network system model, calculate the user-side signal-to-interference-plus-noise ratio (SIN-NMR) of all legitimate users and the eavesdropper-side SIN-NMR of all eavesdroppers. Based on the user-side signal-to-interference-plus-noise ratio, the user link communication rate is calculated using the Shannon capacity formula; Based on the eavesdropper's side signal-to-interference-plus-noise ratio, the communication rate of the eavesdropping link is calculated using the Shannon capacity formula.
4. The method for allocating dense heterogeneous network security communication resources according to claim 3, characterized in that, The user-side signal-to-interference-plus-noise ratio is expressed as: ; The eavesdropper's side signal-to-interference-plus-noise ratio is expressed as: ; Indicates the first The first micro base station service A legitimate user in a discrete time slot User-side signal-to-interference-plus-noise ratio, Indicates the first The first micro base station service A legitimate user in a discrete time slot The transmission power, Indicates the first The first micro base station service A legitimate user in a discrete time slot Channel gain, Indicates the first Index variables of legitimate users of micro base station services , Indicates the first The total number of legitimate users served by each micro base station Indicates the first The first micro base station service A legitimate user in a discrete time slot The transmission power, Indicates the first The first micro base station service A legitimate user in a discrete time slot Channel gain, Indicates the relationship with the first The adjacent micro base station Adjacent micro base stations, Indicates the first The total number of adjacent micro base stations corresponding to each micro base station Indicates adjacent micro base stations The first service A legitimate user, Indicates adjacent micro base stations The total number of all legitimate users served. Indicates the time slot in discrete time No. The neighboring micro base stations are assigned to the first The transmit power of legitimate users, Indicates the first The service of the first adjacent micro base station A legitimate user in a discrete time slot Channel gain, Indicates the first The power of additive white Gaussian noise received by a legitimate user. Indicates the time slot in discrete time No. The eavesdropper attempted to eavesdrop on the first... The micro base station sends to the first The eavesdropper side signal-to-interference-plus-noise ratio during communication between legitimate users Indicates the time slot in discrete time No. From micro base stations to eavesdroppers Channel gain, Indicates the time slot in discrete time No. From adjacent micro base stations to eavesdroppers Channel gain, Indicates eavesdropper The power of the received additive white Gaussian noise.
5. The method for allocating dense heterogeneous network security communication resources according to claim 1, characterized in that, The confidentiality rate model includes: a first confidentiality rate model and a second confidentiality rate model; The first confidentiality rate model is a model for non-collusion eavesdropping scenarios, and is equal to the difference between the current legitimate user's user link communication rate and the largest eavesdropping link communication rate among all eavesdroppers; The second confidentiality rate model is a model for collusion eavesdropping scenarios, and is equal to the difference between the current legitimate user's user link communication rate and the sum of the eavesdropping link communication rates of all eavesdroppers.
6. The method for allocating dense heterogeneous network security communication resources according to claim 1, characterized in that, The state of the micro base station in the multi-agent Markov decision-making process is represented as follows: ; in, Indicates the time slot in discrete time No. The environmental conditions that a micro base station bases use to make decisions. Indicates the first The first micro base station service A legitimate user in a discrete time slot Channel gain, Indicates the time slot in discrete time No. From micro base stations to eavesdroppers Channel gain, Indicates the first The first micro base station service A legitimate user in a discrete time slot The transmission power at that time Indicates the first A legitimate user in a discrete time slot Total security rate at that time; The actions of the micro base station in the multi-agent Markov decision-making process are represented as follows: ; in, Indicates the time slot in discrete time No. The actions of a micro base station Indicates the first The first micro base station service A legitimate user in a discrete time slot The transmission power at that time Indicates the first The first micro base station service A legitimate user, Indicates the first The total number of legitimate users served by each micro base station; The reward function of the micro base station in the multi-agent Markov decision-making process is expressed as: Indicates the time slot in discrete time No. Each micro base station performs actions The corresponding reward at that time Indicates the first weight. Indicates the second weight. Indicates the third weight. Expressing the fourth weight, Indicates the first Each micro base station in discrete time slots Legitimate users serving it Provided user link communication rate, This represents the minimum communication rate threshold. Indicates the first Maximum transmit power budget for a micro base station Indicates the first The total number of neighboring micro base stations corresponding to each micro base station.
7. The method for allocating dense heterogeneous network security communication resources according to claim 1, characterized in that, The process of collaboratively optimizing the local policy model parameters based on the federated learning framework and the macro base station until the convergence condition is met includes: S201, The macro base station generates and distributes the initial global strategy model parameters to all micro base stations; S202. Each of the micro base stations updates its local policy model parameters based on the received current global policy model parameters to obtain locally updated policy model parameters. S203. Each micro base station uploads the locally updated strategy model parameters to the macro base station; S204. The macro base station performs aggregation calculations on all uploaded locally updated policy model parameters to generate new global policy model parameters. S205. Determine whether the convergence condition is met; if not, send the new global strategy model parameters to each micro base station as the current global strategy model parameters in S202 of the next iteration. S206. Repeat steps S202-S205 until the convergence condition is met.
8. The method for allocating dense heterogeneous network security communication resources according to claim 1, characterized in that, The convergence condition is that the norm of change of the global policy model parameters, which are composed of the local policy model parameters, is less than a preset threshold in multiple consecutive collaborative optimization processes.
9. A dense heterogeneous network security communication resource allocation device, characterized in that, The dense heterogeneous network security communication resource allocation device includes: a model building unit, a computing unit, an updating unit, a collaborative optimization unit, and an output unit; The model building unit is used to construct a dense heterogeneous network system model based on the current dense heterogeneous network; the dense heterogeneous network system model includes: macro base stations, multiple micro base stations, and a set of legitimate users served by each micro base station; The computing unit is used to calculate, based on the dense heterogeneous network system model and Shannon capacity formula, the user link communication rate of all legitimate users and the eavesdropping link communication rate of all eavesdroppers in the current dense heterogeneous network. The model building unit is further configured to build a security rate model based on the user link communication rate and the eavesdropping link communication rate; the security rate model is an optimization problem with the transmit power allocated to the corresponding legitimate user by multiple micro base stations as the optimization variable and the goal of maximizing the total security rate of the dense heterogeneous network system model; the security rate model is modeled as a multi-agent Markov decision process, and each micro base station is regarded as an agent in the multi-agent Markov decision process; The updating unit is used to update the policy parameters corresponding to the agents in the multi-agent Markov decision-making process based on the policy gradient method, so as to obtain the local policy model parameters corresponding to each micro base station. The collaborative optimization unit is used to perform collaborative optimization processing on the parameters of the local policy model based on the federated learning framework and the macro base station until the convergence condition is reached. The output unit is used to determine the transmit power allocation strategy based on the local policy model parameters corresponding to the convergence condition.
10. A dense heterogeneous network security communication resource allocation device, characterized in that, include: The device includes a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the dense heterogeneous network security communication resource allocation device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the dense heterogeneous network security communication resource allocation method as described in any one of claims 1-8.