Deterministic time delay cellular-free resource allocation method based on hierarchical distributed security reinforcement learning

By employing a hierarchical distributed security reinforcement learning approach, the problems of high computational complexity and insufficient adaptability in resource allocation in non-cellular networks are solved, achieving low-latency and high-spectral-efficiency wireless access that adapts to dynamic wireless environments.

CN121772010APending Publication Date: 2026-03-31SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing non-cellular networks suffer from high computational complexity and insufficient adaptability in resource allocation in latency-sensitive scenarios, making it difficult to meet the heterogeneous latency requirements of multi-user networks. Furthermore, traditional methods struggle to respond quickly in dynamic environments.

Method used

A hierarchical distributed security reinforcement learning approach is adopted to construct a multi-point collaborative problem of joint time-frequency resource allocation and beam optimization for cellular-free downlink communication. Through a hierarchical distributed multi-agent reinforcement learning algorithm framework, the problem is decomposed into resource block allocation on the upper-layer central processing unit (CU) side and beam selection on the lower-layer distributed unit (DU) side, and optimization is carried out by combining security reinforcement learning and multi-agent reinforcement learning algorithms.

Benefits of technology

It significantly reduces computational complexity, achieves efficient resource scheduling and dynamic optimization, meets strict packet latency constraints, improves spectrum efficiency, and adapts to dynamically changing wireless environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121772010A_ABST
    Figure CN121772010A_ABST
Patent Text Reader

Abstract

The invention discloses a deterministic time delay cellular-free resource allocation method based on hierarchical distributed security reinforcement learning. A layered and distributed security reinforcement learning framework under the constraint of a time delay satisfaction rate is provided, a joint time-frequency resource allocation and beam optimization problem is decomposed into an upper layer sub-task and a lower layer sub-task, and modeling is performed as a constraint Markov decision process (CMDP) and a Markov decision process (MDP). An upper layer central processing unit (CU) realizes time-frequency resource block allocation oriented to time delay satisfaction rate determinacy guarantee by introducing time delay cost and security constraints; and a plurality of intelligent agents at the lower-layer distributed access unit (DU) side cooperatively execute beam selection to ensure the transmission reliability. According to the method, the calculation complexity is remarkably reduced through hierarchical decoupling, the system resource utilization rate is remarkably improved while the time delay determinacy constraint is ensured by introducing a safety reinforcement learning mechanism, and the method has good system expandability and engineering application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deterministic wireless transmission and AI communication fusion technology, and in particular to a deterministic latency-free cellular resource allocation method based on hierarchical distributed security reinforcement learning. Background Technology

[0002] In recent years, non-cellular networks have become a key technology for meeting deterministic latency requirements due to their spatial diversity gain and uniform quality of service brought by their distributed architecture. However, resource allocation remains a core challenge restricting their performance. In latency-sensitive scenarios, user equipment often has heterogeneous and strict latency constraints. Any resource scheduling error can lead to communication conflicts, resulting in latency defaults and severely impacting the reliability of critical applications such as industrial control and vehicle-to-everything (V2X) communication. Even more challenging is that dynamic channel conditions and user mobility change the network state in real time, making it difficult for traditional optimization methods to respond quickly to these changes.

[0003] Existing research primarily addresses latency-constrained problems through dynamic programming and convex optimization. For example, Markov decision processes are used to model packet scheduling, or convex differential programming is employed to maximize throughput under QoS constraints. However, these methods suffer from two major drawbacks: firstly, computational complexity increases exponentially with network size; and secondly, they lack adaptability in rapidly changing environments. While reinforcement learning has been introduced into this field due to its ability to handle nonlinear problems, conventional algorithms often focus on long-term cumulative rewards and lack explicit control mechanisms for instantaneous latency default probabilities. Therefore, considering the dynamic coordination requirements of distributed APs in practical non-cellular networks and the strict constraints of latency-sensitive services, an intelligent and efficient resource allocation mechanism is needed that can meet the heterogeneous latency requirements of multi-user networks while maintaining low computational complexity in distributed decision-making. Summary of the Invention

[0004] Technical Problem: To address the shortcomings of existing technologies, this invention proposes a deterministic latency-free cellular resource allocation method based on hierarchical distributed security reinforcement learning, in order to achieve low-latency wireless access and improve network resource utilization.

[0005] Technical solution: A deterministic, latency-free cellular resource allocation method based on hierarchical distributed security reinforcement learning, comprising the following steps:

[0006] S1 addresses the deterministic latency constraints in non-cellular network scenarios by constructing a joint time-frequency resource allocation and beam optimization problem for multi-point collaborative non-cellular downlink communication.

[0007] S2, construct a time-delay-aware hierarchical distributed multi-agent reinforcement learning algorithm framework, decompose the joint optimization problem into upper and lower sub-tasks. The upper-layer central processing unit (CU) uses security reinforcement learning to solve the resource block allocation. After determining the resource allocation result, each access point (AP) agent on the lower-layer distributed unit (DU) completes beam selection.

[0008] S3 trains the parameters of the constructed time-delay-aware hierarchical distributed multi-agent reinforcement learning algorithm framework to obtain the optimal policy.

[0009] Furthermore, the non-cellular network scenario includes a central processing unit deployed on the CU side, distributed access points (APs) deployed on the DU side, and several users; considering an OFDM downlink system, where B APs each have M antennas, collaboratively serving U single-antenna users; the system bandwidth is divided into K subbands, each with a bandwidth of... Each subband is the smallest schedulable frequency unit, containing a resource block RB, which consists of C consecutive subcarriers. Each time slot contains N OFDM symbols, meaning the time slot length is... Therefore, the minimum scheduling time-frequency unit (TFU) contains C*N resource elements (RE); each is represented by a different resource element (RE). , and This represents the set of access points (APs), users, and subbands; using an equal power allocation scheme, the power allocated by AP b to user u is... , Total transmit power for each access point; for the same user All arriving data packets have the same deadline. Data packets in each user buffer should be scheduled in a first-in, first-out (FIFO) manner; user data packets must be transmitted before the deadline, otherwise they will be lost.

[0010] The constraints of the joint time-frequency resource allocation and beam optimization problem for multi-point cooperative downlink communication include the user's maximum delay violation probability constraint; the objective function of the optimization problem is to maximize spectral efficiency; the optimization variables are resource block allocation and the beam selected by each AP for the user. In step S1, the joint time-frequency resource allocation and beam optimization problem for multi-point cooperative downlink communication is specifically as follows: ;

[0011] in, Indicates user exist Is the time in the sub-band? If scheduled, then The value is 1 if it is set to 1, and 0 otherwise. For DFT codebook set, Indicates AP For users Selected codeword index; Indicates user In the time slot The achievable rate at that time The bandwidth of each subband, where (·) indicates an indicator function, which takes the value 1 when the condition inside the parentheses is true, and takes the value 0 otherwise. Represents a time set; (·) represents a probability function, that is, the probability of the event within the parentheses occurring; Indicates user exist The time reached The latency experienced by each data packet; Indicates user The packet delay violates the probability requirement; constraints This indicates that the probability of a user data packet delay violation does not exceed [a certain value]. ;constraint express Use 0 and 1 to indicate whether to schedule; constraints Indicates AP For users The selected codeword index comes from the DFT codebook set. constraint Indicates user In the time slot It was scheduled to a sub-band.

[0012] Furthermore, step S2 is as follows:

[0013] Step S21: Based on the original optimization problem, the original problem is decomposed into two interrelated sub-problems, and a hierarchical architecture is designed to solve the time-frequency resource allocation and beam selection problems respectively.

[0014] Step S22: The time-frequency resource allocation problem related to time delay is modeled as a constrained Markov decision process (CMDP), and the beam selection problem is modeled as a Markov decision process (MDP).

[0015] Step S23: For the Constrained Markov Decision Process (CMDP), a secure reinforcement learning algorithm is used to solve the time-frequency resource allocation problem, and the allocation results are transmitted to each AP.

[0016] In step S24, each AP agent selects the optimal beam for the user based on the resource block allocation result using a multi-agent reinforcement learning algorithm.

[0017] The layered architecture designed in step S21 is as follows:

[0018] The hierarchical distributed architecture consists of an upper-layer central processing unit (CU) and a lower-layer distributed unit (DU), which respectively complete resource allocation and beam selection tasks. The upper-layer CU deploys several user agents, and each access point (AP) in the lower layer deploys one agent. The CU is responsible for centrally acquiring and processing information and allocating resource blocks to users. The DU then selects the optimal beam for transmission to users based on the allocation result.

[0019] Furthermore, in step S22, the constrained Markov decision process (CMDP) and the Markov decision process (MDP) are specifically modeled as follows:

[0020] The state space of each agent in the CU-side resource allocation module is defined as the joint state of all users, i.e. ,in, Representing the status of each user, by composition; Depend on Composition, representing users Channel state information for all APs, Indicates AP With users In sub-band Channel state information; Indicates user The length of the buffer queue, Indicates user The number of data packets dropped at the previous time point. Indicates user The maximum tolerable latency for the current buffer queue header packets; actions of each user agent. Instructing users In the time slot The assigned sub-band index; rewards Defined as a time slot Real-time spectral efficiency; cost function, which is the degree to which all users violate latency requirements, per time slot. The immediate cost at the end is expressed as , Indicates user The packet delay violates the probability requirement;

[0021] The state space of each agent in the DU side beam selection module Defined as AP In time slots between each user Equivalent channel gain at time, AP In time The encoding index of the beamcodebook selected for each user, and the system in time. and time The number of RBs used, and the amount of data received by each user from the access point AP. Total interference power of signals sent to other users and AP Between each user in time Select channel state information on the selected resource block; the global state should reveal information about all users on each access point; therefore, integrate the local state of each agent into the global state; each AP agent's action is... ,in AP represents time t Assigned to user The codeword index; consistent with the reward defined by the resource allocation module, instant reward. Defined as the system in time slot Spectral efficiency.

[0022] Furthermore, in step S23, the security reinforcement learning algorithm used on the CU side is specifically designed as follows:

[0023] Step S231: Apply the Lagrange relaxation method to construct the Lagrange function, transforming the Constrained Markov Decision Process (CMDP) problem into an equivalent unconstrained optimization problem.

[0024] Step S232: The Lagrange optimization method is combined with the Dueling DQN reinforcement learning algorithm to alternately update the Lagrange multipliers and policy network parameters to obtain the optimal resource allocation.

[0025] Furthermore, in step S24, the QPLEX multi-agent reinforcement learning algorithm is used on the DU side:

[0026] On the DU side, each access point is equipped with a QPLEX-based agent, which is trained under a centralized training and distributed execution CTDE framework; during the execution phase, each agent selects the optimal beam for the user based on its local state.

[0027] Furthermore, in step S3, the training process of the time-delay-aware hierarchical distributed multi-agent reinforcement learning framework is as follows:

[0028] Step S3.1: Initialize all static environment parameters and agent network parameters;

[0029] Step S3.2: Initialize dynamic environment parameters and start a new round, with time slot t=0;

[0030] Step S3.3: The intelligent agent of the CU-side resource allocation module observes the environmental status. Allocate resource blocks to users and share the resource allocation results. Distribute the data to each AP agent on the DU side;

[0031] Step S3.3: Each AP agent, based on the allocation information received from the CU side and the local observation status of each AP, Select the optimal beam for each user Perform data transmission;

[0032] Step S3.4: After each agent performs an action, calculate the reward based on the feedback from the environment. and cost incentives and obtain the next state. ;

[0033] Step S3.5, CU-side user agent storage , , , , >To their respective experience replay caches, DU-side AP agent storage< , , , >Redirect to their respective experience replay caches;

[0034] In step S3.6, each agent in both layers samples data from the experience replay buffer, calculates the gradient, and updates the parameters;

[0035] Step S3.7: Update the target network parameters of each agent at fixed time slots;

[0036] Step S3.8: If the maximum number of time slots for this round is reached, continue execution; otherwise, jump to S3.3.

[0037] In step S3.9, if the maximum number of rounds is reached, the training ends; otherwise, go to S3.2 and continue with a new round of training.

[0038] The deterministic delay non-cellular resource allocation method based on hierarchical distributed secure reinforcement learning of the present invention comprehensively considers the joint time-frequency resource allocation and beam optimization problem of non-cellular networks under delay requirement scenarios. It significantly reduces the computational complexity through a hierarchical distributed intelligent decision-making mechanism, and achieves efficient resource scheduling and dynamic optimization of delay-sensitive services by introducing a secure reinforcement learning mechanism. While meeting strict data packet delay constraints, it achieves higher spectrum efficiency and can adapt to dynamically changing wireless environments. Attached Figure Description

[0039] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0040] Figure 1A flowchart of a deterministic latency non-cellular resource allocation method based on hierarchical distributed security reinforcement learning according to an embodiment of the present invention;

[0041] Figure 2 This is a flowchart of step S2 in an embodiment of the present invention;

[0042] Figure 3 This is a schematic diagram of the hierarchical distributed multi-agent reinforcement learning algorithm according to an embodiment of the present invention;

[0043] Figure 4 This is a flowchart of step S3 in an embodiment of the present invention;

[0044] Figure 5 This is a convergence graph of the algorithm in an embodiment of the present invention;

[0045] Figure 6 This is a comparison diagram of latency protection in embodiments of the present invention;

[0046] Figure 7 This is a comparison chart of spectral efficiency optimization in embodiments of the present invention. Detailed Implementation

[0047] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0048] Figure 1 The flowchart of the deterministic latency non-cellular resource allocation method based on hierarchical distributed security reinforcement learning provided by the present invention specifically includes the following steps:

[0049] S1 addresses the deterministic latency constraints in non-cellular network scenarios by constructing a multi-point collaborative joint time-frequency resource allocation and beam optimization problem for non-cellular downlink communication.

[0050] In step S1, the cellular-free network scenario includes a central processing unit deployed on the CU side, distributed access points (APs) deployed on the DU side, and several users. Considering an OFDM downlink system, B access points (APs) each have M antennas, collaboratively serving U single-antenna users. The system bandwidth is divided into K subbands, each with a bandwidth of [missing information]. Each subband is the smallest schedulable frequency unit, which may contain a resource block (RB) consisting of C consecutive subcarriers. Each time slot contains N OFDM symbols, meaning the time slot length is... Therefore, the Minimum Scheduling Time-Frequency Unit (TFU) contains C*N resource elements (REs). These are respectively... , and This represents the set of access points, users, and subbands. An equal power allocation scheme is used, i.e. , Total transmit power for each access point. For the same user. All arriving data packets have the same deadline. Data packets in each user's buffer should be scheduled in a first-in, first-out (FIFO) manner. User data packets must be transmitted before the deadline, otherwise they will be lost. The constraints of the optimization problem model include the maximum delay violation probability constraint for users; the objective function of the optimization problem model is to maximize spectral efficiency; the optimization variables are resource block allocation and the beam selected by each AP for the user. In step S1, the optimization problem is specifically as follows:

[0051]

[0052] in, Indicates user exist Is the time in the sub-band? If scheduled, then The value is 1 if it is set to 1, and 0 otherwise. For DFT codebook set, Indicates AP For users Selected codeword index; Indicates user In the time slot The achievable rate at that time The bandwidth of each subband, where The parentheses (·) represent an indicator function that takes the value 1 if the condition within the parentheses is true, and 0 otherwise. (Constraints here) This indicates that the probability of user data packet delay default does not exceed ,constraint Indicates user In the time slot It was scheduled to a sub-band.

[0053] With the surge in user numbers and the diversification of service demands, the limited availability of spectrum resources has become a key factor restricting the development of wireless communication. Maximizing spectrum efficiency aims to transmit more information within limited spectrum resources. Optimizing spectrum efficiency requires not only focusing on the amount of information transmitted per unit bandwidth but also considering the real-time requirements of latency-sensitive services. Through dynamic resource scheduling, while improving spectrum utilization, it is ensured that data packets for critical services are transmitted within latency requirements, thereby improving overall communication efficiency and enhancing communication quality and reliability during transmission.

[0054] S2 constructs a time-delay-aware hierarchical distributed multi-agent reinforcement learning algorithm framework, decomposing the joint optimization problem into upper and lower sub-tasks. The upper CU side uses security reinforcement learning to solve resource block allocation. After determining the resource allocation result, each AP agent in the lower DU side completes beam selection.

[0055] In embodiments of the present invention, such as Figure 2 As shown, step S2 specifically includes:

[0056] S21. Based on the original optimization problem, the original problem is decomposed into two interrelated sub-problems, and a hierarchical architecture is designed to solve the time-frequency resource allocation and beam selection problems respectively.

[0057] S22 models the time-frequency resource allocation problem related to time delay as a constrained Markov decision process (CMDP) and the beam selection problem as a Markov decision process (MDP).

[0058] S23. For CMDP, a secure reinforcement learning algorithm is used to solve the time-frequency resource allocation problem, and the allocation results are transmitted to each AP.

[0059] S24, each AP agent selects the optimal beam for the user based on the resource block allocation result using a multi-agent reinforcement learning algorithm.

[0060] In step S21, to efficiently solve the joint optimization problem, a Joint Resource Allocation and Beam Selection (JRABS) scheme is proposed for the problem structure. The core idea is to decompose the original joint optimization into two interrelated sub-tasks and handle their coupling relationship through a collaborative learning framework. Specifically, a hierarchical distributed architecture is constructed: the upper-layer central processing unit (CU) is responsible for resource block allocation, and the lower-layer distributed access unit (DU) completes the beam selection task. This decomposition strategy fully preserves the resource block-beam coupling characteristics of the original problem, while significantly reducing the solution complexity through modular design.

[0061] In step S22, to address the critical impact of latency on resource allocation, a latency-aware cost function is introduced, and the resource allocation problem is modeled as a constrained Markov decision process. A Markov decision process includes a state space, action space, reward function, and discount factor. The constrained Markov decision process adds a cost function to the Markov decision process to ensure that the strategy meets strict latency requirements while optimizing the objective (such as spectral efficiency).

[0062] The state space of each agent in the CU-side resource allocation module is defined as the joint state of all users, i.e. ,in, Representing the status of each user, by composition; Depend on Success means user Channel state information for all APs, Indicates AP With users In sub-band Channel state information; Indicates user The length of the buffer queue, Indicates user The number of data packets dropped at the previous time point. Indicates user The maximum tolerable latency for the current buffer queue header packets; actions of each user agent. Instructing users In the time slot The assigned sub-band index; the reward is defined as the time slot. Instantaneous spectral efficiency:

[0063]

[0064] The cost function represents the extent to which all users violate latency requirements. Therefore, each time slot... The immediate cost at the end is expressed as:

[0065]

[0066] Probability of delay violation It is possible Calculation, where and They represent The number of packets dropped within a time slot and the total number of packets processed.

[0067] In wireless communication systems, beam selection directly impacts signal quality, interference management, and spectral efficiency. Since resource block scheduling, which affects latency, is already completed on the CU side, the beam selection problem, aimed at reducing interference between users, does not involve strict latency constraints. Therefore, it is modeled as a standard Markov decision process (MDP) to improve the user's signal-to-interference-plus-noise ratio (SINR).

[0068] The state space of each agent in the DU side beam selection module Defined as AP In time slots between each user Equivalent channel gain at time, AP In time The encoding index of the beamcodebook selected for each user, and the system in time. and time The number of RBs used, and the amount of data received by each user from the access point AP. Total interference power of signals sent to other users and AP Between each user in time Select channel state information on the selected resource block; the global state should reveal information about all users on each access point; therefore, integrate the local state of each agent into the global state; each AP agent's action is... ,in AP represents time t Assigned to user The codeword index; consistent with the reward defined by the resource allocation module, the instant reward is defined as the system's reward in the time slot. Spectral efficiency.

[0069] In step S23, the optimal strategy for solving the CMDP problem using security reinforcement learning is designed, including:

[0070] S231, applying the Lagrange relaxation method, constructs the Lagrange function, transforming the CMDP problem into an equivalent unconstrained optimization problem:

[0071]

[0072] in, For Lagrange functions, For Lagrange multipliers, Indicates by the parameter The following strategy This indicates a long-term discount reward. This indicates a long-term discount.

[0073] S232 proposes a secure reinforcement learning algorithm by incorporating the Lagrangian method into Dueling DQN to learn the optimal policy for the CMDP problem. In Dueling DQN, the Q-network output is decoupled into a state-value function and an action-advantage function:

[0074]

[0075] Among them, subscript Represents the user agent index. These are the parameters of the agent's general network, state-value network, and advantage network, respectively. Represents the state-action value function. Represents the dominance function. This represents the state-value function. By incorporating the Lagrange multiplier method, the Lagrange multipliers are updated during training. And based on the latest update Update the weights of the Dueling DQN network at each time step. ,award and cost All data are stored in playback memory. The proposed model needs to approximate the target Q-value, which is calculated in the following way:

[0076]

[0077] in, Indicates in time slot For users The target Q value is constructed. Indicates the discount factor. Indicates user The state in the next moment, here This indicates the Q-value estimated using the target Q-network. Indicates the state at the next moment. Below, among all possible actions in the action space, making The biggest movement, This represents the parameters of the target network in Dueling DQN at the current time. These represent the user's intelligent agent at the current moment. The parameters of the state-value network and the advantage network can be estimated by the model after training. All parameters are updated using gradient descent to minimize the mean squared error. The details are as follows:

[0078]

[0079] in, This represents the learning rate of the Dueling DQN network. The optimal value of the Lagrange multiplier depends on the time delay constraint and can be learned online using the stochastic subgradient method.

[0080]

[0081] in Indicates that the Lagrange multipliers Keep in [0, Projection operator within the interval, This represents the update step size of the Lagrange multipliers, set during the agent's learning phase. Greater than Make the strategy weight The update speed is faster than that of the Lagrange multiplier. The update speed is faster. The algorithm periodically copies the parameters of the Dueling DQN evaluation network to the Dueling DQN target network.

[0082] In step S24, considering the distributed deployment of access points and to avoid excessive signal overhead, the QPLEX algorithm based on the CTDE framework is adopted for the beam selection MDP problem. Inspired by Dueling DQN, the QPLEX algorithm decomposes the state-action value Q into state value V and advantage function A, and introduces an advantage-based individual-global maximum (IGM) value decomposition structure, making it particularly suitable for scenarios that require a combination of autonomous agent operation and global coordination.

[0083] like Figure 3 As shown, QPLEX consists of three types of networks: a duel-hybrid network, a transformation network, and a value network for each agent. The value network input for each agent includes its current observation and previous actions. The transformation network receives the value of a single agent and, using weights generated from global state information, calculates a weighted sum of the local value function and the advantage function. The duel-hybrid network module integrates the outputs of the transformation networks into a joint reward.

[0084]

[0085] in, This represents the global trajectory, i.e., the joint actions of all agents – the observation history. Indicates the AP agent index. Let represent the joint action vector composed of all agents, where For the first The actions of each AP agent; Indicates multi-agent trajectory Execute joint actions The global action value function at that time. This represents the global state value function, corresponding to the joint action. The global advantage function, Indicates the first An intelligent agent in the trajectory Select local action The individual action value function, No. Individual advantage function of each agent The positive importance weights are calculated using global trajectories and joint actions, aiming to ensure consistency between the joint advantage and the individual advantage IGM. During centralized training, a batch of empirical data is uniformly drawn from the memory bank, and the parameters are updated using gradient descent. ( (Parameters for the QPLEX algorithm) to minimize the mean squared error:

[0086]

[0087] in, This represents the temporal difference (TD) loss function. The reward value at the current moment. Indicates the discount factor. It is the parameter set of the target network. This represents the global trajectory at the next moment. This indicates a connection to history in the next moment. Below, the global action value output by the target network is increased. The largest joint operation.

[0088] S3 trains the parameters of the hierarchical distributed multi-agent reinforcement learning algorithm to obtain the optimal policy.

[0089] In embodiments of the present invention, such as Figure 4 As shown, step S3 specifically includes:

[0090] Step S3.1: Initialize all static environment parameters and agent network parameters;

[0091] Step S3.2: Initialize dynamic environment parameters and start a new round, with time slot t=0;

[0092] Step S3.3: The intelligent agent of the CU-side resource allocation module observes the environmental status. Allocate resource blocks to users and share the resource allocation results. Distribute the data to each AP agent on the DU side;

[0093] Step S3.3: Each AP agent, based on the allocation information received from the CU side and the local observation status of each AP, Select the optimal beam for each user Perform data transmission;

[0094] Step S3.4: After each agent performs an action, calculate the reward based on the feedback from the environment. and cost incentives and obtain the next state. ;

[0095] Step S3.5, CU-side user agent storage , , , , >To their respective experience replay caches, DU-side AP agent storage< , , , >Redirect to their respective experience replay caches;

[0096] In step S3.6, each agent in both layers samples data from the experience replay buffer, calculates the gradient, and updates the parameters;

[0097] Step S3.7: Update the target network parameters of each agent at fixed time slots;

[0098] Step S3.8: If the maximum number of time slots for this round is reached, continue execution; otherwise, jump to S3.3.

[0099] In step S3.9, if the maximum number of rounds is reached, the training ends; otherwise, go to S3.2 and continue with a new round of training.

[0100] In a further embodiment, based on the known system deployment and considering user latency constraints, an optimization problem model for downlink joint time-frequency resource allocation and beam selection in a cellular system is constructed. Taking into account the distributed processing characteristics of cellular systems, the downlink deterministic latency communication problem is decomposed into a two-layer solution: a latency-constrained resource allocation problem is solved at the upper-layer CU, and a beam selection problem is solved at each lower-layer AP, ultimately yielding the solution to the original problem. Simulation results show that the proposed method achieves high spectral efficiency while satisfying strict data packet latency constraints.

[0101] To verify the effectiveness of this invention, the following simulation comparison experiments were conducted:

[0102] Set up a 600m x 600m non-cellular scenario. Configure 4 access points (APs), each equipped with 4 antennas, located at (0 m, 300 m), (300 m, 0 m), (300 m, 600 m), and (600 m, 600 m). Additionally, 3 users are randomly distributed within the area, with latency constraints of 4 ms, 3 ms, and 3.5 ms, respectively. The transmit power of each access point is [not specified]. = 20 dBm, scheduling time slot length set to = 500. CF networks operate at carrier frequency = 3.5 GHz. The number of sub-bands is set to 4, and the bandwidth of each sub-band is... = 360 kHz. Each subband contains C = 12 consecutive subcarriers, and a single time slot contains N = 14 OFDM symbols. Set to -129 dBm. Packet arrival rates for each user are set to 88%, 86%, and 87%, respectively. The latency default probability for each user is limited to [value missing]. When assessing the delay default rate, each simulation should consist of at least 10,000 time slots, and the result should be the average of 50 simulations.

[0103] In this embodiment, both the Dueling DQN and QPLEX networks use the ReLU function as the nonlinear activation unit. The network learning rates are set to... and The discount factors were set to 0.7 and 0 respectively; the memory pool capacity was 1000 for both, and the mini-batch samples collected during training were set to 128 and 256 respectively.

[0104] First, to verify the rationality of the proposed Joint Resource Allocation and Beam Selection Optimization Scheme (JRABS), it is compared with three benchmark schemes: (1) Random Optimization: All agents randomly select actions; (2) RB Allocation Optimization: Only the resource block allocation module is optimized; (3) Beam Selection Optimization: Only the beam selection module is optimized. Figure 5 As shown, the proposed joint scheme outperforms all benchmark schemes, achieving a nearly 50% improvement in system performance compared to the stochastic optimization scheme. These results demonstrate that the joint optimization method can effectively improve the performance of non-cellular downlink systems.

[0105] To thoroughly evaluate the latency guarantee performance of the proposed solution, it is compared with two benchmark solutions:

[0106] (1) A modified Maximum Weighted Delay First (MLWDF) algorithm; this scheme is derived from the literature [C. Mohanram and S. Bhashyam, “Joint subcarrier and power allocation in channel-aware queue-aware scheduling for multiuser OFDM,” IEEE Transactions on Wireless Communications, vol. 6, no. 9, pp. 3208–3213, 2007.]

[0107] (2) Proportional Fairness (PF) Algorithm:

[0108] from Figure 6 As can be seen, under a packet size of 32 bits (corresponding to a packet arrival rate of 16.45 bit / s / Hz), the performance of the proposed algorithm compared with the MLWDF and PF algorithms in terms of latency default rate is shown. It can be seen that, under the condition of a system load rate (packet arrival rate divided by maximum achievable spectral efficiency) of 92.96%, the proposed algorithm outperforms the two benchmark algorithms, satisfying the requirements. The latency default rate requirement. Specifically, the algorithm achieves an average user latency default rate of 0.0000338 (…). The default rate of the MLWDF algorithm is 0.0004768. These results demonstrate that, compared to the MLWDF algorithm, the proposed algorithm reduces the delay default rate by an order of magnitude and is significantly superior to the PF algorithm. This advantage stems from the fact that the PF algorithm only considers the user channel and average throughput, while MLWDF, although considering user data packet delay, fails to fully utilize the data packet timeout mechanism.

[0109] Furthermore, the spectral efficiency under different loads was compared, and the simulation results are as follows: Figure 7 As shown, the spectral efficiency initially increases with system load under low throughput conditions. However, as the arrival rate further increases, the actual spectral efficiency of the system tends to stabilize and approach its maximum achievable value. It can be seen that, under the same load, the proposed solution achieves higher spectral efficiency.

[0110] This invention proposes a deterministic delay-aware, distributed, deep reinforcement learning (DRL) framework for resource allocation in acellular networks, addressing the joint optimization problem of deterministic delay resource allocation and beam selection. This framework effectively satisfies delay constraints through secure reinforcement learning algorithms, enabling distributed decision-making and improving network scalability.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and do not constitute a limitation. Modifications, variations, or equivalent substitutions made by those skilled in the art without departing from the spirit and essence of the present invention, and subject to the scope defined by the claims, should all be included within the protection scope of the present invention.

Claims

1. A deterministic, latency-delay-free resource allocation method based on hierarchical distributed security reinforcement learning, characterized in that, Includes the following steps: Step S1: For deterministic latency constraints in non-cellular network scenarios, construct a joint time-frequency resource allocation and beam optimization problem for multi-point collaborative non-cellular downlink communication. Step S2: Construct a time-delay-aware hierarchical distributed multi-agent reinforcement learning algorithm framework, decompose the joint time-frequency resource allocation and beam optimization problem into upper and lower sub-tasks. The upper-layer central processing unit (CU) uses security reinforcement learning to solve the resource block allocation. After determining the resource allocation result, each access point (AP) agent on the lower-layer distributed unit (DU) completes beam selection. Step S3: Train the parameters of the constructed time-delay-aware hierarchical distributed multi-agent reinforcement learning algorithm framework to obtain the optimal policy.

2. The deterministic latency non-cellular resource allocation method based on hierarchical distributed security reinforcement learning according to claim 1, characterized in that, The cellular-free network scenario includes a central processing unit deployed on the CU side, access points (APs) distributed on the DU side, and several users. Considering an OFDM downlink system, B APs each have M antennas, collaboratively serving U single-antenna users. The system bandwidth is divided into K subbands, each with a bandwidth of [missing information]. Each subband is the smallest schedulable frequency unit, containing a resource block RB, which consists of C consecutive subcarriers. Each time slot contains N OFDM symbols, meaning the time slot length is... Therefore, the minimum scheduling time-frequency unit (TFU) contains C*N resource elements (RE); each is represented by a different resource element (RE). , and This represents the set of access points (APs), users, and subbands; using an equal power allocation scheme, the power allocated by AP b to user u is... , Total transmit power for each access point; for the same user All arriving data packets have the same deadline. Data packets in each user buffer should be scheduled in a first-in, first-out (FIFO) manner; user data packets must be transmitted before the deadline, otherwise they will be lost. The constraints of the joint time-frequency resource allocation and beam optimization problem for multi-point cooperative downlink communication include the user's maximum delay violation probability constraint; the objective function of the optimization problem is to maximize spectral efficiency; the optimization variables are resource block allocation and the beam selected by each AP for the user. In step S1, the joint time-frequency resource allocation and beam optimization problem for multi-point cooperative downlink communication is specifically as follows: ; in, Indicates user exist Is the time in the sub-band? If scheduled, then The value is 1 if it is set to 1, and 0 otherwise. For DFT codebook set, Indicates AP For users Selected codeword index; Indicates user In the time slot The achievable rate at that time The bandwidth of each subband, where (·) indicates an indicator function, which takes the value 1 when the condition inside the parentheses is true, and takes the value 0 otherwise. Represents a time set; (·) represents a probability function, that is, the probability of the event within the parentheses occurring; Indicates user exist The time reached The latency experienced by each data packet; Indicates user The packet delay violates the probability requirement; constraints This indicates that the probability of a user data packet delay violation does not exceed [a certain value]. ;constraint express Use 0 and 1 to indicate whether to schedule; constraints Indicates AP For users The selected codeword index comes from the DFT codebook set. constraint Indicates user In the time slot It was scheduled to a sub-band.

3. The deterministic delay-free cellular resource allocation method based on hierarchical distributed security reinforcement learning according to claim 2, characterized in that, Step S2 specifically includes the following steps: Step S21: Based on the original optimization problem, the original problem is decomposed into two interrelated sub-problems, and a hierarchical architecture is designed to solve the time-frequency resource allocation and beam selection problems respectively. Step S22: The time-frequency resource allocation problem related to time delay is modeled as a constrained Markov decision process (CMDP), and the beam selection problem is modeled as a Markov decision process (MDP). Step S23: For the Constrained Markov Decision Process (CMDP), a secure reinforcement learning algorithm is used to solve the time-frequency resource allocation problem, and the allocation results are transmitted to each AP. In step S24, each AP agent selects the optimal beam for the user based on the resource block allocation result using a multi-agent reinforcement learning algorithm.

4. The deterministic delay-free cellular resource allocation method based on hierarchical distributed security reinforcement learning according to claim 3, characterized in that, The layered architecture designed in step S21 is as follows: The hierarchical distributed architecture consists of an upper-layer central processing unit (CU) and a lower-layer distributed unit (DU), which respectively complete resource allocation and beam selection tasks. The upper-layer CU deploys several user agents, and each access point (AP) in the lower layer deploys one agent. The CU is responsible for centrally acquiring and processing information and allocating resource blocks to users. The DU then selects the optimal beam for transmission to users based on the allocation result.

5. The deterministic delay-free cellular resource allocation method based on hierarchical distributed security reinforcement learning according to claim 3, characterized in that, In step S22, the constrained Markov decision process (CMDP) and the Markov decision process (MDP) are specifically modeled as follows: The state space of each agent in the CU-side resource allocation module is defined as the joint state of all users, i.e. ,in, Representing the status of each user, by composition; Depend on Composition, representing users Channel state information of all APs, Indicates AP With users In sub-band Channel state information; Indicates user The length of the buffer queue, Indicates user The number of data packets dropped at the previous time point. Indicates user The maximum tolerable latency for the current buffer queue header packets; actions of each user agent. Instructing users In the time slot The assigned sub-band index; rewards Defined as a time slot Real-time spectral efficiency; cost function, which is the degree to which all users violate latency requirements, per time slot. The immediate cost at the end is expressed as , Indicates user The packet delay violates the probability requirement; The state space of each agent in the DU side beam selection module Defined as AP Between each user in time slots Equivalent channel gain at time, AP In time The encoding index of the beamcodebook selected for each user, and the system in time. and time The number of RBs used, and the amount of data received by each user from the access point AP. Total interference power of signals sent to other users and AP Between each user in time Select channel state information on the selected resource block; the global state should reveal information about all users on each access point; therefore, integrate the local state of each agent into the global state; each AP agent's action is... ,in AP represents time t Assigned to user The codeword index; consistent with the reward defined by the resource allocation module, instant reward. Defined as the system in time slot Spectral efficiency.

6. The deterministic delay-free cellular resource allocation method based on hierarchical distributed security reinforcement learning according to claim 3, characterized in that, In step S23, the security reinforcement learning algorithm used on the CU side is specifically designed as follows: Step S231: Apply the Lagrange relaxation method to construct the Lagrange function, transforming the Constrained Markov Decision Process (CMDP) problem into an equivalent unconstrained optimization problem. Step S232: The Lagrange optimization method is combined with the Dueling DQN reinforcement learning algorithm to alternately update the Lagrange multipliers and policy network parameters to obtain the optimal resource allocation.

7. The deterministic delay-free cellular resource allocation method based on hierarchical distributed security reinforcement learning according to claim 2, characterized in that, In step S24, the QPLEX multi-agent reinforcement learning algorithm is used on the DU side: On the DU side, each access point is equipped with a QPLEX-based agent, which is trained under a centralized training and distributed execution CTDE framework; during the execution phase, each agent selects the optimal beam for the user based on its local state.

8. The deterministic delay-free cellular resource allocation method based on hierarchical distributed security reinforcement learning according to claim 3, characterized in that, In step S3, the training process of the time-delay-aware hierarchical distributed multi-agent reinforcement learning framework is as follows: Step S3.1: Initialize all static environment parameters and agent network parameters; Step S3.2: Initialize dynamic environment parameters and start a new round, with time slot t=0; Step S3.3: The intelligent agent of the CU-side resource allocation module observes the environmental status. Allocate resource blocks to users and share the resource allocation results. Distribute the data to each AP agent on the DU side; Step S3.3: Each AP agent, based on the allocation information received from the CU side and the local observation status of each AP, Select the optimal beam for each user Perform data transmission; Step S3.4: After each agent performs an action, calculate the reward based on the feedback from the environment. and cost incentives and obtain the next state. ; Step S3.5, CU-side user agent storage , , , , >To their respective experience replay caches, DU-side AP agent storage< , , , >Redirect to their respective experience replay caches; In step S3.6, each agent in both layers samples data from the experience replay buffer, calculates the gradient, and updates the parameters; Step S3.7: Update the target network parameters of each agent at fixed time slots; Step S3.8: If the maximum number of time slots for this round is reached, continue execution; otherwise, jump to S3.

3. In step S3.9, if the maximum number of rounds is reached, the training ends; otherwise, go to S3.2 and continue with a new round of training.