Cellular-free network resource optimization method for attention-driven hierarchical multi-agent reinforcement learning

Through the attention-driven hierarchical multi-agent reinforcement learning algorithm, the problem of inefficient resource allocation in cellular systems is solved, collaborative decision-making and global optimization between access points are realized, and system performance is improved.

CN120264429APending Publication Date: 2025-07-04TONGJI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510616535.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Cellulous systems are inefficient in resource allocation in high-dimensional state space and dynamic scenarios, insufficient collaboration between agents, and easy to fall into local optimality, and traditional methods are difficult to effectively adapt to complex multi-user and multi-service concurrent scenarios.

Method used

The attention-driven hierarchical multi-agent reinforcement learning algorithm (HAD-MARL) is used to split the optimization problem into two layers: the upper controller generates the access point and the user pairing matrix, the lower controller performs fine-grained resource allocation, and introduces an attention mechanism to make the access point selectively pay attention to the information of other access points to enhance coordination.

Benefits of technology

It significantly improves resource allocation efficiency, achieves global optimization through hierarchical architecture and attention mechanism, avoids local optimization, and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264429A_ABST
    Figure CN120264429A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cellular-free large-scale multiple-input multiple-output (CF-mMIMO) resource allocation and artificial intelligence, in particular to a method for optimizing cellular-free network resource allocation efficiency by using an attention-driven hierarchical multi-agent enhancement algorithm. Aiming at the problems of low resource allocation efficiency and insufficient cooperation among agents in a high-dimensional state space and a dynamic scene, the invention provides an attention-driven hierarchical multi-agent reinforcement learning algorithm (HAD-MARL), a source optimization problem is divided into two layers, an upper layer controller grasps an overall training target of a system from a global perspective, and a lower layer controller grasps an overall training target of the system from a global perspective; the access point and user pairing matrix is output and transmitted to a lower-layer controller, and all agents in the lower-layer controller cooperate to complete a finer-grained resource allocation task. By introducing an attention mechanism, each access point can selectively pay attention to information from other access points in a decision-making process, so that the strategy is promoted to develop towards a better overall direction by enhancing collaboration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of cell-free massive multiple-input multiple-output (CF-mMIMO) resource allocation and artificial intelligence, and particularly to a method for optimizing the resource allocation efficiency of a cell-free network by using an attention-driven hierarchical multi-agent reinforcement algorithm. Background Art

[0002] With the evolution of wireless communication systems, cell-free massive multiple-input multiple-output has become a key candidate architecture for future 6G networks due to its excellent spectral efficiency and user experience. By deploying a large number of distributed access points in a wide area and enabling their cooperative transmission, cell-free systems effectively alleviate the long-standing intra-cell interference problem in traditional cellular networks and significantly improve the performance of cell-edge users. In the actual 6G application scenarios with the characteristics of multi-user and multi-service concurrent access, wireless communication systems need to simultaneously meet heterogeneous service requirements such as high throughput, precise positioning and sensing, and ultra-high reliable connection. This poses new challenges for resource allocation - especially in cell-free systems supporting integrated sensing and communication (ISAC), due to the combined effects of high-dimensional state space, dynamic user behavior, and diverse service requirements, this problem has inherent complexity. Traditional optimization methods based on static models or heuristic rules are difficult to effectively adapt to such highly dynamic large-scale networking environments.

[0003] Reinforcement learning is a machine learning method in which an agent interacts with the environment and learns the optimal policy through trial and error to maximize long-term rewards. Among them, multi-agent reinforcement learning (MARL) can effectively handle complex and realistic application problems through distributed intelligent decision-making and collaborative optimization, and has now received extensive attention. For example, Ammar et al. verified in the paper "Downlink Resource Allocation in Multiuser Cell-Free MIMO Networks with User-Centric Clustering" that the multi-agent algorithm can achieve higher energy efficiency than traditional algorithms in a cache-enabled cell-free network. However, the high-dimensional search space and sparse reward problems caused by complex scenarios are major factors restricting the proposed algorithm. Liu et al. combined the weighted K-nearest neighbor algorithm with MARL in the paper "Cooperative Multi-Target Positioning for Cell-Free Massive MIMO with Multi-Agent Reinforcement Learning" to solve the balance problem between the complexity of high-dimensional signal processing and positioning accuracy in a cell-free massive MIMO system. However, this algorithm treats each access point as an independent individual during decision-making and does not consider the cooperation between access points, resulting in the optimization process being easily trapped in a local optimum. Luo et al. proposed a two-layer MARL network architecture in the paper "Deep Reinforcement Learning-Based Sum Rate Fairness Trade-off for Cell-Free mMIMO" to improve the decision-making efficiency. This algorithm requires each agent to interact with all access points to train the value network. Although it initially reflects the interaction between access points, not all access points have a cooperative relationship in the actual scenario, and the addition of redundant information restricts the scalability of this method in scenarios where the number of access points changes dynamically.

[0004] In summary, there are problems of low resource allocation efficiency and insufficient cooperation among agents in the dynamic scenarios of existing cell-free systems. In a cell-free system, users can receive services from multiple access points, and the high-dimensional state space caused by the concurrency of multiple users and multiple services makes resource allocation extremely difficult. On the other hand, existing methods treat access points, that is, agents, as independent individuals for decision-making, which makes the optimization easily fall into a local optimum and does not give full play to the advantages of distributed decision-making in cell-free networks. Summary of the Invention

[0005] In view of the problems of low resource allocation efficiency in the above-mentioned high-dimensional state space and dynamic scenarios, and insufficient cooperation among agents, the present invention proposes an attention-driven hierarchical multi-agent reinforcement learning algorithm (Hierarchical Attention-Driven Multi-Agent Reinforcement Learning, HAD-MARL), which fully exploits the cooperation potential among multi-agents, thereby significantly improving the resource allocation efficiency in complex environments. Specifically, the HAD-MARL method splits the source optimization problem into two layers. The upper-layer controller grasps the overall training objective of the system from a global perspective, outputs the access point-user pairing matrix, and transmits it to the lower-layer controller. Each agent in the lower-layer controller then collaborates to complete a more fine-grained resource allocation task. To solve the sub-optimal local decision problem caused by insufficient information sharing among access points, an attention mechanism is introduced, enabling each access point to selectively focus on information from other access points during the decision-making process, thereby promoting the policy to develop towards an overall better direction through enhanced cooperation.

[0006] Technical Solution

[0007] A method for optimizing the resources of a cell-free network based on attention-driven hierarchical multi-agent reinforcement learning includes the following steps:

[0008] S1. Construct a system model for a cell-free integrated communication and sensing multi-user multi-service concurrent scenario;

[0009] S2. Construct communication and sensing metrics based on Shannon's theorem and the Cramer-Rao bound respectively, and establish an optimization problem;

[0010] S3. Design of the HAD-MARL algorithm, including the design of the hierarchical architecture and the attention mechanism;

[0011] S4. Train the model according to the HAD-MRAL algorithm and update the parameters.

[0012] Beneficial Effects

[0013] The positive and progressive effects of the present invention are as follows:

[0014] 1. Consider a cell-free massive MIMO system with an integrated communication and sensing multi-user multi-service concurrent scenario, construct communication and sensing metrics and related optimization problems, and solve them using reinforcement learning methods.

[0015] 2. Use a hierarchical architecture to decompose the optimization problem into multiple sub-problems that are easier to solve, effectively addressing the deficiencies of high complexity, high-dimensional state space, and low efficiency of the original problem. The upper-layer controller outputs the access point-user pairing instructions, and the lower-layer controller then makes a fine-grained resource allocation strategy based on this instruction.

[0016] 3. To address the problem of insufficient information sharing between access points, which is prone to falling into the overall local optimum, an attention mechanism is introduced. When each access point makes a decision, it selectively observes information from other access points according to the attention weights to complete tasks in a collaborative manner. Description of the Drawings

[0017] Figure 1 is a system block diagram of the scenario considered in the present invention;

[0018] Figure 2 is a schematic diagram of the algorithm proposed by the present invention;

[0019] Figure 3 is a schematic diagram of the attention-driven actor-critic multi-agent reinforcement learning algorithm proposed by the present invention;

[0020] Figure 4 is a processing flowchart of step S4 of the method of the present invention;

[0021] Figure 5 is an experimental comparison diagram of the embodiments of the present invention;

[0022] Figure 6 Attention visualization diagram of the embodiments of the present invention (where (a) is the access point-user equipment pairing relationship diagram, (b) is the distribution diagram of the attention weights of each access point to the other access points, and (c) is the distribution diagram of the entropy values of the attention heads inside each access point). Detailed Embodiments

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] A cell-free network resource optimization method for attention-driven hierarchical multi-agent reinforcement learning includes the following steps:

[0025] S1. Construct a system model for a cell-free communication and sensing integrated multi-user and multi-service concurrent scenario;

[0026] S2. Construct communication and sensing metrics respectively based on Shannon's theorem and the Cramer-Rao bound, and establish an optimization problem;

[0027] S3. Design the HAD-MARL algorithm, including the design of a hierarchical architecture and the design of an attention mechanism;

[0028] S4. Train the model according to the HAD-MRAL algorithm and update the parameters.

[0029] Specifically, in step S1, consider the downlink communication scenario of a cell-free system, as Figure 1 shown. The system includes: a central controller, a number of access points, and a number of single-antenna users.

[0030] Among them, the access point set Each access point is equipped with N t transmit antennas and N r receive antennas to serve the single-antenna user equipment set Assume that each access point can adaptively send radar signals or communication signals according to the heterogeneous service requirements of user equipment.

[0031] The service types of users are divided into three categories: communication-only users, sensing-only users, and communication-sensing integrated users.

[0032] All access points are connected to the central controller through high-speed fronthaul links to achieve real-time information sharing. On the other hand, assume that the system is a time-division duplex system, using the same spectrum resources, and its uplink and downlink channels have reciprocity, which means that the channel state information obtained by uplink pilot estimation can be directly used for downlink data transmission. For the convenience of subsequent modeling, define the set to represent users with communication service requirements, and the set to represent users with sensing service requirements. The sets and satisfy

[0033] In step S2, it is necessary to characterize the service metrics and thus establish an optimization problem.

[0034] S2.1. Construct the communication service metrics as follows:

[0035]

[0036] Among them, τ c represents the coherence time, τ p represents the pilot sequence length, b mk is the bandwidth allocated by access point k to user m, and η mk is the signal-to-noise ratio and is related to the allocated power.

[0037] S2.2. Construct the sensing service metrics as follows:

[0038]

[0039] Among them, CRB(τ jk ) and CRB(θ jkrepresent the Cramér-Rao bounds of the transmission delay and the angle of arrival respectively, and λ1 and λ2 are the weighting coefficients respectively.

[0040] S2.3. Construct the optimization problem:

[0041]

[0042] R min ≤R mk

[0043] ω min ≤ω jk

[0044]

[0045] where U(·) is the system utility function, and the optimization variable is the pairing matrix between the access point and the user, and 1 or 0 is used to represent whether the access point is connected to the user; the optimization variables b nk and P nk represent the bandwidth and power resources allocated by the access point to the user respectively. R min and ω min represent the lower bounds of the communication metric and the sensing metric respectively, w nk represents the beamforming matrix, and represent the minimum and maximum values of the transmission power of the k-th access point respectively, and represent the minimum and maximum values of the bandwidth resources allocated by the k-th access point to the n-th user respectively. B and P represent the total bandwidth and the total power respectively.

[0046] The above optimization problem is a non-convex problem, and as the number of access points and users increases, the dimension of the solution space will also increase rapidly, resulting in the inability of traditional reinforcement learning algorithms to solve it efficiently.

[0047] In step S3, based on the above proposed optimization problem, an algorithm design is carried out, and a hierarchical attention-driven multi-agent reinforcement learning algorithm (Hierarchical Attention-Driven Multi-Agent Reinforcement Learning, HAD-MARL) is proposed, including the design of the hierarchical architecture and the design of the attention mechanism. The algorithm block diagram is as Figure 2 shown.

[0048] S3.1. Design of the hierarchical architecture

[0049] The present invention proposes an innovative hierarchical architecture to disassemble problems. The hierarchical architecture includes a "commander controller" and an "employee controller". Among them, the upper-layer "commander controller" is deployed in the central controller, and the lower-layer "employee controller" is deployed at the access point. One agent is deployed at each access point, with a total of K agents. The "commander controller" is responsible for grasping the overall situation and generating the access point and user pairing matrix instruction; the "employee controller" then makes fine-grained power and bandwidth allocation strategies based on this instruction.

[0050] Specifically, the "commander controller" obtains the global environment state observation information and generates an instruction c t , that is, the access point and user pairing matrix, which is used to indicate whether there is a connection between access point k and user n. It should be noted that the "commander controller" updates the instruction every f time slots. Between two update cycles, the instruction remains unchanged.

[0051] The state space of the "commander controller" is defined as where g nk (t) is the estimated channel state information, including the user service demand; the instruction can be defined as The reward is defined as Finally, all experiences are stored in the experience replay buffer in the form of a tuple .

[0052] The "commander controller" uses a classic Dueling DQN network, which consists of three layers. The first two hidden layers each contain 256 neurons, and the output layer has a total of 128 neurons. One neuron is responsible for the output of the advantage function, and the remaining neurons are responsible for the output of the value function.

[0053] For the lower-layer "employee controller", after receiving the instruction c t from the upper layer, it distributes it to each agent. Each agent will output a local policy based on the attention mechanism.

[0054] For each agent, the state space is defined as The action space is defined as The reward function is defined as:

[0055]

[0056] where λ5 is the penalty coefficient, is the penalty function.

[0057] Then the total reward function can be expressed as

[0058] The experience tuple ​Stored in the k-th experience replay buffer inside

[0059] S3.2. Attention mechanism design (as Figure 3 shown)

[0060] The lower-level "employee controller" uses an attention-driven actor-critic multi-agent reinforcement learning algorithm, introducing the attention mechanism into the actor-critic multi-agent reinforcement learning algorithm to solve the problems of insufficient access point information sharing, insufficient cooperation, and easy to fall into local optimality in the traditional resource allocation method in the cell-free scenario.

[0061] The working steps of the attention mechanism include:

[0062] S3.2.1 The central attention mechanism collects the states and actions of each agent, and obtains the corresponding embedding vectors through a single-layer neural network: [e1,...,e k ,...,e K , where all the embedding vectors except the k-th agent are represented as e \k .

[0063] S3.2.2 For the attention head u, calculate the attention weights:

[0064]

[0065] Among them, W K and W Q are the linear mapping matrices corresponding to "Key" and "Query", D is the matrix dimension; the function Θ(·) is the softmax function.

[0066] S3.2.3 Calculate the attention value:

[0067]

[0068] Among them, W V is the linear mapping matrix corresponding to "Value"; the function L(·) is a non-linear activation function, and ReLU is used here.

[0069] S3.2.4 Generate the Q function based on the attention value:

[0070]

[0071] Among them, U represents the number of attention heads, represents the critic network of the k-th agent, which contains a three-layer multi-layer perceptron, as its parameter. It is worth noting that the Q function is an important parameter for the decision-making and update of the actor and critic. Here, it is not directly generated by a multi-layer perceptron as in the traditional method, but an attention mechanism is added. That is, the generation of the Q function is based on the selective observation of the state-action information of the remaining access points based on the attention weights, enabling the efficient circulation of information among various access points and reducing the possibility of falling into local optima.

[0072] Furthermore, for the upper-layer "commander controller", Dueling DQN is used, and the Q function is defined as:

[0073]

[0074] where and are the value function and the advantage function generated by the neural network in step S3.1 respectively. Thus, the optimization objective can be expressed as:

[0075]

[0076] where represents the standard mathematical expectation.

[0077] Thus, the gradient can be expressed as:

[0078]

[0079] Then the network parameters can be updated by the ADAM optimizer:

[0080]

[0081] where ξ h is the learning rate, τ << 1 is the soft update coefficient, m t and represent the first and second moments related to the bias respectively.

[0082] Furthermore, for the lower-layer "employee controller", an attention-driven actor-critic multi-agent algorithm is used.

[0083] The optimization objective of the actor can be defined as:

[0084]

[0085] The gradient is:

[0086]

[0087] Parameter update:

[0088]

[0089] where is the learning rate.

[0090] The optimization objective of the Critic can be defined as:

[0091]

[0092] The gradient is:

[0093]

[0094] Parameter update:

[0095]

[0096] φ l- ← τφ l + (1 - τ)φ l-

[0097] where ξ l is the learning rate and τ << 1 is the soft update coefficient.

[0098] In step S4, the model is trained according to the HAD-MRAL algorithm to update the parameters. The training process of the HAD-MRAL algorithm is as Figure 4 shown as follows:

[0099] S4.1. Initialize each environmental parameter and agent parameter φ h φ l and set φ h- = φ h φ l- = φ l ;

[0100] S4.2. Start a new episode. The upper-level "commander controller" observes the environmental state and generates an instruction c t ;

[0101] S4.3. The "commander controller" distributes the instruction to each agent in the lower-level "employee controller". Each agent generates a corresponding resource allocation strategy a k,t according to the environment and the instruction;

[0102] S4.4. Each agent executes the action a k,t and obtains the reward fed back by the environment and the next state

[0103] S4.5. Each agent stores the record of this attempt into the corresponding experience replay buffer

[0104] S4.6. Each agent samples from the experience replay buffer to calculate the gradient and update the parameter φ l , φ l- ;

[0105] S4.7. Determine whether the number of time slots is divisible by the update interval f. If so, the upper-layer "commander controller" calculates the obtained reward r t h ; If not, directly jump to S4.10.;

[0106] S4.8. The upper-layer controller stores the record of this attempt in the experience replay buffer

[0107] S4.9. The upper-layer controller samples from the experience replay buffer to calculate the gradient and update the parameter φ h , φ h- ;

[0108] S4.10. Determine whether the maximum number of time slots is reached. If so, continue to execute; if not, jump to S4.3;

[0109] S4.11. Determine whether the maximum number of episodes is reached. If so, end the training; if not, jump to S4.2 to continue the training of the next episode.

[0110] A cell-free network resource optimization method based on attention-driven hierarchical multi-agent reinforcement learning proposed by the present invention aims to solve the problems of low resource allocation efficiency and insufficient cooperation among access points in complex scenarios.

[0111] To fully illustrate the effectiveness of the present invention, comparative experiments and attention visualization were carried out.

[0112] Assume a 0.8km * 0.8km cell-free massive MIMO system, in which multiple access points cooperate to provide three types of services for user equipment: independent communication, independent sensing, and integrated communication and sensing services. The system adopts a fixed access point deployment architecture, and the user positions remain relatively static during the service period, while the user service requirements follow a random generation mechanism.

[0113] In this embodiment, Dueling DQN adopts a double hidden layer architecture and selects the ReLU function as the non-linear activation unit. The network learning rate is set to 1*10 -3 , and the discount factor is 0.95. In the low-layer attention-driven actor-critic multi-agent model, both the actor and critic networks adopt a single hidden layer structure (128 neurons), the embedding dimension is set to 256, and the learning rates are both 1*10-2 Meanwhile, the discount factor is kept at 0.95. The soft update coefficients of each module are all set to 0.01. There are 5 access points and 20 users in the system. The total number of time slots is fixed at 200, and the total number of iteration rounds is 150.

[0114] Four control models are adopted in the experiment:

[0115] 1) Hierarchical Multi-Agent Reinforcement Learning (HMARL): This solution is from the literature [F. Ye, J. Wang, J. Li, P. Zhu, D. Wang, and X. You, “Intelligent Hierarchical Network Slicing Based on Dynamic Multi-Connectivity in Cell-Free Distributed Massive MIMO Systems,” IEEE Transactions on Vehicular Technology, vol. 72, no. 9, pp. 11 855–11 870, 2023.]. It uses a hierarchical reinforcement learning architecture to optimize network slicing in a cell-free massive system. Its upper-layer controller uses a Double Deep Q-Network (Double DQN), and the lower-layer controller uses Multi-Agent Deep Deterministic Policy Gradient (MADDPG). The core difference from the present invention is that the lower-layer agents of this method operate completely independently, without a cooperation mechanism or information interaction. This isolated decision-making characteristic easily leads the system to fall into a local optimal dilemma.

[0116] 2) Hierarchical Deep Q-Network (H-DQN): As a classic algorithm for solving complex tasks in the field of reinforcement learning, H-DQN decouples the main task into multi-level subtasks by introducing a dual DQN hierarchical modeling architecture. In the scenario of this study, the discrete output of the lower-layer controller needs to be converted into continuous variables. However, due to its off-policy characteristic, compared with multi-agent algorithms, H-DQN shows weaker adaptive ability in non-steady-state environment training.

[0117] 3) Soft Actor-Critic algorithm (SAC): As an offline algorithm that integrates entropy maximization, the core architecture of SAC includes an actor-critic framework and a temperature parameter adaptive adjustment module. This algorithm significantly improves the robustness of the learning process through a stochastic policy and a soft Bellman update mechanism. However, SAC essentially belongs to a non-hierarchical architecture and needs to synchronously output all discrete and continuous variables in the original problem. Although this algorithm can effectively balance the policy robustness and exploration efficiency, it often performs poorly when dealing with complex tasks.

[0118] 4) The Deep Deterministic Policy Gradient algorithm (DDPG), as one of the most widely used algorithms currently, adopts a deterministic policy mechanism, which can achieve efficient and stable policy learning in a continuous action space, demonstrating excellent convergence characteristics and high sample utilization. However, this algorithm also faces similar limitation challenges as the SAC algorithm.

[0119] First, a comparative experiment was conducted on the system utility value, as Figure 5 shown. It can be seen that in terms of convergence speed, the algorithms adopting a hierarchical architecture are generally superior to non-hierarchical methods. Among them, the HAD-MARL algorithm proposed in the present invention shows a slight advantage, which benefits from its hierarchical framework that can decouple the complex action space, effectively alleviate the problem of reward sparsity and improve sample efficiency. In terms of the average utility value, HAD-MARL significantly exceeds all benchmark algorithms with a convergence value of 160.15: HMARL (147.73), H-DQN (141.81), SAC (129.19), and DDPG (125.21), improving the performance by 7.68% compared to the optimal benchmark scheme. This performance improvement not only stems from the advantages of the hierarchical architecture but, more crucially, from the effective use of the attention mechanism. Different from the traditional mode of isolated decision-making of agents, this solution realizes collaborative decision-making by dynamically aggregating the information of other agents, thereby achieving the effect of globally optimized resource allocation.

[0120] For the attention-driven multi-agent actor-critic module proposed in the present invention, although the effectiveness of its attention mechanism is difficult to be strictly verified by mathematical methods, it can be visually demonstrated through experimental visualization means how this mechanism promotes collaborative decision-making among agents in complex tasks. Under the default configuration, first, the user locations and service demands are randomly generated, and then the pairing relationship between access points and user devices is visually presented based on the converged model (as Figure 6 (a) shown). The experimental results show that all user devices have successfully accessed at least one accessible access point within the coverage range. On this basis, we further visually presented the attention weight distribution among access points (as Figure 6 (b) shown), as well as the entropy value distribution characteristics of different attention heads within a single access point (i.e., agent) (as Figure 6 (c) shown).

[0121] Explanation: The attention weight characterizes the correlation strength between input elements. When there is a high attention weight between access point i and access point j, it indicates that access point i needs to focus on the information shared by access point j during decision-making, and vice versa. The attention entropy is used to quantify the disorder degree of the weight distribution within a single attention head and reflects the functional role diversity among different attention heads: when the entropy value approaches 0, it means that this attention head highly focuses on extracting features from specific access points; a higher entropy value indicates that this attention head comprehensively considers multi-source information during feature extraction.

[0122] From Figure 6 (b), it can be observed that the more users are jointly served by two access points, the more necessary it is to consider their information interaction during resource allocation, so they will assign higher attention weights to each other. For example, access point 1 has extensive direct cooperation relationships with access point 2 and access point 3, so the attention weights assigned to them both exceed 0.25; while due to limited cooperation or no cooperation with access point 4 and access point 5, the attention of access point 1 to them is significantly reduced. On the other hand, access point 3 located in the center of the scenario undertakes more service tasks for multi-connected users, and it assigns relatively balanced attention weights to all other access points - this indicates that access point 3 needs to make decisions by integrating diverse information, so as to optimize the overall performance of the system during the resource allocation process.

[0123] To deeply analyze the operation mechanism of the attention mechanism, the entropy value distribution of different attention heads within each access point during the training process was further visualized ( Figure 6 (c)). It is found that the access points with more balanced attention weight distributions tend to have higher entropy values of their attention heads - this phenomenon is highly consistent with the previous analysis, and the entropy value distribution of access point 3 is a typical example. On the contrary, for access point 5, since it serves fewer users and mainly cooperates with access point 3 and access point 5, its attention weights are unevenly distributed, and the entropy values of each attention head show an increasing trend.

[0124] The experimental results fully verify the effectiveness of the adopted attention mechanism: each access point does dynamically integrate the information of co-serving peer nodes (i.e., adjacent access points sharing users) and achieves global optimization of resource allocation through a cooperation mechanism.

[0125] Although the specific implementation manners of the present invention are described above, those skilled in the art should understand that these are only examples, and the protection scope of the present invention is defined by the appended claims. Without departing from the principles and essence of the present invention, those skilled in the art can make various changes or modifications to these implementation manners, but these changes and modifications all fall within the protection scope of the present invention.

Claims

1. A cell-free network resource optimization method based on attention-driven hierarchical multi-agent reinforcement learning, characterized in that It includes the following steps: S1. Construct a system model for the co-located communication and sensing integrated multi-user and multi-service concurrent scenario; S2. Construct communication and sensing metrics respectively based on the Shannon theorem and the Cramer-Rao bound, and establish an optimization problem; S3. Design the HAD-MARL algorithm, including the hierarchical architecture design and the attention mechanism design; S4. Train the model according to the HAD-MRAL algorithm and update the parameters.

2. The method for optimizing the resource of a cell-free network based on attention-driven hierarchical multi-agent reinforcement learning according to claim 1, characterized in that In step S1, the system for the co-located communication and sensing integrated multi-user and multi-service concurrent scenario includes: a central controller, a number of access points, and a number of single-antenna users; Among them, the set of access points Each access point is equipped with N t transmitting antennas and N r receiving antennas to serve a set of single-antenna user equipments Each access point can adaptively transmit radar signals or communication signals according to the heterogeneous service requirements of user equipments; The service types of users are divided into three categories: communication-only users, sensing-only users, and communication and sensing integrated users; All access points are connected to the central controller through high-speed fronthaul links to achieve real-time information sharing. On the other hand, assume that the system is a time-division duplex system that uses the same spectrum resources and its uplink and downlink channels are reciprocal. For the convenience of subsequent modeling, define the set to represent users with communication service requirements, and the set to represent users with sensing service requirements. The sets and satisfy 3. The method for optimizing the resource of a cell-free network based on attention-driven hierarchical multi-agent reinforcement learning according to claim 2, wherein, In step S2, the optimization problem is specifically as follows: S2.

1. Construct the communication service metrics as follows: Among them, τ c represents the coherence time, τ p represents the length of the pilot sequence, b mk is the bandwidth allocated by access point k to user m, η mk is the signal-to-noise ratio and is related to the allocated power; S2.

2. Construct the sensing service metrics as follows: where CRB(τ jk ) and CRB(θ jk ) denote the Cramer-Rao bounds of the transmission delay and the angle of arrival respectively, and λ1 and λ2 are the weighting coefficients respectively; S2.

3. Construct the optimization problem: R min ≤R mk ω min ≤ ω jk Among them, U(·) is the system utility function, and the optimization variables are the pairing matrix of access points and users, represented by 1 or 0 to indicate whether the access point is connected to the user or not; the optimization variables b nk and P nk respectively represent the bandwidth and power resources allocated by the access point to the user; R min and ω min respectively represent the lower limits of the communication index and the sensing index, and w nk represents the beamforming matrix, and respectively represent the minimum and maximum values of the transmission power of the k-th access point, and respectively represent the minimum and maximum values of the bandwidth resources allocated by the k-th access point to the n-th user. B and P respectively represent the total bandwidth and the total power.

4. The method for optimizing the resource of a cell-free network based on attention-driven hierarchical multi-agent reinforcement learning according to claim 2, characterized in that In step S3, the hierarchical architecture design is specifically as follows: The hierarchical architecture includes a "commander controller" and a "staff controller"; among them, the upper-layer "commander controller" is deployed in the central controller, and the lower-layer "staff controller" is deployed in the access points. One agent is deployed in each access point, with a total of K agents; the "commander controller" is responsible for grasping the overall situation and generating the pairing matrix instruction of the access points and users; the "staff controller" then makes fine-grained power and bandwidth allocation strategies based on this instruction; Specifically, the "Commander Controller" obtains the global environmental state observation information and generates instruction c t , that is, the access point and user pairing matrix, which is used to indicate whether there is a connection between access point k and user n; the "Commander Controller" updates the instruction every f time slots, and the instruction remains unchanged between two update cycles; The state space of the "Commander Controller" is defined as where g nk (t) is the estimated channel state information, including the user service requirements; the instruction can be defined as The reward is defined as Finally, all experiences are stored in the form of tuples in the experience replay buffer ; The "commander controller" uses the classic Dueling DQN network, which has three layers. The hidden layers of the first two layers each contain 256 neurons, and the output layer has a total of 128 neurons. One neuron is responsible for the output of the advantage function, and the remaining neurons are responsible for the output of the value function; For the lower-level "employee controller", after receiving instruction c from the upper layer t it distributes it to each agent; each agent outputs a local policy based on the attention mechanism; For each agent, the state space is defined as The action space is defined as The reward function is defined as: where λ5 is the penalty coefficient, is the penalty function; The total reward function can be expressed as Store the experience tuple in the k-th experience replay buffer .

5. The method for optimizing the resource of a cell-free network based on attention-driven hierarchical multi-agent reinforcement learning according to claim 2, wherein, In step S3, the attention mechanism design is as follows: The lower-layer "staff controller" uses the attention-driven actor-critic multi-agent reinforcement learning algorithm. The working steps of the attention mechanism include: S3.2.1 The central attention mechanism collects the states and actions of each agent and obtains the corresponding embedding vectors through a single-layer neural network: [e1,..., e k ,..., e K , where all the embedding vectors except the k-th agent are denoted as e \k ; S3.2.2 For the attention head u, calculate the attention weights: Among them, W K and W Q are the linear mapping matrices corresponding to "Key" and "Query", D is the matrix dimension; the function Θ(·) is the softmax function; S3.2.3 Calculate the attention value: Among them, W V is the linear mapping matrix corresponding to "Value"; the function L(·) is a non-linear activation function, and ReLU is used here; S3.2.4 Generate the Q function based on the attention value: Among them, U represents the number of attention heads, and f k (·; θ k ) represents the critic network of the k-th agent, which includes a three-layer multi-layer perceptron, and θ k are its parameters.

6. The method for optimizing the resource of a cell-free network based on attention-driven hierarchical multi-agent reinforcement learning according to claim 4, characterized in that The upper-layer "commander controller" uses Dueling DQN: Define the Q function as: Among them, and are the value function and the advantage function generated by the neural network, respectively; The optimization objective is expressed as: Among them, represents the standard mathematical expectation; The gradient is expressed as: The network parameters are updated by the ADAM optimizer: φ h- ← τφ h- +(1 - τ)φ h- Among them, ξ h is the learning rate, τ << 1 is the soft update coefficient, m t and respectively represent the first-order and second-order momenta related to the bias.

7. The method for optimizing the resource of a cell-free network based on attention-driven hierarchical multi-agent reinforcement learning according to claim 4, wherein The lower-layer "staff controller" uses the attention-driven actor-critic multi-agent algorithm: Among them, The optimization objective of the Actor is defined as: The gradient is: Parameter update: wherein, is the learning rate; The optimization objective of the Critic is defined as: The gradient is: Parameter update: φ l- ← τφ l +(1 - τ)φ l- where ξ l is the learning rate, and τ << 1 is the soft update coefficient.

8. The method for optimizing the resource of a cell-free network based on attention-driven hierarchical multi-agent reinforcement learning according to claim 2, characterized in that In step S4, the training process of the HAD-MRAL algorithm is as follows: S4.

1. Initialize each environmental parameter and agent parameter φ h ,φ l ,and set φ h- = φ h and φ l- = φ l ; S4.

2. Start a new round. The upper-layer "Commander Controller" observes the environmental status and generates instruction c t ; S4.

3. The "Commander Controller" distributes the instructions to each agent in the lower-level "Employee Controller", and each agent generates the corresponding resource allocation strategy a according to the environment and the instructions k,t ; S4.

4. Each agent executes action a k,t , and obtains the reward feedback from the environment and the next state S4.

5. Each agent stores the record of this attempt into the corresponding experience replay buffer S4.

6. Each agent samples from the experience replay buffer to calculate the gradients and update the parameters φ l , φ l- ; S4.

7. Determine whether the number of time slots divides the update interval f. If so, the upper-layer "Commander Controller" calculates and obtains the reward r t h ; If not, directly jump to S4.10.; S4.

8. The upper-layer controller stores the record of this attempt into the experience replay buffer S4.

9. The upper-layer controller samples from the experience replay buffer to calculate the gradients and update the parameter φ h , φ h- ; S4.

10. Judge whether the maximum number of time slots is reached. If so, continue to execute; if not, jump to S4.3; S4.

11. Judge whether the maximum number of episodes is reached. If so, end the training; if not, jump to S4.2 to continue the training of the next episode.

Citation Information

Cited By

  • Power distribution network dynamic cluster planning control method and system based on electrical indexes

    CN120528034A

  • Cellular network uplink power control method based on graph attention-condition generation diffusion model

    CN121037958A