A proportional fairness power allocation method and system for a cell-free massive MIMO system
Patent Information
- Application Number
- CN202610276832.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-09
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-03-09
AI Technical Summary
[0005]约束处理不当:现有约束处理方法难以直接保证约束满足,往往需要额外的后处理步骤,且存在梯度消失、训练速度慢等问题;
[0034]本发明提供的一种去蜂窝大规模MIMO系统的比例公平性功率分配方法及系统,基于双延迟深度确定性策略梯度(TD3)算法框架,通过构建用户位置与信道状态的马尔可夫时空演进模型,精准捕捉移动场景下的非平稳动态特性,有效克服了传统迭代优化算法计算复杂度高及现有DDPG等方法普遍存在的Q值过估计问题;同时,结合数学全可导的安全层机制,在摒弃传统硬性截断与罚函数操作的前提下,严格保障了物理功率约束并保留了完整的梯度信息,从而实现了兼顾系统鲁棒性与多用户长期比例公平性的最优功率分配。
Smart Images

Figure CN122269426B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mobile communication technology, and in particular to a proportional fairness power allocation method and system for decellularized massive MIMO (multiple-input multiple-output) systems. Background Technology
[0002] Cellular-free massive MIMO systems are one of the key architectures of 6G mobile communication, widely used in scenarios such as low-altitude drone coverage and vehicle-to-everything (V2X) networks. With the increasing mobility of users and the diversification of devices, networks are becoming denser and more heterogeneous, making power allocation an important means to improve system performance, suppress interference, and extend device lifespan.
[0003] Traditional power optimization algorithms, such as fractional programming and weighted least mean square error (LMS), can improve system performance, but they require iterative inversion of high-dimensional matrices, and their computational complexity increases cubically with the number of users, making it difficult to meet the real-time requirements of mobile scenarios. Therefore, researchers have introduced deep reinforcement learning methods to achieve intelligent power control in dynamic environments.
[0004] However, existing deep reinforcement learning-based solutions still have the following drawbacks:
[0005] Improper constraint handling: Existing constraint handling methods cannot directly guarantee constraint satisfaction, often requiring additional post-processing steps, and suffer from problems such as gradient vanishing and slow training speed;
[0006] Algorithm limitations: For example, the DDPG algorithm suffers from problems such as insufficient exploration and overestimation of Q-values, leading to unstable training.
[0007] Insufficient mobility modeling: It often assumes that the user is stationary and does not consider dynamic features such as position and speed, making it difficult to adapt to real-world mobile scenarios;
[0008] Single optimization objective: Existing research focuses on single performance indicators such as maximizing the total system rate and the maximum and minimum rates, neglecting the importance of long-term proportional fairness in mobile scenarios, resulting in uneven resource allocation among users.
[0009] In summary, existing technologies have significant shortcomings in terms of constraint handling, basic algorithm models, mobility modeling, and optimization objectives. There is an urgent need for a power constraint method that is highly compatible with downlink power allocation, models user mobility characteristics, and supports a stable intelligent power allocation method that provides long-term proportional fairness. Summary of the Invention
[0010] This invention provides a proportionally fair power allocation method and system for decellularized massive MIMO systems, addressing the technical problem of how to achieve efficient, fair, and robust power allocation in decellularized massive MIMO systems in mobile scenarios.
[0011] To address the above technical problems, this invention provides a proportional fairness power allocation method for decellularized massive MIMO systems, comprising:
[0012] Establish a physical model of a decellularized massive MIMO system containing M access points and K mobile users;
[0013] Construct the input state vector of the deep reinforcement learning agent based on the physical model;
[0014] The deep neural network receives the input state vector and outputs the original action value;
[0015] The original action values are processed by power constraints based on a differentiable safety layer to generate power allocation coefficients;
[0016] Downlink data transmission is performed based on the power allocation coefficient, and a system reward value is calculated to train the deep neural network.
[0017] Furthermore, the input state vector includes: a large-scale fading coefficient matrix between all users and all access points at the current moment after logarithmic compression, normalized user location coordinates and velocity vectors, and a historical transmission rate sequence.
[0018] Furthermore, the physical model of the decellularized massive MIMO system includes network architecture, channel settings, user mobility settings, power constraints, and transmission sequence buffer settings. The network architecture includes the number and distribution of access points, the number of mobile users, and the coverage area. The channel settings include methods for estimating path loss, fading, and channel state information. The user mobility settings include setting the mobility methods of mobile users. The power constraints stipulate that the total power allocated to all users for each access point does not exceed its rated maximum transmit power. The transmission sequence buffer settings include determining the length of the user's historical transmission rate sequence that needs to be maintained and updating the user's long-term average transmission rate.
[0019] Furthermore, the original action values are processed by power constraints based on a differentiable safety layer to generate power allocation coefficients, specifically including:
[0020] Reshape the original action value (user count multiplied by access point count) into a two-dimensional matrix. And introduce temperature coefficient The original motion values are scaled to obtain , The Middle Line 1 Column element values Indicates the first The access point to the first The zoom action value for each user;
[0021] For each access point The original action values assigned to all users are subjected to a normalized exponential function along the user dimension to calculate the access point. Assigned to user power percentage ;
[0022] Finally, normalize the probability distribution. Multiply by the rated maximum transmit power of the access point Get access point Assigned to user Power allocation coefficient .
[0023] Further, downlink data transmission is performed according to the power allocation coefficient, and a system reward value is calculated to train the deep neural network, specifically including:
[0024] Based on the power allocation coefficient and channel state information, calculate the user At the present moment instantaneous transmission rate ;
[0025] The long-term average transmission rate for each user is updated using the exponential moving average method. ;
[0026] The sum of the logarithmic utility functions of the long-term average transmission rates of all users is calculated as the cumulative reward value at the current moment. ;
[0027] The double-delay deep deterministic policy gradient algorithm maximizes this cumulative reward value. To update the parameters of the deep neural network.
[0028] Furthermore, the long-term average transmission rate per user , For smoothing coefficients, This is the long-term average transmission rate calculated from the previous moment.
[0029] Furthermore, the cumulative reward value at the current moment. , is the numerical stability constant.
[0030] Furthermore, the policy network of the deep neural network receives the input state vector and outputs raw action values with a dimension of K×M. The policy network adopts a fully connected deep neural network structure, and the output layer of the policy network is configured to directly output a set of raw action values that have not undergone activation processing.
[0031] Furthermore, the policy network comprises two hidden layers, each with 256 neurons, and uses ReLU as the activation function.
[0032] This invention also provides a proportionally fair power allocation system for a decellularized massive MIMO system, used to implement the proportionally fair power allocation method for a decellularized massive MIMO system. The key feature of this system is that it includes a physical model building unit and a neural network unit; the physical model building unit is used to establish a physical model of a decellularized massive MIMO system containing M access points and K mobile users.
[0033] The neural network unit is used to: construct an input state vector for a deep reinforcement learning agent based on the physical model; receive the input state vector through the deep neural network and output the original action value; process the original action value with power constraints based on a differentiable security layer to generate power allocation coefficients; and perform downlink data transmission according to the power allocation coefficients and calculate the system reward value to train the deep neural network.
[0034] This invention provides a proportionally fair power allocation method and system for decellularized massive MIMO systems. Based on the dual-delay deep deterministic policy gradient (TD3) algorithm framework, it accurately captures the non-stationary dynamic characteristics in mobile scenarios by constructing a Markov spatiotemporal evolution model of user location and channel state. This effectively overcomes the high computational complexity of traditional iterative optimization algorithms and the Q-value overestimation problem commonly found in existing methods such as DDPG. At the same time, combined with a mathematically fully differentiable safety layer mechanism, it strictly guarantees physical power constraints and retains complete gradient information while abandoning traditional hard truncation and penalty function operations. Thus, it achieves optimal power allocation that balances system robustness and long-term proportional fairness among multiple users. Attached Figure Description
[0035] Figure 1 This is a flowchart of a proportional fairness power allocation method for a decellularized massive MIMO system provided in an embodiment of the present invention;
[0036] Figure 2 This is a comparison chart of the training performance of the five schemes provided in the embodiments of the present invention;
[0037] Figure 3 This is a comparison chart of the cumulative rewards of the five schemes provided in the embodiments of the present invention. Detailed Implementation
[0038] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. The embodiments are given for illustrative purposes only and should not be construed as limiting the present invention. The accompanying drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0039] This invention provides a proportional fairness power allocation method for decellularized massive MIMO systems, such as... Figure 1 The flowchart shown includes the following steps:
[0040] I. Establish a physical model of a decellularized massive MIMO system that includes multiple access points and multiple mobile users;
[0041] II. Constructing the input state vector of a deep reinforcement learning agent based on a physical model;
[0042] III. Deep neural networks receive input state vectors and output raw action values;
[0043] Fourth, the original action values are processed by power constraints based on a differentiable safety layer to generate power allocation coefficients;
[0044] 5. Perform downlink data transmission according to the power allocation coefficient and calculate the system reward value to train the deep neural network.
[0045] This method effectively solves the problems of high computational complexity and gradient loss caused by hard truncation in existing technologies by introducing a differentiable safety layer and a state space design containing historical rate sequences. It utilizes millisecond-level forward inference of deep neural networks to meet the real-time requirements of mobile scenarios, and the differentiable safety layer rigorously guarantees physical power constraints in mathematical principles while preserving complete gradient information, significantly improving training stability. Furthermore, by combining explicit modeling of users' historical service quality with long-term logarithmic utility optimization, this invention can proactively coordinate multi-user resource competition, effectively preventing marginal users from starving, thus achieving optimal power allocation in mobile environments that balances system capacity and long-term proportional fairness.
[0046] The following provides a more detailed explanation of each step.
[0047] (1) Step 1: Construct the physical model of the decellularized massive MIMO system
[0048] The physical model mainly includes core elements such as network architecture, channel settings, user mobility settings, power constraints, and transmission sequence buffer settings. Network architecture refers to network scale, including the number and distribution of access points, the number of mobile users, and coverage area. Channel settings include methods for estimating path loss, fading, and channel state information. User mobility settings refer to configuring the mobility methods of mobile users. Power constraints ensure that the total power allocated to all users at each access point does not exceed its rated maximum transmit power. Transmission sequence buffer settings include determining the length of the historical transmission rate sequence for each user that needs to be maintained and the method for updating the long-term average transmission rate of users.
[0049] As an example, the physical model settings for the decellularized massive MIMO system constructed in this embodiment are as follows:
[0050] Network architecture: M access points and K mobile users, downlink communication scenario, coverage area is The 15 access points are evenly distributed throughout.
[0051] Channel setup: The Hata-COST231 channel model is adopted, and channel state information is obtained through uplink pilot estimation. In the Hata-COST231 channel model, large-scale fading (path loss exponent is set to 3.7 and shadowing fading is included) is combined with small-scale Rayleigh fading.
[0052] User movement settings: A Gaussian-Markov hybrid movement model is used to simulate the continuous movement of users, and the movement speed of all users is kept within a range (e.g., 1~5 m / s, to simulate pedestrian or low-speed movement scenarios).
[0053] Power constraint: The total power allocated to all users at each access point shall not exceed its rated maximum transmit power (e.g., 1.0W).
[0054] Transmission sequence caching settings: The system simultaneously maintains the historical transmission rate sequence for each user in the background in real time. The sequence length is set to 10, which records the instantaneous transmission rate of each user in the most recent 10 time steps. This is used to assist the agent in identifying vulnerable users with long-term insufficient service. The long-term average transmission rate of each user is updated using the exponential moving average method.
[0055] (2) Step 2: Construct the input state vector of the deep reinforcement learning agent based on the physical model
[0056] At each decision-making time, an input state vector is constructed for the deep reinforcement learning agent. The input state vector consists of the following three parts (concatenated):
[0057] Channel state information, i.e., the large-scale fading coefficient matrix between all users and all access points at the current moment, and the numerical values of this matrix are... Logarithmic compression, i.e. , is the element in the m-th row and k-th column of the large-scale fading coefficient matrix.
[0058] User location and velocity information, i.e., normalized user location coordinates and velocity vectors, is used to provide mobility characteristics. It is generated by obtaining the user's two-dimensional coordinate position and velocity vector, and then normalizing them to... Interval.
[0059] The historical transmission rate sequence, which is the normalized instantaneous transmission rate of each user over the most recent 10 time steps, is used to characterize the temporal evolution of user service quality. It is generated by reading the historical transmission rate data of length 10 maintained in the first step and performing normalization processing.
[0060] These three pieces of information together constitute a composite state vector that can characterize the physical interference environment, user mobility characteristics, and the evolution of service quality.
[0061] (3) Step 3: The deep neural network receives the input state vector and outputs the original action value.
[0062] The policy network of a deep neural network receives an input state vector and outputs raw action values with a dimension of K×M. This output represents the degree to which each access point tends to allocate power to each user.
[0063] The policy network in this embodiment employs a fully connected deep neural network structure, containing two hidden layers, each with 256 neurons, and using ReLU as the activation function. The output layer of the policy network is configured to directly output a set of raw action values (Logits) without processing by the activation function. The dimension of this output value is... (Corresponding to 5 users) (15 access points), numerically representing the original tendency of each access point to allocate power to each user.
[0064] (4) Step 4: Generate power allocation coefficients based on the original action values
[0065] The original action values are input into the differentiable safety layer module, and the final power allocation coefficients that satisfy the physical constraints are generated through the following sub-steps:
[0066] First, multiply the number of users by the number of access points ( The original action values are reshaped into a two-dimensional matrix. To balance exploration and utilization, a temperature coefficient was introduced. The original motion values are scaled to obtain , The Middle Line 1 Column element values Indicates the first The access point to the first The zoom action value for each user;
[0067] Secondly, for each access point The original action values assigned to all users are subjected to a normalized exponential function along the user dimension to calculate the access point. Assigned to user power percentage The calculation formula is as follows: ,in Given the total number of users, this step ensures that for any access point, the sum of the power allocated to all users is strictly 1.
[0068] Finally, the power ratio Multiply by the rated maximum transmit power of the access point (e.g., 1.0 W), to obtain the final power allocation factor. .
[0069] Through the above processing, the system automatically satisfies the total power constraint condition for each access point, and the entire process is differentiable, ensuring effective gradient propagation. In this embodiment, the temperature coefficient... Setting it to 5.0 controls the concentration of power allocation, and the calculation process of the differentiable security layer module is mathematically differentiable throughout, preserving the complete gradient information from the backpropagation of the final power allocation coefficients to the policy network parameters.
[0070] (v) Step 5: Perform the action and calculate the reward to train the deep neural network
[0071] This step performs downlink data transmission based on the final power allocation coefficient and calculates the system's reward value to train the policy network. The specific process is as follows:
[0072] First, based on the power allocation coefficient and channel state information, calculate the user... At the present moment Instantaneous transmission rate ;
[0073] Secondly, the long-term average transmission rate for each user is updated using the exponential moving average method. The updated formula is as follows: ,in This is a smoothing coefficient used to control the degree of smoothing for instantaneous fluctuations. The long-term average transmission rate calculated at the previous moment;
[0074] Then, the sum of the logarithmic utility functions of the long-term average transmission rates of all users is calculated as the reward value at the current moment. ,in It is the numerical stability constant;
[0075] Finally, the TD3 (Dual Delay Deep Deterministic Policy Gradient) algorithm is used to update the deep neural network parameters by maximizing the cumulative reward value, thereby directly optimizing the long-term proportional fairness of the system.
[0076] In this embodiment, the smoothing coefficient Setting it to 0.05 ensures the system prioritizes long-term fairness and effectively suppresses instantaneous rate fluctuations.
[0077] It should be noted that the various processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved. This embodiment does not impose any limitations on these steps.
[0078] Based on the above-described proportional fair power allocation method for decellularized massive MIMO systems, this embodiment of the invention also provides a proportional fair power allocation system for decellularized massive MIMO systems, which includes a physical model building unit and a neural network unit. The physical model building unit is used to execute step one of the above-described method, and the neural network unit is used to execute steps two to five of the above-described method.
[0079] The embodiments described in this invention can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0080] Computer programs for implementing the methods and systems of the present invention may be written in any combination of one or more programming languages and stored in a computer-readable storage medium. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0081] Computer-readable storage media can be tangible media that may contain or store computer programs for use by or in conjunction with an instruction execution system, apparatus, or device. Computer-readable storage media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0082] The following comparison of this scheme with four other schemes verifies the performance improvement effect of the differentiable security layer proposed in this invention on deep reinforcement learning algorithms. Scheme 1 (Equal Power Allocation) is the baseline scheme for equal power allocation, where each access point distributes power evenly to all users. Scheme 2 (DDPG) uses the DDPG algorithm, where the policy network output is first mapped to the (0,1) interval using the Sigmoid function to ensure non-negative power, and then normalized column-wise to ensure that the total power constraint is met. Scheme 3 (TD3) uses the TD3 algorithm, with the same constraint processing method as Scheme 2, and the policy network output satisfies the power constraint condition through traditional numerical processing methods. Scheme 4 (DDPG + Security Layer) uses the DDPG algorithm combined with a differentiable security layer. After the policy network outputs the original action value, the power allocation coefficient is generated through Softmax normalization with a temperature coefficient τ=5.0. Scheme 5 (TD3 + Safety Layer) is the complete implementation of this scheme. The constraint processing method is the same as that of Scheme 4. It adopts the TD3 algorithm and combines it with a differentiable safety layer. After the policy network outputs the original action value, the power allocation coefficient is generated by Softmax normalization operation with temperature coefficient τ=5.0. While ensuring that the power constraint conditions are met, the complete gradient information is preserved, so as to achieve efficient and stable policy optimization and long-term proportional fairness goals.
[0083] Keeping the system environment parameters unchanged, we tested the five schemes and compared their performance. The results are as follows: Figure 2 , Figure 3 As shown. Figure 2 For the experimental comparison of different power allocation methods, the horizontal axis represents the number of training steps (unit: million steps), and the vertical axis represents the long-term proportional fairness utility value (i.e., the sum of the logarithmic utility functions of the long-term average transmission rate of all users). Figure 3 To compare our method with four other comparative schemes after training convergence using the cumulative distribution function (CDF) of user rates, the horizontal axis represents the instantaneous user transmission rate (in Mbps), and the vertical axis represents the CDF value, indicating the proportion of users at different rate levels. The specific implementation of the comparative experiment in this embodiment is as follows:
[0084] Step 1: Configure five schemes.
[0085] Step 2: Execute the training and evaluation process. Schemes 2, 3, 4, and 5 all undergo policy optimization over 6 million training steps, with performance evaluation performed every 20,000 steps. Scheme 5 strictly follows the complete process from Step 1 to Step 5. All schemes use the same neural network architecture, the same reward function, and the same training hyperparameter configuration.
[0086] Step 3: Record and compare performance metrics. For example... Figure 2 As shown, Option 1 has a long-term proportional fairness utility of 2.33, serving as the performance baseline. Option 2 has a utility of 2.54, representing a 9.0% improvement over the baseline. Option 3 has a utility of 2.57, a 10.3% improvement over the baseline. Option 4 has a utility of 3.78, a 62.2% improvement over the baseline. Option 5 achieves a utility of 4.08, a 75.1% improvement over the baseline.
[0087] Regarding training convergence speed, such as Figure 2 As shown, Scheme 5 achieves a utility value of over 3.8 after approximately 2 million training steps and maintains a stable increase; Scheme 4 achieves a utility value of over 3.5 after approximately 2.5 million training steps and remains stable; while Schemes 2 and 3 consistently maintain a utility value of around 2.5 throughout the entire 6 million training steps, indicating a significant lag in convergence speed. The training curves of Schemes 2 and 3 exhibit periodic oscillations, indicating that traditional numerical processing methods suffer from training instability. Comparing the two implementations of Schemes 4 and 5 reveals that the Softmax safety layer significantly outperforms traditional numerical processing methods in both the TD3 and DDPG algorithms, fully validating the universality and effectiveness of the safety layer of this invention.
[0088] Further analysis of user rate distribution fairness, such as Figure 3As shown, after training convergence, the four schemes were tested for 20 evaluation rounds, and the instantaneous transmission rate distribution of all users was statistically analyzed. The average user rate for Scheme 1 was 1.38 Mbps, Scheme 2 was 2.41 Mbps, Scheme 3 was 2.44 Mbps, Scheme 4 was 3.41 Mbps, and Scheme 5 reached 3.52 Mbps. From the cumulative distribution function curves, it can be observed that Scheme 5's curve is furthest to the right, indicating that it can provide higher transmission rates for more users, and the rate distribution is more concentrated, reflecting better fairness among users.
[0089] Step 4: Quantifying the performance improvement. Quantitative comparative analysis revealed that: compared to Scheme 3, Scheme 5 improved the long-term proportional fairness utility value by 58.8% (from 2.57 to 4.08); compared to Scheme 2, Scheme 4 improved the long-term proportional fairness utility value by 48.8% (from 2.54 to 3.78). These results demonstrate that the Softmax safety layer of this invention achieves significant performance improvements on both the TD3 and DDPG algorithms, validating the effectiveness of the scheme. Furthermore, the TD3 algorithm outperforms the DDPG algorithm in terms of training stability and convergence speed: when using the Softmax safety layer, Scheme 5 improved performance by 7.9% compared to Scheme 4 (from 3.78 to 4.08), and the convergence speed improved by approximately 25%.
[0090] The experimental results above show that the differentiable security layer proposed in this embodiment of the invention can significantly improve the performance of deep reinforcement learning algorithms in power allocation tasks in mobile scenarios, and the security layer has good versatility and can be used in combination with various deep reinforcement learning algorithms, thus verifying the effectiveness of the technical solution of this invention.
[0091] In summary, the proportional fairness power allocation method and system for decellularized massive MIMO systems provided by the embodiments of the present invention have the following beneficial effects:
[0092] 1. Through a differentiable safety layer, mathematically rigorous power constraint satisfaction was achieved for the first time in a decellularized MIMO system without sacrificing gradient information, solving the problems of unstable training by the penalty function method and inability to backpropagate by the truncation method.
[0093] 2. By introducing user location, speed, and historical rate sequences, a spatiotemporal evolution state space is constructed, enabling intelligent agents to perceive user movement trends, achieve proactive power pre-allocation, and significantly improve service quality in mobile scenarios;
[0094] 3. The Softmax safety layer of this invention has good versatility and can be used in conjunction with various deep reinforcement learning algorithms (such as TD3, DDPG, etc.). It can significantly improve the performance of different algorithms. As shown in the experimental section, the long-term proportional fairness utility values of the DDPG and TD3 algorithms (Schemes 4 and 5) combined with this safety layer are improved by 41.5% and 58.8% respectively compared with the corresponding algorithms using traditional constraint processing methods (Schemes 2 and 3).
[0095] 4. By designing the logarithmic utility function and the exponential moving average rate, long-term proportional fairness is directly optimized, solving the problem of ensuring fairness among users in mobile scenarios and avoiding long-term starvation for marginal users.
[0096] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A proportionally fair power allocation method for a decellularized massive MIMO system, characterized in that, include: Establish a physical model of a decellularized massive MIMO system containing M access points and K mobile users; Construct the input state vector of the deep reinforcement learning agent based on the physical model; The deep neural network receives the input state vector and outputs the original action value; The original action values are processed by power constraints based on a differentiable safety layer to generate power allocation coefficients, specifically including: Reshape the original action value (user count multiplied by access point count) into a two-dimensional matrix. And introduce temperature coefficient The original motion values are scaled to obtain , The Middle Line 1 Column element values Indicates the first The access point to the first The zoom action value for each user; For each access point The original action values assigned to all users are subjected to a normalized exponential function along the user dimension to calculate the access point. Assigned to user power percentage ; Finally, the power ratio Multiply by the rated maximum transmit power of the access point Get access point Assigned to user Power allocation coefficient ; Downlink data transmission is performed based on the power allocation coefficient, and a system reward value is calculated to train the deep neural network.
2. The proportional fairness power allocation method for a decellularized massive MIMO system according to claim 1, characterized in that, The input state vector includes: a large-scale fading coefficient matrix between all users and all access points at the current moment after logarithmic compression, normalized user location coordinates and velocity vectors, and a historical transmission rate sequence.
3. The proportional fairness power allocation method for a decellularized massive MIMO system according to claim 2, characterized in that: The physical model of the decellularized massive MIMO system includes network architecture, channel settings, user mobility settings, power constraints, and transmission sequence buffer settings. The network architecture includes the number and distribution of access points, the number of mobile users, and the coverage area. The channel settings include methods for estimating path loss, fading, and channel state information. The user mobility settings include setting the mobility methods for mobile users. The power constraint is that the total power allocated to all users for each access point does not exceed its rated maximum transmit power. The transmission sequence buffer settings include determining the length of the historical transmission rate sequence of users that needs to be maintained and methods for updating the long-term average transmission rate of users.
4. The proportional fairness power allocation method for a decellularized massive MIMO system according to claim 1, characterized in that, Performing downlink data transmission based on the power allocation coefficient and calculating the system reward value to train the deep neural network specifically includes: Based on the power allocation coefficient and channel state information, calculate the user At the present moment Instantaneous transmission rate ; The long-term average transmission rate for each user is updated using the exponential moving average method. ; The sum of the logarithmic utility functions of the long-term average transmission rates of all users is calculated as the cumulative reward value at the current moment. ; The double-delay deep deterministic policy gradient algorithm maximizes this cumulative reward value. To update the parameters of the deep neural network.
5. The proportional fairness power allocation method for a decellularized massive MIMO system according to claim 4, characterized in that: Long-term average transmission rate per user , For smoothing coefficients, This is the long-term average transmission rate calculated from the previous moment.
6. The proportional fairness power allocation method for a decellularized massive MIMO system according to claim 5, characterized in that: Cumulative reward value at the current moment , is the numerical stability constant.
7. The proportional fairness power allocation method for a decellularized massive MIMO system according to claim 1, characterized in that: The policy network of the deep neural network receives the input state vector and outputs raw action values with a dimension of K×M. The policy network adopts a fully connected deep neural network structure, and the output layer of the policy network is configured to directly output a set of raw action values that have not undergone activation processing.
8. The proportional fairness power allocation method for a decellularized massive MIMO system according to claim 7, characterized in that: The policy network consists of two hidden layers, each with 256 neurons, and uses ReLU as the activation function.
9. A proportionally fair power allocation system for a decellularized massive MIMO system, used to implement the proportionally fair power allocation method for a decellularized massive MIMO system according to any one of claims 1 to 8, characterized in that: The system includes a physical model building unit and a neural network unit; the physical model building unit is used to establish a physical model of a decellularized massive MIMO system containing M access points and K mobile users. The neural network unit is used to: construct an input state vector for a deep reinforcement learning agent based on the physical model; receive the input state vector through the deep neural network and output the original action value; and process the original action value through power constraints based on a differentiable safety layer to generate power allocation coefficients. Furthermore, downlink data transmission is performed based on the power allocation coefficient, and a system reward value is calculated to train the deep neural network.
Citation Information
Patent Citations
Cellular-removed large-scale MIMO uplink total rate first-order optimization method
CN114760647A
Multi-agent learning method for non-cellular network user scheduling and resource configuration
CN117221925A