Construction methods and resource management methods of wireless network resource allocation systems

By constructing a resource allocation system based on imperfect CSI in wireless networks, and utilizing DQN, DDQN, Dueling DQN, and DDPG network models to optimize channel and power allocation strategies, the resource allocation problem in imperfect CSI environments in wireless communication is solved, improving spectral efficiency and convergence speed.

CN116406004BActive Publication Date: 2026-04-07INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing learning-based wireless communication resource allocation methods perform poorly and have low convergence speeds in environments with imperfect global channel state information, failing to effectively address co-channel interference and spectral efficiency issues in wireless networks.

Method used

A wireless network resource allocation system is constructed. By acquiring imperfect global channel state information, an initial resource allocation system is trained using reinforcement learning, including channel allocation strategies and power allocation strategies. DQN, DDQN, Dueling DQN, and DDPG network models are used to optimize wireless network resource allocation to improve spectral efficiency and convergence speed.

Benefits of technology

In imperfect CSI environments, it improves the convergence rate and target performance of the wireless network resource allocation system, reduces co-channel interference, and enhances spectrum efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116406004B_ABST
    Figure CN116406004B_ABST
Patent Text Reader

Abstract

This invention provides a method for constructing a wireless network resource allocation system. The system is used to obtain a wireless network resource allocation strategy based on the wireless network state. The method includes: S1, obtaining a non-convex optimization objective with interruption probability constraints corresponding to wireless communication requirements under an imperfect global channel state information (CSI) environment; S2, transforming the obtained non-convex optimization objective to obtain a non-convex optimization objective without interruption probability constraints; S3, obtaining imperfect global channel state information of the wireless network; S4, using the obtained non-convex optimization objective as the training target and the imperfect global channel state information from step S3 as input, training the initial resource allocation system to convergence using reinforcement learning. This invention uses more realistic CSI to train the learning-based initial allocation system, improving the convergence rate of the wireless network resource allocation system and enhancing its performance in achieving the optimization objective.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication, specifically to the field of wireless communication network resource allocation, and more specifically to a method for constructing a wireless network resource allocation system, a wireless network resource management method based thereon, and a wireless communication system. Background Technology

[0002] In existing technologies, increasing the spatial spectrum reuse rate in wireless communication networks and deploying a large number of wireless access points (APs) can improve user transmission rates and network capacity. However, in scenarios with dense and irregular deployment of wireless access points, co-channel interference (CCI) can be particularly severe. Furthermore, as the number of base stations (BSs) deployed in a wireless communication network increases, unreasonable allocation of wireless network resources can further increase CCI and reduce communication performance, such as spectral efficiency. Therefore, it is necessary to optimize wireless network resource allocation (e.g., channel allocation strategies and power allocation strategies) to reduce CCI and improve communication performance, such as spectral efficiency.

[0003] There are two main types of existing methods for solving resource allocation problems in wireless networks: model-driven optimization algorithms and learning-based optimization algorithms.

[0004] Among them, model-driven optimization algorithms typically assume perfect global channel state information (CSI) to optimize resource allocation problems. When applied to actual wireless communication environments, they have excessively high computational complexity, resulting in large latency and high energy consumption. Their performance in solving resource allocation problems in wireless networks is suboptimal, making them difficult to deploy and apply in practice.

[0005] Learning-based optimization algorithms typically rely on deep reinforcement learning (DRL) for optimization. DRL leverages the powerful perceptual capabilities of deep learning to handle complex, high-dimensional environmental features and interacts with the environment using reinforcement learning principles to complete the decision-making process. Therefore, DRL has been successfully applied in various fields (autonomous driving decision-making, industrial robot control, and recommender systems). In the field of wireless communication, due to the dynamic nature of the wireless communication environment, resource allocation can also be modeled as a dynamic decision-making process. Therefore, applying deep reinforcement learning-based wireless resource management methods to wireless resource allocation tasks can solve the problems of traditional wireless resource allocation methods. Compared to model-driven resource optimization algorithms, learning-based optimization algorithms can effectively reduce the computational complexity of resource allocation and are more likely to be deployed and applied in future wireless network architectures. Currently, in the field of wireless communication technology, learning-based optimization algorithms commonly use perfect CSI (Computer-In-Size) for resource allocation in wireless networks. However, due to the objective existence of channel estimation errors and channel feedback delays, it is difficult to obtain a truly perfect CSI. Therefore, in wireless resource management tasks, it is necessary to consider a more realistic imperfect CSI in the wireless environment. This can be seen from the research in references [1]-[8], where optimization based on imperfect CSI is more realistic. However, as mentioned above, existing learning-based optimization methods are generally based on perfect CSI. For example, references [3]-[7] and [9] both design optimization targets based on perfect CSI. The algorithms have a slow convergence speed and low performance such as spectral efficiency. Moreover, as can be seen from the research in references

[10] -

[12] , perfect CSI is difficult to obtain in the actual environment.

[0006] In summary, existing learning-based methods are not designed for imperfect CSI (Channel Sequence Indicator). Furthermore, in real-world communication environments, channel estimation errors are inherent and cannot be completely eliminated. Directly applying existing learning-based algorithms in imperfect CSI environments results in poor optimization performance (i.e., communication performance) and slow convergence speed. Therefore, a more effective DRL (Dedicated Resource Allocation) architecture is urgently needed to optimize resource allocation strategies in wireless networks based on imperfect CSI.

[0007] References:

[0008] [1]Y.Teng,M.Liu,F.R.Yu,V.C.M.Leung,M.Song,and Y.Zhang,“Resourceallocation for ultra-dense networks:A survey,some research issues andchallenges,”IEEE Commun.Surv.Tut.,vol.21,no.3,pp.2134–2168,Jul.–Sep.2019.

[0009] [2]L.Liu,Y.Zhou,W.Zhuang,J.Yuan,and L.Tian,“Tractable coverageanalysis for hexagonal macrocell-based heterogeneous UDNs with adaptiveinterference-aware CoMP,”IEEE Trans.Wireless Commun.,vol.18,no.1,pp.503–517,Jan.2019.

[0010] [3]Y.Zhang,C.Kang,T.Ma,Y.Teng,and D.Guo,“Power allocation in multi-cell networks using deep reinforcement learning,”in Proc.IEEE 88thVeh.Technol.Conf.(VTC-Fall),2018,pp.1–6.

[0011] [4]S.Lahoud,K.Khawam,S.Martin,G.Feng,Z.Liang,and J.Nasreddine,“Energy-efficient joint scheduling and power control in multicell wirelessnetworks,”IEEE J.Sel.Areas Commun.,vol.34,no.12,pp.3409–3426,Dec.2016.

[0012] [5]K.Shen and W.Yu,“Fractional programming for communicationsystems—Part I:Power control and beamforming,”IEEE Trans.Signal Process.,vol.66,no.10,pp.2616—2630,May 2018.

[0013] [6]F.Meng,P.Chen,L.Wu,and J.Cheng,“Power allocation in multi-usercellular networks:Deep reinforcement learning approaches,”IEEE Trans.WirelessCommun.,vol.19,no.10,pp.6255–6267,Oct.2020.

[0014] [7]J.Tan,Y.-C.Liang,L.Zhang,and G.Feng,“Deep reinforcement learningfor joint channel selection and power control in D2D networks,”IEEETrans.Wireless Commun.,vol.20,no.2,pp.1363–1378,Feb.2021.

[0015] [8]Y.Guo,F.Zheng,J.Luo,and X.Wang,“Optimal resource allocation viamachine learning in coordinated downlink multi-cell OFDM networks underimperfect CSI,”in Proc.Veh.Technol.Conf.(VTC-Spring),2020,pp.1–6

[0016] [9] YSNasir and D.Guo, "Deep Reinforcement Learning for JointSpectrum and Power Allocation in Cellular Networks," 2021 IEEE GlobecomWorkshops (GC Wkshps), 2021, pp.1-6.

[0017]

[10] T.Yoo and A.Goldsmith, "Capacity and power allocation for fadingMIMO channels with channel estimation error," IEEE Trans.Inf.Theory, vol.52, no.5, pp.2203–2214, May 2006.

[0018]

[11] F.Fang, H.Zhang, J.Cheng, S.Roy, and VCMLeung, "Joint users scheduling and power allocation optimization for energy-efficient NOMAsystems with imperfect CSI," IEEE J.Sel.Areas Commun., vol.35, no.12, pp.2874–2885, Dec.2017.

[0019]

[12] X.Wang,F.-C.Zheng,P.Zhu,and Summary of the Invention

[0020] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a method for constructing a wireless network resource allocation system, a wireless network resource management method based thereon, and a wireless communication system.

[0021] The objective of this invention is achieved through the following technical solution:

[0022] According to a first aspect of the present invention, a method for constructing a wireless network resource allocation system is provided. The wireless network resource allocation system is used to obtain a wireless network resource allocation strategy based on the wireless network state. The method includes: S1, obtaining a non-convex optimization objective with interruption probability constraints corresponding to wireless communication requirements under an imperfect global channel state information environment; S2, transforming the non-convex optimization objective obtained in step S1 to obtain a non-convex optimization objective without interruption probability constraints; S3, obtaining imperfect global channel state information of the wireless network; S4, using the non-convex optimization objective in step S2 as a training objective and the imperfect global channel state information in step S3 as input, training an initial resource allocation system to convergence using reinforcement learning. The initial resource allocation system is a system constructed based on an agent and used to generate an action set based on the wireless network state. The action set includes a channel allocation strategy and a power allocation strategy.

[0023] In some embodiments of the present invention, the wireless communication requirement is to maximize the spectral efficiency of the wireless network, and the non-convex optimization objective with outage probability constraints is:

[0024]

[0025] in,

[0026]

[0027]

[0028]

[0029]

[0030]

[0031]

[0032] in, This represents the average spectral efficiency of the wireless network in time slot t, where K represents the total number of links and N represents the total number of sub-channels. This represents the set of sub-channel indices. This represents the scheduling spectral efficiency of the k-th link in selecting the n-th sub-channel in time slot t. This represents the maximum spectral efficiency of the k-th link when selecting the n-th sub-channel in time slot t. This represents the estimated small-scale fading component when the k-th link selects the n-th sub-channel in time slot t. This indicates the estimation of small-scale fading components. Under the condition The probability, p represents the power of the k-th link in selecting the n-th sub-channel in time slot t. t Indicates all The collection of power, α represents the identifier value of the k-th link after selecting the n-th sub-channel in time slot t. t Indicates all The set of identifier values, ε out P represents the expected probability of interruption. max The constraint M1 represents the power threshold of the link, and the constraint M1 represents the estimation of small-scale fading components. Under the condition that the probability of any link being interrupted after selecting any sub-channel in time slot t is less than the expected interruption probability, constraint M2 means that the transmit power on each link cannot exceed the power threshold of the link, and constraints M3 and M4 mean that each link can only select one sub-channel in each time slot.

[0033] In some embodiments of the present invention, in step S2, the non-convex optimization objective is transformed by parameter transformation to obtain a non-convex optimization objective without interruption probability constraints:

[0034]

[0035] in,

[0036]

[0037]

[0038]

[0039]

[0040]

[0041] Among them, Ω t The value represents the average spectral efficiency of the wireless network in time slot t after parameter transformation.

[0042] In some embodiments of the present invention, the initial resource allocation system includes: a channel allocation model and a power allocation model. The channel allocation model is used to predict the channel allocation strategy for a time slot based on the imperfect global channel state information of that time slot, and is configured as a DQN network, a DDQN network, or a Dueling DQN network. The power allocation model is used to predict the power allocation strategy for a time slot based on the imperfect global channel state information of that time slot, and is configured as a DDPG network.

[0043] In some embodiments of the present invention, step S4 includes steps S41, S42, and S43. Step S41 includes: acquiring imperfect global channel state information of the input time slot and performing the following steps: S411, the channel allocation model predicts the channel allocation strategy for the input time slot based on the imperfect global channel state information of the input time slot; updates the imperfect global channel state information of the input time slot based on the predicted channel allocation strategy; and the power allocation model predicts the power allocation strategy for the input time slot based on the updated imperfect global channel state information of the input time slot; the predicted channel allocation strategy and power allocation strategy of the input time slot are interacted with the wireless network to obtain the imperfect global channel state information of the next time slot of the input time slot; and the channel allocation model predicts the channel allocation strategy of the next time slot of the input time slot based on the imperfect global channel state information of the next time slot of the input time slot. The allocation strategy is as follows: S411: Update the channel allocation strategy for the next time slot of the input time slot based on the channel allocation strategy of the input time slot; S412: Calculate the spectral efficiency bonus for the input time slot based on the channel allocation strategy and power allocation strategy of the input time slot; S413: Store a channel allocation experience as a channel selection replay pool using the imperfect global channel state information of the input time slot, the channel allocation strategy of the input time slot, the spectral efficiency bonus of the input time slot, and the imperfect global channel state information of the next time slot of the input time slot; Store a power allocation experience as a power selection replay pool using the updated imperfect global channel state information of the input time slot, the power allocation strategy of the input time slot, the spectral efficiency bonus of the input time slot, and the updated imperfect global channel state information of the next time slot of the input time slot. Step S42 includes: Using the imperfect global channel state information of the next time slot of the previous input time slot as the new imperfect global channel state information of the input time slot. Step S43 includes: Updating the initial resource allocation system parameters based on the channel allocation experience in the channel selection replay pool and the power allocation in the power selection replay pool until convergence.

[0044] In some embodiments of the present invention, in step S43, the parameters of the channel allocation model are updated when there is a channel allocation experience in the channel selection replay pool; the parameters of the power allocation model are updated when there is a power allocation experience in the power selection replay pool.

[0045] In some embodiments of the present invention, in step S43: when the channel allocation experience in the channel selection replay pool reaches a preset number of experience points, the parameters of the channel allocation model are updated multiple times until convergence. Each update involves randomly sampling from the channel selection replay pool to obtain multiple channel allocation experiences, and updating the parameters of the channel allocation model using gradient descent based on the sampled channel allocation experiences. Similarly, when the power allocation experience in the power selection replay pool reaches a preset number of experience points, the parameters of the power allocation model are updated multiple times until convergence. Each update involves randomly sampling from the power selection replay pool to obtain multiple power allocation experiences, and updating the parameters of the power allocation model using gradient descent based on the sampled power allocation experiences.

[0046] In some embodiments of the present invention, in step S41, the imperfect global channel state information of the input time slot includes imperfect global channel state information of multiple links selecting different sub-channels in the input time slot:

[0047]

[0048] in,

[0049]

[0050] in, This represents the state set of the k-th link when selecting the n-th sub-channel in time slot t. This represents the independent channel gain of the k-th link when selecting the n-th sub-channel in time slot t, given the existence of channel estimation error. This represents the channel power of the k-th link when selecting the n-th sub-channel in time slot t. This represents the identifier value of the k-th link after selecting the nth sub-channel in time slot t-1. This represents the power of the k-th link when selecting the n-th sub-channel in time slot t-1. This represents the spectral efficiency of the k-th link in time slot t-1. This represents the estimated small-scale fading component corresponding to the nth sub-channel selected by the kth link in time slot t. The ratio of the total interference power to the total interference power, ranked across all channels. This indicates co-channel interference when the k-th link selects the n-th sub-channel in time slot t, using the sub-channel allocation scheme and power allocation scheme of the previous time slot. k′ represents other links different from k. The variance representing the channel estimation error. It is a large-scale fading component that takes into account both shadow fading and geometric decay. This indicates that the mean is 0 and the variance is . The complex Gaussian distribution.

[0051] In some embodiments of the present invention, the spectral efficiency bonus is calculated in the following manner:

[0052]

[0053] in,

[0054]

[0055] in, ε represents the spectral efficiency of the k-th link when selecting the n-th sub-channel in time slot t. out This represents the expected probability of interruption. Let φ represent the scheduling spectral efficiency of the k-th link in selecting the n-th sub-channel in time slot t, where φ is the interference weighting coefficient, and k′ represents other links different from k. This indicates the external interference of the k-th link selecting the n-th sub-channel in time slot t. This represents the spectral efficiency of link k′ in the nth subchannel of time slot t, where there is no interference from the kth link. This represents the spectral efficiency of the k′-th link when selecting the nth subchannel in time slot t.

[0056] In some embodiments of the present invention, the DDPG network includes an Actor network and a Critic network, and the final resource allocation system is: a DQN network trained to convergence, a DDQN network, or a Dueling DQN network and an Actor network.

[0057] According to a second aspect of the present invention, a wireless network resource management method is provided, the method comprising: T1, obtaining the wireless network state of the wireless communication system in the previous time slot; T2, based on the wireless network state of the previous time slot obtained in step T1, predicting the resource allocation strategy for the next time slot using the resource allocation system obtained by the method of the first aspect of the present invention; T3, allocating wireless network resources in the wireless communication system based on the resource allocation strategy for the next time slot obtained in step T2.

[0058] According to a third aspect of the present invention, a wireless communication system is provided, the system comprising a plurality of base stations, each base station comprising a wireless resource management unit configured to allocate wireless network resources in the base station using the method described in the second aspect of the present invention.

[0059] Compared with the prior art, the advantages of the present invention are as follows: it adopts a non-convex optimization objective with interruption probability constraints corresponding to wireless communication requirements under imperfect global channel state information environment as the training objective, which can fully consider the channel estimation error in the actual communication environment. That is, it uses more realistic CSI (imperfect global channel state information) to train the learning-based initial resource allocation system, thereby improving the convergence rate of the wireless network resource allocation system and improving the performance of achieving the optimization objective. Attached Figure Description

[0060] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:

[0061] Figure 1 This is a flowchart illustrating a method for constructing a wireless network resource allocation system according to an embodiment of the present invention.

[0062] Figure 2 This is a schematic diagram of the model training and parameter update architecture of an initial allocation system composed of a Dueling DQN network and a DDPG network according to an embodiment of the present invention;

[0063] Figure 3 This is a schematic diagram of a wireless network resource management method according to an embodiment of the present invention;

[0064] Figure 4 This is a schematic diagram comparing the convergence performance of the algorithm proposed in this patent according to an embodiment of the present invention and the four baseline algorithms mentioned above.

[0065] Figure 5 This is a schematic diagram comparing the relationship between the spectral efficiency and the variance of the channel estimation error that can be achieved by the algorithm proposed in the patent according to the embodiment of the present invention and the above four baseline algorithms.

[0066] Figure 6 This diagram illustrates the performance comparison of the spectral efficiency of the algorithm proposed in this patent and the four baseline algorithms described above under different numbers of sub-channels, according to embodiments of the present invention. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of this invention clearer, the invention is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0068] As mentioned in the background section, existing learning-based methods are not designed for imperfect CSI (Channel Sequence Indicator). Furthermore, in real-world communication environments, channel estimation errors are inherent and cannot be completely eliminated. Directly applying existing learning-based algorithms in imperfect CSI environments results in poor optimization performance and slow convergence speed. To address these issues, this invention proposes a wireless network resource allocation scheme based on the characteristics of imperfect CSI. By modeling the wireless network resource allocation problem under imperfect CSI as a problem of solving a non-convex optimization objective with outage probability constraints corresponding to wireless communication requirements, and considering the difficulty in solving optimization objectives with probability constraints, this invention further transforms the non-convex optimization objective with outage probability constraints into a non-convex optimization objective without outage probability constraints and uses a learning-based method to solve it. This improves the optimization performance and convergence speed achievable in imperfect CSI environments.

[0069] To better understand the present invention, the following detailed description of the invention is provided in conjunction with the accompanying drawings and embodiments.

[0070] According to an embodiment of the present invention, a method for constructing a wireless network resource allocation system is provided. The wireless network resource allocation system is used to obtain a wireless network resource allocation strategy based on the wireless network status, such as... Figure 1 As shown, the method includes: S1, obtaining a non-convex optimization objective with interruption probability constraints corresponding to wireless communication requirements under an imperfect global channel state information environment; S2, transforming the non-convex optimization objective obtained in step S1 to obtain a non-convex optimization objective without interruption probability constraints; S3, obtaining imperfect global channel state information of the wireless network; S4, using the non-convex optimization objective in step S2 as the training objective and the imperfect global channel state information in step S3 as input, training the initial resource allocation system to convergence using reinforcement learning, wherein the initial resource allocation system is a system built based on an agent and used to generate an action set based on the wireless network state, the action set including channel allocation strategy and power allocation strategy. To better illustrate the specific scheme of this invention, the following details the establishment of the non-convex optimization objective without interruption probability constraints, model training, and experimental verification.

[0071] I. Establishment of a Non-convex Optimization Objective Without Interruption Probability Constraints

[0072] Since existing learning-based network resource allocation methods are not designed for imperfect CSI (Communications for Integrated Services), the following section will explain in detail the establishment and transformation of optimization objectives for wireless communication networks for better understanding. Specifically, for ease of comprehension, this process is illustrated using formula derivation in this embodiment.

[0073] This invention first provides a mathematical description of the wireless network environment under imperfect CSI, and then models the wireless network environment based on this mathematical description. The wireless communication network includes multiple communication areas, each with a base station and multiple users. All users in the multiple communication areas share multiple sub-channels. Each base station is located at the center of its area, and licensed users are randomly distributed within the communication areas. All users and base station transceivers are equipped with a single antenna, and each formed link can only select one sub-channel in a time slot. For example, in a wireless network scenario with multiple cells and multiple users in a downlink, K links are distributed across M cells and share N orthogonal sub-channels. and These represent the link index set, cell index set, and sub-channel index set, respectively.

[0074] In a wireless communication environment, considering a fully synchronized system with time intervals, the independent channel gain of the k-th link choosing the n-th sub-channel in time slot t can be expressed as:

[0075]

[0076] in, This represents the large-scale fading component considering both shadow fading and geometric decay, where it is assumed that... It remains unchanged across multiple time slots; This represents the estimated small-scale fading component when the k-th link selects the n-th sub-channel in time slot t.

[0077] In a wireless communication environment, considering normalized bandwidth and perfect CSI, the maximum spectral efficiency of the k-th link in selecting the n-th sub-channel in time slot t is:

[0078]

[0079] in, This represents the identifier value of the k-th link after selecting the n-th sub-channel in time slot t, for example... This indicates that the k-th link selected the n-th sub-channel in time slot t; otherwise... σ represents the power of the k-th link in selecting the n-th sub-channel in time slot t. 2 This represents the power of additive white Gaussian noise. This represents the co-channel interference experienced by the k-th link when the n-th sub-channel is selected in time slot t.

[0080] In wireless communication environments, due to unavoidable channel estimation errors, perfect CSI... Assuming true values ​​ignores channel estimation errors in actual communication environments. Therefore, an objective estimation of small-scale fading components is needed. While the base station can perfectly estimate large-scale fading coefficients due to their slow changes, small-scale fading coefficients cannot be perfectly estimated due to their rapid changes. Thus, in one embodiment of this invention, based on imperfect CSI, the estimated small-scale fading component of the k-th link selecting the n-th sub-channel in time slot t is expressed as:

[0081]

[0082] in,

[0083] in, This represents the estimated small-scale fading component when the k-th link selects the n-th sub-channel in time slot t. This represents the error in estimating the small-scale fading component when the k-th link selects the n-th sub-channel in time slot t, and each They are independent of each other. This indicates that the mean is 0 and the variance is . The complex Gaussian distribution, This indicates that the mean is 0 and the variance is . The complex Gaussian distribution, This represents the variance of the channel estimation error. It should be noted that the existing perfect CSI information has a defect, mainly referring to the fact that small-scale fading coefficients usually cannot be perfectly estimated, as shown in formula (3). Due to the existence of channel estimation error and other factors, the channel estimate of small-scale fading coefficients... It is usually not equal to the true value. Applying an existing perfect CSI algorithm directly to a non-perfect CSI environment is equivalent to directly applying the estimated value... Treating the estimated value as the true value for resource allocation is problematic because of the error between the estimated and true values ​​(i.e., channel estimation error). Therefore, directly using an algorithm based on perfect CSI for network resource allocation generally only improves transmission performance and network capacity. In reality, channel estimation errors and other factors inevitably prevent perfect CSI estimation. Therefore, this imperfect CSI factor must be considered. Furthermore, directly using an existing resource allocation algorithm based on perfect CSI is equivalent to substituting the estimated value for the actual value, which degrades the algorithm's performance.

[0084] Following the mathematical description of the wireless network environment described above, the optimization problem (i.e., wireless communication requirements) will now be modeled. The optimization problem includes at least maximizing throughput and maximizing spectral efficiency. Since maximizing throughput and maximizing spectral efficiency are interchangeable by formula, this embodiment will use maximizing spectral efficiency as an example for modeling and explanation; the modeling of maximizing throughput will not be elaborated here.

[0085] Due to the influence of imperfect CSI, the spectral efficiency of scheduling may exceed the maximum achievable spectral efficiency defined by the Shannon capacity formula. Therefore, when the spectral efficiency of scheduling exceeds the achievable spectral efficiency under imperfect CSI, the outage probability is used as a performance metric. The scheduling spectral efficiency of the k-th link selecting the n-th sub-channel in time slot t is expressed as: The average spectral efficiency of the wireless network in time slot t is given by the following formula:

[0086]

[0087] Furthermore, under imperfect CSI in time slot t, the non-convex optimization objective with outage probability constraints corresponding to maximizing the spectral efficiency of the wireless network is:

[0088]

[0089] in,

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096] in, This represents the average spectral efficiency of the wireless network in time slot t, where K represents the total number of links and N represents the total number of sub-channels. This represents the set of sub-channel indices. This represents the scheduling spectral efficiency of the k-th link in selecting the n-th sub-channel in time slot t. This represents the maximum spectral efficiency of the k-th link when selecting the n-th sub-channel in time slot t. This represents the estimated small-scale fading component when the k-th link selects the n-th sub-channel in time slot t. This indicates the estimation of small-scale fading components. Under the condition The probability, p represents the power of the k-th link in selecting the n-th sub-channel in time slot t. t Indicates all The collection of power, α represents the identifier value of the k-th link after selecting the n-th sub-channel in time slot t.t Indicates all The set of identifier values, ε out P represents the expected probability of interruption. max The constraint M1 represents the power threshold of the link, and the constraint M1 represents the estimation of small-scale fading components. Under the condition that the probability of any link being interrupted after selecting any sub-channel in time slot t is less than the expected interruption probability, constraint M2 means that the transmit power on each link cannot exceed the power threshold of the link, and constraints M3 and M4 mean that each link can only select one sub-channel in each time slot.

[0097] Since non-convex optimization objectives with outage probability constraints are proven to be NP-hard problems (problems that can be reduced to polynomial time complexity for all nondeterministic polynomial problems) even when the sub-channel strategy in wireless network resources is fixed and only power allocation is considered, the optimal solution of non-convex optimization objectives with outage probability constraints is difficult to solve directly through mathematical derivation. This invention addresses this problem by using parameter transformation to convert the original optimization objective (i.e., the non-convex optimization objective with outage probability constraints) into a non-convex optimization objective without outage probability constraints (through constraint replacement and corresponding solution transformation). This allows the solution of the non-convex optimization objective with outage probability constraints that maximizes the spectral efficiency of the wireless network to be solved. The following sections will detail the process of transforming the original optimization objective using parameter transformation, focusing on constraint replacement and optimization problem transformation.

[0098] In the constraint replacement process, the inventors considered a stricter constraint R1 to replace the interruption probability constraint M1, such that constraint R1 always satisfies the interruption probability constraint M1 in the above non-convex optimization objective. Constraint R1 is:

[0099]

[0100] in, This represents the noise and interference signal strength of the k-th link in time slot t when selecting the n-th sub-channel, as defined by Shannon's formula. This represents the noise and interference signal strength of the k-th link when selecting the n-th sub-channel in time slot t, under actual schedulable conditions. This represents the useful signal strength of the k-th link in time slot t when selecting the n-th sub-channel, under actual schedulable conditions. Let Rk represent the useful signal strength of the k-th link in time slot t when selecting the n-th subchannel, as defined by Shannon's formula. Constraint R1-1 represents... Under the condition that for all k and n, Less than The probability cannot be greater than Constraint R1-2 indicates that in Under the condition that for all k and n, Less than The probability of each is equal to

[0101] The following will explain the proof that constraint R1 is more stringent than interruption constraint M1. The proof consists of two parts: parameter definition and proof reasoning.

[0102] The parameter definition section is as follows: First, it is defined according to Shannon's formula.

[0103]

[0104] Similarly, in the case of imperfect CSI, the scheduling spectral efficiency of the k-th link selecting the n-th sub-channel in time slot t is defined as:

[0105]

[0106] in, This represents the channel interference ratio of the k-th link in time slot t when selecting the n-th sub-channel under actual schedulable conditions.

[0107] From formulas (2) and (7), we can obtain

[0108]

[0109] but From formula (8), we can obtain The original interruption probability constraint M1 can then be written as:

[0110]

[0111] Substituting formulas (7) and (8) into formula (9), we can obtain:

[0112]

[0113] According to the law of total probability, we can obtain:

[0114]

[0115] in,

[0116]

[0117] Where Pr(E1) represents the value of E1. and Under the conditions Less than The probability, Pr(E2), represents the probability in and Under the conditions Less than The probability of.

[0118] The proof and reasoning part is as follows:

[0119] The proof of constraint R1-1 in constraint R1 is as follows:

[0120] constrain R1-1 Replace with but Therefore, we can obtain For Pr(E2), since Then there is It must be less than

[0121] The proof of constraint R1-2 in constraint R1 is as follows:

[0122] According to the law of total probability, we can obtain... Since Pr(E1)≤1, then we have

[0123] The proofs of constraints R1-1 and R1-2 in the above constraint R1 show that constraint R1 is a more stringent constraint than constraint M1.

[0124] During the optimization problem transformation process, the original optimization problem is transformed according to the stricter constraint R1. The following explains the specific derivation process from the transformation of constraint R1-1, constraint R1-2, and the transformation of the non-convex optimization objective with interruption probability constraint.

[0125] Based on the more stringent constraint R1-1 above, we can obtain:

[0126]

[0127] According to Markov's inequality, we can obtain from formula (12):

[0128]

[0129] Let the right side of formula (13) equal to Then we have:

[0130]

[0131] Based on the more stringent constraint R1-2 above, we can obtain:

[0132]

[0133] Where F represents the cumulative distribution function of the chi-square distribution, let formula (15) equal to We can obtain:

[0134]

[0135] Among them, F -1 This represents the inverse cumulative distribution function (CDF) of the chi-square distribution. Because... and Substituting these two terms and formula (16) into formula (14) yields:

[0136]

[0137] Therefore, we can conclude that:

[0138]

[0139] Formula (18) is equivalent to:

[0140]

[0141] Therefore, the average spectral efficiency of the wireless network in time slot t after parameter transformation is expressed as:

[0142]

[0143] Among them, F -1 The inverse cumulative distribution function (CDF) represents the chi-square distribution.

[0144] In summary, the non-convex optimization objective with interruption probability constraints is transformed into a non-convex optimization objective without interruption probability constraints as follows:

[0145]

[0146] in,

[0147]

[0148]

[0149]

[0150]

[0151]

[0152] Among them, Ω tThe value represents the average spectral efficiency of the wireless network in time slot t after parameter transformation. It should be noted that this invention considers imperfect Channel Instability (CSI) caused by channel estimation errors in resource allocation for more practical scenarios. Since imperfect CSI leads to an outage probability, constraints with outage probabilities in the optimization model cannot be directly solved using existing algorithms based on perfect CSI. Therefore, after parameter transformation of the optimization model, a new learning algorithm based on imperfect CSI is designed for the transformed model. Imperfect CSI, channel estimation error, and other parameters are designed as part of the state set, enabling the deep reinforcement learning network to effectively learn the impact of imperfect CSI, thereby improving the optimization objective and algorithm performance achievable by the learning-based algorithm in imperfect CSI environments.

[0153] II. Model Training

[0154] After the above steps, the non-convex optimization objective with interruption probability constraints is transformed to obtain a non-convex optimization objective without interruption probability constraints (Formula 21), which still belongs to the NP-Hard problem. Traditional algorithms, such as the solution scheme recorded in reference [5] mentioned in the background section, require multiple iterations to converge and cannot be well expanded as the number of user links increases. In addition, it is very challenging for the centralized controller in the communication system to obtain the instantaneous global CSI and send the allocation scheme back to the BS. In order to make the non-convex optimization objective without interruption probability constraints solvable, the joint wireless communication requirements (this embodiment takes the wireless communication requirement as maximizing the spectrum efficiency of the wireless network as an example for explanation, but it does not mean that the wireless communication requirement is only to maximize the spectrum efficiency) are first decoupled into two sub-problems, namely the sub-channel selection sub-problem and the power allocation sub-problem. Then, the learning model (initial resource allocation system) that can handle these two sub-problems at the same time is used to handle the problem of maximizing the spectrum efficiency of the wireless network to improve the convergence performance of the final resource allocation system and the effect of the optimization objective. The final resource allocation system is obtained by training the initial resource allocation system with a training set consisting of a non-convex optimization objective and a training set composed of imperfect global channel state information and resource allocation strategies related to the training objective.

[0155] According to an embodiment of the present invention, the initial resource allocation system includes: a channel allocation model (also referred to as a first-layer network in this embodiment) and a power allocation model (also referred to as a second-layer network in this embodiment); the channel allocation model is used to predict the channel allocation strategy for a time slot based on imperfect global channel state information of that time slot; preferably, the channel allocation model is configured as a DQN network, a DDQN network, or a Dueling DQN network; the power allocation model is used to predict the power allocation strategy for a time slot based on imperfect global channel state information of that time slot, preferably, the power allocation model is configured as a DDPG network. It should be noted that the channel allocation subproblem is a discrete task, while the power allocation subproblem is a continuous task. The two-layer learning network architecture composed of the channel allocation model and the power allocation model described in the previous embodiment can avoid introducing quantization errors. Specifically, for channel allocation, a DQN network, a DDQN network, or a Dueling DQN network is used to handle discrete variable resources; for power allocation, the channel power is determined by P... max For constrained continuous scalars (for some algorithms, such as value-based DQN algorithms, the action space must be finite, and transmit power may be discretized; discretization of continuous variables inevitably leads to quantization errors), to avoid channel power discretization, the second-layer network in this invention employs a DDPG network. The DDPG network includes an Actor network and a Critic network. The Actor network outputs the allocated power, while the Critic network evaluates the actions of the Actor network and updates the parameters in the Actor network. The second-layer network, through the Actor network in the DDPG, can output a power allocation strategy composed of deterministic power allocation actions based on the imperfect global channel state information of a time slot. Therefore, using a DQN network, DDQN network, or Dueling DQN can learn the optimal sub-channel actions more quickly. Combined with DDPG to handle continuous variable types of resources (channel power allocation), it can achieve a faster convergence rate and higher spectral efficiency compared to existing algorithms that handle non-convex optimization objectives without interruption probability constraints.

[0156] According to an embodiment of the present invention, during the training of the initial resource allocation system, step S4 includes steps S41, S42, and S43. Step S41 includes: acquiring imperfect global channel state information of the input time slot and performing the following steps: S411, the channel allocation model predicts the channel allocation strategy for the input time slot based on the imperfect global channel state information of the input time slot; updates the imperfect global channel state information of the input time slot based on the predicted channel allocation strategy; and the power allocation model predicts the power allocation strategy for the input time slot based on the updated imperfect global channel state information of the input time slot; the predicted channel allocation strategy and power allocation strategy of the input time slot interact with the wireless network to obtain the imperfect global channel state information of the next time slot of the input time slot; and the channel allocation model predicts the channel allocation strategy of the next time slot of the input time slot based on the imperfect global channel state information of the next time slot of the input time slot. The allocation strategy is as follows: S411: Update the channel allocation strategy for the next time slot of the input time slot based on the channel allocation strategy of the input time slot; S412: Calculate the spectral efficiency bonus for the input time slot based on the channel allocation strategy and power allocation strategy of the input time slot; S413: Store a channel allocation experience as a channel selection replay pool using the imperfect global channel state information of the input time slot, the channel allocation strategy of the input time slot, the spectral efficiency bonus of the input time slot, and the imperfect global channel state information of the next time slot of the input time slot; Store a power allocation experience as a power selection replay pool using the updated imperfect global channel state information of the input time slot, the power allocation strategy of the input time slot, the spectral efficiency bonus of the input time slot, and the updated imperfect global channel state information of the next time slot of the input time slot. Step S42 includes: Using the imperfect global channel state information of the next time slot of the previous input time slot as the new imperfect global channel state information of the input time slot. Step S43 includes: Updating the initial resource allocation system parameters based on the channel allocation experience in the channel selection replay pool and the power allocation in the power selection replay pool until convergence. It should be noted that the spectral efficiency reward for the input time slot is calculated based on the channel allocation strategy and power allocation strategy. This spectral efficiency reward represents the overall contribution of channel allocation and power allocation to the optimization objective. This allows the channel allocation model and the power allocation model to share the same reward function and work together to maximize the spectral efficiency of the wireless network.

[0157] According to one embodiment of the present invention, in step S43, the parameters of the channel allocation model are updated when there is a channel allocation experience in the channel selection replay pool; the parameters of the power allocation model are updated when there is a power allocation experience in the power selection replay pool.

[0158] According to one embodiment of the present invention, in step S43, after the channel allocation experience in the channel selection replay pool reaches a preset number of experience points, the parameters of the channel allocation model are updated multiple times until convergence. Each update involves randomly sampling multiple channel allocation experiences from the channel selection replay pool, and updating the parameters of the channel allocation model using gradient descent based on the sampled channel allocation experiences. Similarly, after the power allocation experience in the power selection replay pool reaches a preset number of experience points, the parameters of the power allocation model are updated multiple times until convergence. Each update involves randomly sampling multiple power allocation experiences from the power selection replay pool, and updating the parameters of the power allocation model using gradient descent based on the sampled power allocation experiences. It should be noted that when the number of experience entries in the channel selection replay pool reaches a threshold, the channel allocation experience in the channel selection replay pool is stored in the channel selection replay pool in a manner that replaces the channel allocation experience stored first in the current channel selection replay pool (i.e., first-in, first-out). When the number of experience entries in the power selection replay pool reaches a threshold, the power allocation experience in the power selection replay pool is stored in the power selection replay pool in a manner that replaces the power allocation experience stored first in the current channel selection replay pool (i.e., first-in, first-out). Specifically, the practice of using a preset number of experience entries in the channel selection replay pool or power selection replay pool before random sampling can accelerate the convergence speed of the channel allocation model or power allocation model. Setting a threshold for the number of experience entries can reduce the hardware requirements for model training. The first-in, first-out (FIFO) approach for storing experience in the channel selection replay pool or power selection replay pool allows newly generated better experience to effectively replace relatively poor experience, thus ensuring that the experience in the channel selection replay pool or power selection replay pool is in an optimal storage state during sampling, thereby further accelerating the convergence speed of the channel allocation model or power allocation model.

[0159] To better train the initial resource allocation system, during the training process, imperfect global channel state information of the input time slots is first acquired. According to one embodiment of the present invention, during the training of the initial resource allocation system, the imperfect global channel state information of the input time slots includes imperfect global channel state information (sometimes referred to as a state set) of multiple links selecting different sub-channels in the input time slots. The imperfect global channel state information of one link after selecting different sub-channels in the input time slot constitutes a state set. The imperfect global channel state information of the multiple links after selecting different sub-channels in the input time slots is as follows:

[0160]

[0161] in,

[0162]

[0163] in, This represents the state set of the k-th link when selecting the n-th sub-channel in time slot t. This represents the independent channel gain of the k-th link when selecting the n-th sub-channel in time slot t, given the existence of channel estimation error. This represents the channel power of the k-th link when selecting the n-th sub-channel in time slot t. This represents the identifier value of the k-th link after selecting the nth sub-channel in time slot t-1. This represents the power of the k-th link when selecting the n-th sub-channel in time slot t-1. This represents the spectral efficiency of the k-th link in time slot t-1. This represents the estimated small-scale fading component corresponding to the nth sub-channel selected by the kth link in time slot t. The ratio of the total interference power to the total interference power, ranked across all channels. This indicates co-channel interference when the k-th link selects the n-th sub-channel in time slot t, using the sub-channel allocation scheme and power allocation scheme of the previous time slot. k′ represents other links different from k. The variance representing the channel estimation error. It is a large-scale fading component that takes into account both shadow fading and geometric decay. This indicates that the mean is 0 and the variance is .

[0164] The complex Gaussian distribution. According to one embodiment of the present invention, when the imperfect global channel state information of the input time slot is a set of states, the first layer network is configured as a channel allocation model with the same number of links, and the second layer network is configured as a power allocation model with the same number of links, wherein one channel allocation model and one power allocation model process one set of states. According to one embodiment of the present invention, when the imperfect global channel state information of the input time slot is a set of states, the first layer network is configured as a channel allocation model, and the second layer network is configured as a power allocation model, wherein one channel allocation model and one power allocation model process each set of states in the state set sequentially. Configuring the first layer network as a channel allocation model with the same number of links and the second layer network as a power allocation model with the same number of links can improve the processing speed of the initial resource allocation model. According to an embodiment of the present invention, in step S411, when the imperfect global channel state information of the input time slot is a set of states, the imperfect global channel state information of the input time slot is updated based on the predicted channel allocation strategy in the following manner: based on the channel allocation strategy predicted by the channel allocation model, the set of states corresponding to the predicted channel allocation strategy is selected from the set of states as the imperfect global channel state information of the updated input time slot.

[0165] It is important to note that the selection of the state set is crucial to the training effect of the initial resource allocation system. The state set should reflect the characteristics of imperfect CSI, meaning that relevant channel state information that reflects imperfect CSI must be selected as elements of the state set. This avoids unnecessary channel state information in the imperfect global channel state information, thereby improving the training effect of the initial resource allocation model. Among these, the variance of the channel estimation error, the estimated channel gain (independent channel gain), and the ranking of the ratio of the estimated small-scale fading component to the total interference power for each link in a given time slot across all channels are the key features that best reflect the channel state information corresponding to imperfect CSI. Therefore, this invention introduces these information into the state set and designs corresponding rewards to enable the resource allocation model system to achieve better gains even with the introduction of channel estimation errors.

[0166] According to one embodiment of the present invention, the spectral efficiency bonus is calculated in the following manner:

[0167]

[0168] in,

[0169]

[0170] in, ε represents the spectral efficiency of the k-th link when selecting the n-th sub-channel in time slot t. out This represents the expected probability of interruption. Let φ represent the scheduling spectral efficiency of the k-th link in selecting the n-th sub-channel in time slot t, where φ is the interference weighting coefficient, and k′ represents other links different from k. This indicates the external interference of the k-th link selecting the n-th sub-channel in time slot t. This represents the spectral efficiency of link k′ in the nth subchannel of time slot t, where there is no interference from the kth link. This represents the spectral efficiency of the k′-th link when selecting the nth sub-channel in time slot t. It should be noted that by defining a weighting coefficient for interference, the variance of the reward function can be reduced; preferably, φ = 1.

[0171] To better explain the parameter update process of the initial resource allocation system of this invention, the following description uses an initial resource allocation system composed of a Dueling DQN network and a DDPG network, and the parameter update process employing random sampling from the channel selection replay pool and the power selection replay pool, as an example. It should be noted that selecting the Dueling DQN network can distinguish whether the spectral efficiency gain depends on the sub-channel actions taken, or simply because the input state set is good enough to effectively complete the channel allocation task, thereby further improving the performance of the initial resource allocation system after convergence.

[0172] like Figure 2 As shown, the initial resource allocation system consists of a Dueling DQN network, a DDPG (including an Actor network), and a Critic Network. The Actor Network and the Critic Network together form the DDPG network. It should be noted that the channel allocation strategy described above corresponds to the channel selection action for each sub-channel on each link in each time slot, while the power allocation strategy corresponds to the power allocation action for the selected sub-channel on each link in each time slot. The following will illustrate the process of generating one channel allocation experience and one power allocation experience using the initial resource allocation system as an example: the state set of the k-th link in time slot t. Inputting into the Dueling DQN network, the Dueling DQN network predicts the channel selection action of the sub-channels. Channel selection action in the predicted sub-channel Afterwards, based on renew Get the updated state set (Right now, middle Corresponding state set ) and As input to the DDPG network, the Actor network in the DDPG network is based on Predict the power allocation action of the sub-channel The base station performs two actions sequentially at the beginning of time slot t. and To determine its associated sub-channel and the transmit power of that sub-channel, the base station sequentially performs two actions ( and After that, and after interacting with the wireless network environment (i.e., the ultra-dense network environment under imperfect CSI), the state set for the next time slot t+1 will be generated. And channel selection actions of sub-channels predicted by the channel allocation model. And the channel allocation actions of sub-channels predicted by the power allocation model. Calculate spectral efficiency reward Inputting the Dueling DQN network will predict the sub-channel selection action. Channel selection action in the predicted sub-channel Afterwards, based on renew Get the updated state set (Right now, middle Corresponding state set ), among which, This channel allocation experience is stored in the channel selection replay pool. Figure 2 (as shown) and will This power allocation experience is stored in the power selection playback pool. Figure 2 (As shown in the diagram). This allows for the continuous generation of channel allocation and power allocation experience during the training of the initial resource allocation system.

[0173] During parameter updates, the parameters of the Dueling DQN network are updated until convergence based on experience from the channel selection replay pool, and the parameters of the DDPG network are updated until convergence based on experience from the power selection replay pool. It should be noted that the Critic network in DDPG is only used during training. In actual deployment, only the Actor network is needed for power allocation. For example, taking the initial resource allocation system consisting of the Dueling DQN network and the DDPG network as an example, the final resource allocation system is: the Dueling DQN network and the Actor network trained until convergence.

[0174] Since training the Dueling DQN and DDPG networks to convergence is a process well-known to those skilled in the art, the specific convergence conditions will not be elaborated here. Instead, the following explanation will focus on the initial resource allocation system composed of the Dueling DQN and DDPG networks, illustrating the parameter update process of the initial resource allocation system. Preferably, the initial parameter update process of the initial resource allocation system is explained using two methods: first, randomly sampling from the channel selection replay pool to obtain multiple channel allocation experiences and calculating gradients based on these experiences to update the parameters of the Dueling DQN network; and second, randomly sampling from the power selection replay pool to obtain multiple power allocation experiences and calculating gradients based on these experiences to update the parameters of the DDPG network.

[0175] Preferably, a channel allocation experience set B1 is obtained by randomly sampling channel allocation experience from the channel selection replay pool, and the gradient is calculated using the following rules to update the parameters of the Dueling DQN network:

[0176]

[0177] in,

[0178]

[0179] Where, θ c Let represent the trainable parameters of the hidden layer in the Dueling DQN network, β represent the trainable parameters of the fully connected layer with value function V, χ represent the trainable parameters of the fully connected layer with dominance function A, and B1 represent the empirical set of randomly sampled channel allocations. |B1| represents a channel allocation experience in B1, and |B1| represents the number of experience entries in the channel allocation experience set. This represents the target value of Dueling DQN. Let Q' represent the set of state information for time slot t+1. γ represents the Q-function value of time slot t, and γ′ represents the discount coefficient in the Dueling DQN network. These represent the parameters of the target network in the Dueling DQN network. This indicates the sub-channel selection action in time slot t+1. Indicates being in The value function of time, Indicates the state Choose action Advantage function value, Indicates the state Choose action The dominant function value, |A| represents The number of summation calculations.

[0180] For updates to the DDPG network, the DDPG network uses a neural network to approximate the action value function Q(s,a) (Critic network) and the action function u. θ (s)(Actor network).

[0181] Preferably, in order to update the network parameters θ of the Critic network Q The temporal difference (TD) error method is used to obtain the power allocation experience set B2 by randomly sampling power allocation experience from the power selection playback pool, and the parameters of the Critic network are updated by calculating the minimum mean square error under the following rules:

[0182]

[0183] in,

[0184] Where B2 represents the empirical set of power allocation from random sampling, This represents a power allocation experience in the power allocation experience set. This represents the Q-function of the Critic network at the input. The function value at time, (Determined by the Actor network activation function) represents the channel selection action of the k-th link in the Actor network passing through the fixed sub-channel of the Dueling DQN output in time slot t. The output power represents the determined power allocation action, |B2| represents the number of empirical data points in the randomly sampled power allocation empirical set, y′ represents the target value of the DDPG network, and γ represents the discount coefficient in the DDPG network. This represents the Q-function of the Critic network at the input. The function value at that time.

[0185] The parameters of the Actor network are updated based on the collected power allocation experience set B2 and the gradient is calculated using the following rules:

[0186]

[0187] in, This indicates that the Q function of the Critic network is given by the input... The function value at time, This indicates that the Q function of the Critic network is given by the input... The function value at that time.

[0188] It should be noted that for the initial allocation network composed of DQN network or DDQN network and DDPG network, the update of DQN network or DDQN network is mainly different from the calculation method of Q function of Dueling DQN network. The Q function and update process of DQN network or DDQN network are known to those skilled in the art, and the parameter update process of DQN network or DDQN network will not be described in detail here.

[0189] Furthermore, based on the aforementioned method for constructing a wireless network resource allocation system, this invention also provides a wireless network resource management method, such as... Figure 3 As shown, a wireless network resource management method includes mathematically describing the entire wireless network environment by establishing a non-convex optimization objective with outage probability constraints in a method for constructing a wireless network resource allocation system, and forming a non-convex optimization objective with outage probability constraints based on this (i.e., Figure 3 The model is established in the middle, and then the non-convex optimization objective with interruption probability constraints is transformed into a non-convex optimization objective without interruption probability constraints (i.e. Figure 3 The parameter transformation in the process is performed using the initial resource allocation system for solution training (i.e., Figure 3 (A two-layer network architecture in the present invention) When the initial resource allocation system converges, the resource allocation scheme obtained by the resource allocation system can be used to allocate wireless network resources in the wireless communication system. According to an embodiment of the present invention, a wireless network resource management method includes: T1, obtaining the wireless network state of the wireless communication system in the previous time slot; T2, based on the wireless network state of the previous time slot obtained in step T1, predicting the resource allocation strategy for the next time slot using the wireless resource allocation system obtained by the above-described method for constructing a wireless network resource allocation system; T3, allocating wireless network resources in the wireless communication system based on the resource allocation strategy for the next time slot obtained in step T2.

[0190] Based on a wireless network resource management method, this invention also provides a wireless communication system. The system includes multiple base stations, each base station including a wireless resource management unit. The wireless resource management unit is configured to allocate wireless network resources within the base station using the aforementioned wireless network resource management method. Specifically, the wireless network resource allocation system configured in the wireless resource management unit employs a centralized training mode during the training phase, i.e., it uses a set of states from a multi-cell, multi-user network scenario for training. During the deployment phase, the trained wireless network resource allocation system is distributed and deployed to each base station. This improves the efficiency of the wireless communication system in allocating wireless network resources.

[0191] III. Experimental Verification

[0192] To better illustrate the technical effects of the present invention, the following simulation experiments were conducted for verification.

[0193] First, let's introduce the simulation parameter settings: The wireless network scenario is set as a multi-cell, multi-user downlink network, where K links are distributed across M cells and share N orthogonal sub-channels, meaning each cell has M / K users; for cell i, the base station BS is located at the center of cell i and serves M / K users randomly distributed within the cell; the large-scale path loss is calculated as 128.1 + 37.6 log... 10 (d) Calculation, where d is the distance from transmitter to receiver in kilometers; the upper limit of the signal-to-interference-plus-noise ratio (SINR) received by the user is set to 30 dB, and the noise power σ 2 Setting the value to -114dBm, the optimization problem requires maximizing spectral efficiency. The initial allocation network in this invention employs a Dueling DQN network and a DDPG network. Both the Dueling DQN and DDPG networks have three hidden layers with 200, 200, and 100 neurons respectively. In addition to the above settings, the detailed simulation parameters are shown in Table 1.

[0194] Table 1

[0195] Simulation parameters value Simulation parameters value Cell radius 200m Link power threshold 38dB Interruption probability 0.1 Channel estimation error variance 0.1 Time slot interval 20ms Interference weighting coefficient 1

[0196] In this simulation experiment, the method for constructing a wireless network resource allocation system provided by this invention (hereinafter referred to as the algorithm proposed in this patent) was used together with some other baseline algorithms to conduct three sets of comparative experiments (testing the convergence, generalization ability and spectral efficiency performance of the training phase, and the spectral efficiency performance of the algorithm under different numbers of sub-channels). The baseline algorithms include random algorithms and the random, FP (fractional programming algorithm) algorithm recorded in the background art reference [5], the joint learning algorithm recorded in the reference [7], and the distributed learning algorithm recorded in the reference [9]. Specifically, the fractional programming algorithm includes randomly allocated sub-channels and power values ​​and traditional model-driven algorithms, which have high computational complexity. The joint learning algorithm uses a DQN network to optimize two variables, namely sub-channels and power. The distributed learning algorithm uses DQN to optimize sub-channels and DDPG to optimize power under perfect CSI.

[0197] During the training phase, the algorithm proposed in this patent and the four baseline algorithms mentioned above undergo a training process comprising 20 episodes, each containing 2000 time slots. That is, the algorithm stops training and parameter updates after a fixed 2000 time steps within an episode. At the beginning of each episode, a new user distribution is set and parameters such as the learning rate are reset to ensure convergence of the proposed algorithm and the four baseline algorithms. The following section will compare and explain the data corresponding to the proposed algorithm and the four baseline algorithms.

[0198] To test the convergence of the proposed algorithm and the four baseline algorithms mentioned above during the training phase, only parameters such as the learning rate were reset in each episode without updating the user distribution. When the number of users is 25 and the number of base stations and sub-channels is 5, the convergence performance of the proposed algorithm and the four baseline algorithms is as follows: Figure 4 As shown, the algorithm proposed in this patent and the four baseline algorithms mentioned above converge to around 4.0 bps / Hz in each episode (except for the stochastic algorithm and the fractional programming algorithm). Among the convergent algorithms, the algorithm proposed in this patent converges faster than the learning-based baseline algorithms (joint learning algorithm and distributed learning algorithm). Furthermore, the spectral efficiency achievable by the four baseline algorithms mentioned above is lower than that of the algorithm proposed in this patent, and when the number of iterations is small, i.e., 5000-6000 iterations, the spectral efficiency of the four baseline algorithms is significantly lower than that of the proposed scheme. Therefore, the algorithm proposed in this patent has a significant advantage in terms of convergence rate.

[0199] In the generalization ability and spectral efficiency performance tests, since the channel estimation error in real dynamic wireless communication scenarios changes over time, it is impractical to track the environment through frequent online training. Therefore, the algorithm's generalization ability is crucial in constantly changing environments. The training model that achieves the best results when the channel estimation error variance is set to 0.01 in the simulation parameter settings described above is shown. Tests were conducted under different conditions, and the relationship between the spectral efficiency achievable by the algorithm proposed in this patent and the four baseline algorithms mentioned above and the channel estimation error variance is as follows: Figure 5 As shown, the performance of the stochastic algorithm and the fractional programming algorithm degrades with increasing channel estimation error; the spectral efficiency of the joint learning algorithm, the distributed learning algorithm, and the algorithm proposed in this patent remains almost unchanged with changes in channel estimation error. Therefore, the algorithm proposed in this patent has stronger generalization ability and can achieve higher spectral efficiency under different channel estimation errors.

[0200] In the performance test of the algorithm's spectral efficiency under different numbers of sub-channels, with a channel estimation error variance of 0.1, the spectral efficiency performance of the algorithm proposed in this patent and the four baseline algorithms mentioned above under different numbers of sub-channels is as follows: Figure 6As shown, the average spectral efficiency per link achievable by the algorithm proposed in this patent and the four baseline algorithms mentioned above gradually increases with the increase of the number of sub-channels. However, at a fixed number of sub-channels, the algorithm proposed in this patent can achieve a higher spectral efficiency than the four baseline algorithms mentioned above. That is, the algorithm proposed in this patent is easier to scale in multi-cell networks and performs better with the increase of the number of sub-channels.

[0201] In summary, existing learning-based methods are not designed for imperfect CSI, and channel estimation errors are inherent and cannot be completely eliminated in real-world communication environments. Therefore, directly applying existing learning-based algorithms in imperfect CSI environments results in poor optimization performance (i.e., communication performance) and slow convergence speed. The proposed method incorporates the estimated channel gain and its corresponding error (represented by the variance of the channel estimation error) into the state set of the initial resource allocation system, and designs a corresponding reward (e.g., spectral efficiency reward). This allows the initial resource allocation system to achieve better gains even with channel estimation errors, resulting in a faster convergence rate than the baseline algorithm and significantly improved spectral efficiency performance. In other words, the proposed method for constructing a wireless network resource allocation system is more suitable for real-world dynamic wireless communication environments (imperfect CSI).

[0202] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0203] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0204] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.

[0205] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for constructing a wireless network resource allocation system, wherein the wireless network resource allocation system is used to obtain a wireless network resource allocation strategy based on the wireless network status, characterized in that... The method includes: S1. Obtain the non-convex optimization objective with interruption probability constraints corresponding to the wireless communication requirements under an imperfect global channel state information environment; wherein, the wireless communication requirement is to maximize the spectral efficiency of the wireless network, and the non-convex optimization objective with interruption probability constraints is: in, in, Indicates the wireless network in time slots Average spectral efficiency Represents the total number of links Indicates the total number of sub-channels. This represents the set of sub-channel indices. Indicates the first Link in time slot Select the Scheduling spectral efficiency of individual sub-channels. Indicates the first Link in time slot Select the Maximum spectral efficiency of each sub-channel Indicates the first Link in time slot Select the Estimated small-scale fading components of each sub-channel This indicates the estimation of small-scale fading components. Under the condition The probability, Indicates the first Link in time slot Select the The power of each sub-channel Indicates all The collection of power, Indicates the first Link in time slot Select the The identifier value after each sub-channel Indicates all A set of identifier values This represents the expected probability of interruption. Represents the power threshold of the link, a constraint. This indicates the estimation of small-scale fading components. Under the condition that any link in the time slot The probability of being interrupted after selecting any sub-channel must be less than the expected interruption probability. This means that the transmit power on each link cannot exceed the link's power threshold, constraining... This means that each link can only select one subchannel in each time slot; S2. Transform the non-convex optimization objective obtained in step S1 to obtain a non-convex optimization objective without interruption probability constraints; wherein, parameter transformation is used to transform the non-convex optimization objective to obtain a non-convex optimization objective without interruption probability constraints: in, in, The wireless network represented is in the time slot after parameter transformation. Average spectral efficiency; S3. Obtain imperfect global channel state information of the wireless network; S4. Using the non-convex optimization objective in step S2 as the training objective and the imperfect global channel state information in step S3 as input, the initial resource allocation system is trained to convergence using reinforcement learning. The initial resource allocation system is a system built based on an agent and used to generate a set of actions based on the wireless network state. The set of actions includes a channel allocation strategy and a power allocation strategy.

2. The method according to claim 1, characterized in that, The initial resource allocation system includes: A channel allocation model is used to predict the channel allocation strategy for a time slot based on the imperfect global channel state information of that time slot. It can be configured as a DQN network, a DDQN network, or a Dueling DQN network. A power allocation model is used to predict the power allocation strategy for a time slot based on imperfect global channel state information for that time slot, and it is configured for a DDPG network.

3. The method according to claim 2, characterized in that, S4 includes: S41. Obtain the imperfect global channel state information of the input time slot and perform the following steps: S411. The channel allocation model predicts the channel allocation strategy for the input time slot based on the imperfect global channel state information of the input time slot, updates the imperfect global channel state information of the input time slot based on the predicted channel allocation strategy, and the power allocation model predicts the power allocation strategy for the input time slot based on the updated imperfect global channel state information of the input time slot; the predicted channel allocation strategy and power allocation strategy of the input time slot are interacted with the wireless network to obtain the imperfect global channel state information of the next time slot of the input time slot, and the channel allocation model predicts the channel allocation strategy of the next time slot of the input time slot based on the imperfect global channel state information of the next time slot of the input time slot, and updates the channel allocation strategy of the next time slot of the input time slot based on the channel allocation strategy of the next time slot of the input time slot; S412. Calculate the spectral efficiency bonus of the input time slot based on the channel allocation strategy and power allocation strategy of the input time slot; S413. Store a channel allocation experience as a channel selection replay pool using the imperfect global channel state information of the input time slot, the channel allocation strategy of the input time slot, the spectral efficiency bonus of the input time slot, and the imperfect global channel state information of the next time slot of the input time slot; store a power allocation experience as a power selection replay pool using the updated imperfect global channel state information of the input time slot, the power allocation strategy of the input time slot, the spectral efficiency bonus of the input time slot, and the updated imperfect global channel state information of the next time slot of the input time slot. S42. The imperfect global channel state information of the next time slot of the previous input time slot is used as the imperfect global channel state information of the new input time slot. S43. Update the initial resource allocation system parameters based on the channel allocation experience in the channel selection replay pool and the power allocation in the power selection replay pool until convergence.

4. The method according to claim 3, characterized in that, In step S43, the parameters of the channel allocation model are updated when there is a channel allocation experience in the channel selection replay pool; the parameters of the power allocation model are updated when there is a power allocation experience in the power selection replay pool.

5. The method according to claim 3, characterized in that, In S43; Once the channel allocation experience in the channel selection replay pool reaches a preset number of experiences, the parameters of the channel allocation model are updated multiple times until convergence. During each update, multiple channel allocation experiences are randomly sampled from the channel selection replay pool, and the parameters of the channel allocation model are updated based on the sampled channel allocation experiences using gradient descent. Once the power allocation experience in the power selection replay pool reaches a preset number of experiences, the parameters of the power allocation model are updated multiple times until convergence. During each update, multiple power allocation experiences are randomly sampled from the power selection replay pool, and the parameters of the power allocation model are updated based on the sampled power allocation experiences using gradient descent.

6. The method according to any one of claims 3-5, characterized in that, The imperfect global channel state information of the input time slot mentioned in step S41 includes imperfect global channel state information of multiple links selecting different sub-channels in the input time slot: in: , , Indicates the first Link in time slot Select the The state set of each sub-channel This indicates that the first... Link in time slot Select the Independent channel gain of each sub-channel Indicates the first Link in time slot Select the Channel power of each sub-channel Indicates the first Link in time slot Select the The identifier value after each sub-channel, Indicates the first Link in time slot Select the The power of each sub-channel Indicates the first Link in time slot The corresponding spectral efficiency, Indicates the first Link in time slot Select the The estimated small-scale fading components corresponding to each sub-channel The ratio of the total interference power to the total interference power, ranked across all channels. Indicates the first Link in time slot Select the When dealing with multiple sub-channels, co-channel interference is addressed using the sub-channel allocation scheme and power allocation scheme from the previous time slot. Indicates different Other links, The variance representing the channel estimation error. It is a large-scale fading component that takes into account both shadow fading and geometric decay. This indicates that the mean is 0 and the variance is . The complex Gaussian distribution.

7. The method according to claim 6, characterized in that, The spectral efficiency bonus is calculated as follows: in, in, Indicates the first Link in time slot Select the The spectral efficiency corresponding to each sub-channel. This represents the expected probability of interruption. Indicates the first Link in time slot Select the Scheduling spectral efficiency of individual sub-channels. It is the interference weighting coefficient. Indicates different Other links, Indicates the first Link in time slot Select the External interference of each subchannel Indicates in time slot No. The sub-channel has no first Link interference Spectral efficiency, Indicates the first Link in time slot Select the The spectral efficiency corresponding to each sub-channel.

8. The method according to claim 7, characterized in that, The DDPG network includes an Actor network and a Critic network. The final resource allocation system is: a DQN network, a DDQN network, or a Dueling DQN network and an Actor network trained to convergence.

9. A wireless network resource management method, characterized in that, The method includes: T1. Obtain the wireless network status of the wireless communication system in the previous time slot; T2. Based on the wireless network state of the previous time slot obtained in step T1, the resource allocation system trained by the method described in any one of claims 1-8 is used to predict the resource allocation strategy for the next time slot. T3. Allocate wireless network resources in the wireless communication system based on the resource allocation strategy obtained in step T2 for the next time step.

10. A wireless communication system, the system comprising multiple base stations, characterized in that, Each base station includes a radio resource management unit, which is configured to allocate radio network resources in the base station using the method described in claim 9.

11. A computer-readable storage medium, characterized in that, It stores a computer program that can be executed by a processor to implement the steps of the method according to any one of claims 1 to 9.

12. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the electronic device to perform the steps of the method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Cognitive NOMA network stubborn resource allocation method based on energy efficiency

    CN110417496A

  • Power distribution method of NOMA downlink system under non-ideal channel state information

    CN113056015A