A Cognitive Cell-Free Power Allocation Method Based on the MAMFSAC Algorithm

Through the multi-agent reinforcement learning method using the MAMFSAC algorithm in a cognitive cellular-free network, the problem of inaccurate spectrum perceived by secondary users is solved, and efficient power allocation and network resource utilization are achieved.

CN118265144BActive Publication Date: 2025-06-20XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410298579.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-06-20
Estimated Expiration
2044-03-15

AI Technical Summary

Technical Problem

In the existing cognitively cellular cell-free network, secondary users cannot accurately perceive the spectrum void of the primary user, resulting in false alarms or missed detection, affecting network throughput and user service performance.

Method used

The multi-agent reinforcement learning method based on the MAMFSAC algorithm is adopted to construct a cognitive cell-free system model, and the POMDP problem model is solved through the multi-agent average field soft actor critic algorithm to obtain the optimal power allocation coefficient.

Benefits of technology

Solve non-convex power distribution problems in a very short time, improve the performance and scalability of the algorithm, and ensure the signal reception rate of secondary users and the effective utilization of network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118265144B_ABST
    Figure CN118265144B_ABST
Patent Text Reader

Abstract

The present invention discloses a power allocation method for a cognitive cell-free system based on the MAMFSAC algorithm, including: constructing a cognitive cell-free system model, which includes a primary transmitter, primary users, M secondary transmitters, and N secondary users; obtaining an optimal objective model of the signal reception rate of each secondary user according to the cognitive cell-free system model; creating M agents according to the M secondary transmitters, and modeling the optimal objective model as a POMDP problem model; constructing the weight coefficient of the weighted mean field of the multi-agent mean field soft actor-critic algorithm; using the multi-agent mean field soft actor-critic algorithm to solve the POMDP problem model to obtain the optimal power allocation coefficient. By using a data-driven algorithm of multi-agent reinforcement learning, the present invention can solve the original non-convex problem in an extremely short time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power distribution, and particularly relates to a cognitive cell-free power distribution method based on the MAMFSAC algorithm. Background Art

[0002] A cognitive cell-free massive network system consists of an in-network secondary cell-free network and an out-of-network primary user (PU) network. As the number of in-network users gradually increases, the secondary network suffers from a shortage of spectrum resources. To solve this problem, secondary users (SUs) need to sense the spectrum holes of primary users to opportunistically access the primary user network. In traditional power distribution literature based on cognitive cell-free systems, it is usually assumed that SUs can accurately sense the spectrum holes of primary users. However, in actual scenarios, SUs cannot accurately sense the spectrum holes of primary users.

[0003] When there are errors in the spectrum sensing results, false alarms or missed detections may occur. False alarms will cause SUs to lose the opportunity to access the spectrum and reduce network throughput. Missed detections will cause SUs to access unavailable spectra and may interfere with PUs. This may penalize SUs and even prohibit them from using the authorized spectra of primary users again. In addition, within the secondary cell-free network, the cell is liberated from the constraints of the inherent cellular shape and the cell boundary. Compared with cellular networks, the cell-free network allows multiple base stations to serve a user simultaneously, greatly improving the service performance for edge users. However, this also brings interference between users. To comprehensively consider the interference between users within the secondary cell-free network and coordinate the interference between the primary and secondary networks, a power distribution method for cognitive cell-free systems is crucial.

[0004] The existing maximum-minimum fairness problem based on traditional cell-free networks is non-deterministic polynomial hard (NP-hard) and non-convex. In the literature, equivalent convex problems of these optimization problems are obtained through mathematical derivation. The conventional approach is to adopt the idea of iterative optimization, but this method has a very high time complexity and cannot meet the requirement of making power distribution results in a very short time in actual scenarios. In recent years, with the rise of deep learning (DL) technology, model-free reinforcement learning, as a method driven by data and dynamically updating the network through interaction with the environment, has been favored by researchers.

[0005] On the one hand, in the literature related to cognitive cell-free systems, it is mostly considered that secondary users can accurately sense the spectrum of primary users. However, in the actual application process, false alarms and missed detections are inevitable. Therefore, for the actual scenario, the false alarm and missed detection probabilities need to be introduced into the system optimization model. On the other hand, in the literature related to traditional cognition, for the optimization problem of power control, iterative methods are generally used to solve the problem, and the high complexity of this method cannot cope with the rapidly changing channel state in practice. Although there are also literatures using reinforcement learning algorithms, they generally stay at two algorithms: Deep Q-learning Network (DQN) and Deep Deterministic Policy Gradient (DDPG). They have weak exploration ability in terms of algorithm performance and are very sensitive to hyperparameters, and the weak exploration ability leads to large fluctuations in the curve. Summary of the Invention

[0006] To solve the above problems existing in the prior art, the present invention provides a cognitive cell-free power allocation method based on the MAMFSAC algorithm. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] The present invention provides a cognitive cell-free power allocation method based on the MAMFSAC algorithm, including:

[0008] S1: Construct a cognitive cell-free system model, where the cognitive cell-free system model includes a primary transmitter, a primary user, M secondary transmitters, and N secondary users;

[0009] S2: Obtain an optimal target model of the signal reception rate of each secondary user according to the cognitive cell-free system model;

[0010] S3: Create M agents according to the M secondary transmitters to form a multi-agent system, and model the optimal target model as a POMDP problem model;

[0011] S4: Construct the weight coefficient of the weighted mean field of the multi-agent mean field soft actor-critic algorithm;

[0012] S5: Use the multi-agent mean field soft actor-critic algorithm to solve the POMDP problem model to obtain the optimal power allocation coefficient.

[0013] In an embodiment of the present invention, the S2 includes:

[0014] S2.1: Obtain the signal y received by the nth secondary user SU n when the primary user does not occupy the frequency bandn0 :

[0015]

[0016] Among them, represents the desired signal, η mn is expressed as the power coefficient allocated to the m-th secondary transmitter ST, where m ∈ [1, M] m to the n-th secondary user SU, where n ∈ [1, N] n , h mn is expressed as the downlink channel coefficient between ST m and SU n , h mn = {h mn1 ,..., h mnk ,..., h mnK}, h mnk is the downlink channel coefficient of the k-th antenna between ST m and SU n , W mn is the precoding matrix in the conjugate beamforming mode between ST m and SU n , q n represents the data sent to the n-th secondary user, represents the interference between secondary users, q i represents the data sent to the i-th secondary user, w n is Gaussian white noise with a mean of 0 and a variance of σ 2 ;

[0017] S2.2: Obtain the signal-to-interference-plus-noise ratio (SINR) of the n-th secondary user SU n when the primary user does not occupy the frequency band:

[0018]

[0019] Among them, P max represents the maximum transmission power of each secondary emitter;

[0020] S2.3: Obtain the received signal y n of the n-th secondary user SU n1 when the primary user occupies the frequency band:

[0021]

[0022] Among them, P t represents the transmission power of the primary transmitter, h pn represents the channel coefficient between the primary transmitter and the n-th secondary user SU n ;

[0023] S2.4: Obtain the nth secondary user SU n The signal-to-interference-plus-noise ratio corresponding to the primary user occupying the frequency band:

[0024]

[0025] S2.4: Obtain the secondary user SU according to the signal-to-interference-plus-noise ratio when the primary user occupies the frequency band and the signal-to-interference-plus-noise ratio when the primary user does not occupy the frequency band n The signal reception rate R of n is:

[0026] R n =(1 - P f )P r (H0)log2(1 + SINR n0 )+(1 - P d )P r (H1)log2(1 + SINR n1 )

[0027] where P r (H0) is the probability that the primary user does not occupy the frequency band, P r (H1) is the probability that the primary user occupies the frequency band, P f =P r {judgment as H1|H0} represents the false alarm probability, P d =P r {judgment as H1|H1} represents the detection probability;

[0028] S2.5: Construct an optimal objective model for the signal reception rate of the secondary user:

[0029]

[0030] where h mp represents the channel coefficient between the primary user and the mth secondary transmitter, P Imax represents the maximum interference threshold of the primary user.

[0031] In an embodiment of the present invention, the S3 includes:

[0032] S3.1: Create the M secondary transmitters into M agents, and construct the cognitive cell-free system model into (S, O, A, R, P), where S = {s1,..., s m ,..., s M} represents the state set of all agents, O = {o1,..., o m ,..., o M} represents the local observation set of all agents, A = {a1,..., a m,...,a M}, which represents the set of actions of all agents, R = {r1,..., r m ,..., r M}, which represents the set of return values of all agents, and P represents the state transition probability;

[0033] S3.2: Construct the state, action, and return value of each agent according to the power coefficient, downlink channel coefficient between the secondary transmitter and the secondary user, and the signal reception rate of the secondary user.

[0034] In an embodiment of the present invention, the S3.2 includes:

[0035] S3.21: Obtain the local observation o of the m-th agent m :

[0036] o m = {η m1 ,..., η mn ,..., η mN , h m1 ,..., h mn ,..., h mN}}

[0037] Among them, η mn represents the power coefficient allocated to the n-th secondary user SU by the m-th secondary transmitter ST m , and h n represents the downlink channel coefficient between ST mn and SU m ; n

[0038] S3.22: Obtain the set of states of all agents according to the local observation o m :

[0039]

[0040] S3.23: Obtain the action of the m-th agent:

[0041] a m = {η m1 ,..., η mn ,..., η mN}};

[0042] S3.24: Design the return value at time t as:

[0043] r(t) = min{R1,..., R n ,..., R N} + c(t),

[0044] where, min{R1,...,R n ,...,R N} represents the minimum secondary user signal reception rate when taking action a t in state s t , and c(t) represents the cost function with a penalty term.

[0045] In an embodiment of the present invention, S4 includes:

[0046] S4.1: Assume that all secondary transmitters include an N-dimensional pre-connected vector pre j :

[0047]

[0048] where, pre j (n) represents the connection situation between the j-th secondary transmitter and the n-th secondary user, L f represents the path loss, Threshold represents the path loss threshold, and pre j = {pre j (1),..., pre j (n),..., pre j (N)};

[0049] S4.2: Obtain the weight coefficient of agent j on the n-th secondary user with respect to agent i:

[0050]

[0051] where, K represents the number of antennas of the secondary transmitter, and h jnk represents the channel coefficient of the j-th secondary transmitter to the n-th secondary user on the k-th antenna;

[0052] S4.3: Obtain the weight coefficient w i (n) of agent i on the n-th secondary user:

[0053]

[0054] In an embodiment of the present invention, S5 includes:

[0055] S5.1: Create M agents, each agent includes a policy network and an evaluation network, initialize the policy network parameters θ and the agent power allocation coefficients {η1,..., η N}, and initialize the network parameters ω1, ω2,

[0056] S5.2: Obtain the experience information of all agents at different times and store it in the experience pool Among them, the experience information at different times includes the state set, partial observation set, action set, and reward value set at different times;

[0057] S5.3: Train the agent according to the experience information of the agent at different times, update the network parameters of the policy network and the evaluation network, and obtain the trained multi-agent system;

[0058] S5.4: Input the local observation data of the secondary user into the trained multi-agent system to obtain the optimal power allocation coefficient of the secondary user.

[0059] In an embodiment of the present invention, the S5.2 includes:

[0060] S5.21: Each agent obtains the local observation o at the current time t t , and input the local observation o t into the policy network to obtain the action a corresponding to the current time t t ;

[0061] S5.22: Use the action a corresponding to the current time t of each agent t to obtain the current reward value r of each agent t and the observation o at the next time t+1 ;

[0062] S5.23: Obtain the experience information (S t , O t , A t , R t , S t+1 , O t+1 ) of all agents and store it in the experience pool ;

[0063] In an embodiment of the present invention, the S5.3 includes:

[0064] S5.31: Select the experience information of the agent at different times in the experience pool as training data;

[0065] S5.32: Input the training data into the policy network and the evaluation network, calculate the weighted weight coefficient of the current action, and calculate the weighted average action according to the weighted weight coefficient;

[0066] S5.33: Calculate the loss function of entropy according to , obtain the loss function of the evaluation network according to , and obtain according to Obtain the loss function of the policy network, where E represents taking the average, denotes (S t , A t , r t , S t+1 ) is taken from the replay experience pool α is the temperature coefficient, H(π(·|S t )) is the regularization term of entropy, π(A t |S t ) represents the probability of executing action A t at state S t , represents that in the global state S t , the agent's own action is a t , and the weighted average action of the remaining agents is the Q value calculated by the evaluation network with parameter ω, represents that in the global state S t+1 , the state value calculated by the evaluation network with parameter ω - , γ represents the discount factor, π θ (A t |S t ) represents the probability of the policy network with parameter θ outputting action A t at state S t ;

[0067] S5.34: Perform backpropagation operations on the loss function of the entropy, the loss function of the evaluation network, and the loss function of the policy network to obtain the corresponding gradients and Update the parameter ω of the evaluation network i and the parameter θ of the policy network and the parameter α of the entropy regularization term.

[0068] Another aspect of the present invention provides a storage medium, in which a computer program is stored, and the computer program is used to execute the steps of the cognitive cell-free power allocation method based on the MAMFSAC algorithm described in any one of the above embodiments.

[0069] Another aspect of the present invention provides an electronic device, including a memory and a processor, in which a computer program is stored, and when the processor calls the computer program in the memory, it implements the steps of the cognitive cell-free power allocation method based on the MAMFSAC algorithm described in any one of the above embodiments.

[0070] Compared with the prior art, the beneficial effects of the present invention are:

[0071] 1. The present invention provides a cognitive cell-free power allocation method based on the MAMFSAC algorithm. By using a data-driven algorithm of multi-agent reinforcement learning, after the agents are trained, the original non-convex problem can be solved in a very short time.

[0072] 2. In large-scale scenarios, the present invention uses the mean-field theory to replace the action information of other agents in the evaluation network with an equivalent mean action, thereby reducing the original action dimension of O(MN) to O(N), greatly improving the performance and scalability of the algorithm.

[0073] 3. According to the characteristics of the cell-free network scenario, the present invention designs the channel coefficient as the weight of the weighted mean field. This weight design method fully considers the coupling relationship between the secondary transmitter and the secondary user and the small-scale influence. Compared with the traditional method that uses distance as the weight, the equivalent mean action calculated by the weight proposed in the present invention is more accurate.

[0074] The following will further elaborate on the present invention in detail with reference to the accompanying drawings and embodiments. Description of the Drawings

[0075] Figure 1 is a flowchart of a cognitive cell-free power allocation method based on the MAMFSAC algorithm provided by an embodiment of the present invention;

[0076] Figure 2 is a schematic structural diagram of a cognitive cell-free system model provided by an embodiment of the present invention;

[0077] Figure 3 is a schematic structural diagram of an evaluation network provided by an embodiment of the present invention;

[0078] Figure 4 is a schematic structural diagram of a policy network provided by an embodiment of the present invention;

[0079] Figure 5 is a schematic diagram of a cognitive cell-free power allocation method based on the MAMFSAC algorithm provided by an embodiment of the present invention;

[0080] Figure 6 is a comparison graph of the secondary user rates obtained by using different power allocation methods. Detailed Embodiments

[0081] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, with reference to the accompanying drawings and specific embodiments, elaborate in detail on the cognitive cell-free power allocation method based on the MAMFSAC algorithm proposed according to the present invention.

[0082] The foregoing and other technical contents, features and effects of the present invention will be clearly presented in the following detailed description of the specific embodiments in conjunction with the accompanying drawings. Through the description of the specific embodiments, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the predetermined purpose can be obtained. However, the attached drawings are only for reference and illustration, and are not used to limit the technical solution of the present invention.

[0083] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element.

[0084] Embodiment 1

[0085] Please refer to Figure 1 , Figure 1 , which is a flowchart of a cognitive cell-free power allocation method based on the MAMFSAC algorithm provided by an embodiment of the present invention. The power allocation method includes:

[0086] S1: Construct a cognitive cell-free system model.

[0087] Please refer to Figure 2 , Figure 2 , which is a schematic structural diagram of a cognitive cell-free system model provided by an embodiment of the present invention. There are two networks in this cognitive cell-free system model, namely a primary network composed of a primary transmitter (PT) and a primary user (PU), and a secondary cell-free network composed of M secondary transmitters (ST) and N secondary users (SU). Among them, each secondary transmitter is equipped with K antennas, and the secondary users are single-antenna. The secondary cell-free network can sense the frequency band of the primary user and access the primary user frequency band when the sensing result is idle. Assuming that the channel occupancy is represented by H1 and the channel idle is represented by H0, then under non-ideal sensing, the false alarm probability P f =P r {judgment as H1|H0}, that is, the probability that the channel is originally idle but misjudged as occupied, and the detection probability Pd = P r {Judgment is H1|H1}, that is, the probability that the channel is originally occupied and is successfully judged as occupied.

[0088] Since the secondary cell-free network only accesses when it senses idle, there are two probabilities during access. One is P r {Judgment is H0|H0}, expressed as 1 - P with the false alarm probability f , and the other is P r {Judgment is H0|H1}, expressed as 1 - P with the detection probability d . Assume that the data q n sent to the n-th (n ∈ [1, N]) secondary user follows a complex Gaussian random distribution with a mean of 0 and a variance of 1, and P max represents the maximum transmit power of each ST.

[0089] S2: Obtain the optimal target model of the signal reception rate of each secondary user according to the cognitive cell-free system model.

[0090] Specifically, when the primary user does not occupy the frequency band, the n-th (n ∈ [1, N]) secondary user SU n will not be interfered by the primary user, and the received signal y n0 is expressed as:

[0091]

[0092] where represents the desired signal, and η mn represents the power coefficient allocated from the m-th (m ∈ [1, M]) secondary transmitter ST m to the n-th (n ∈ [1, N]) secondary user SU n , h mn represents the downlink channel coefficient between the ST m and the SU n . Since each secondary transmitter has K antennas, h mn = {h mn1 ,..., h mnk ,..., h mnK}, h mnk represents the downlink channel coefficient of the k-th antenna between the ST m and the SU n . W mn is the precoding matrix in the conjugate beamforming mode between the ST m and the SU n , and q n represents the data sent to the n-th secondary user. The second term represents the interference between secondary users, qi Data sent to the \(i\)-th secondary user, \(w\) n is Gaussian white noise with a mean of 0 and a variance of \(\sigma\) 2 .

[0093] When the primary user does not occupy the frequency band, the signal-to-interference-plus-noise ratio (SINR) of the \(n\)-th (\(n\in[1,N]\)) secondary user, SU n is given by:

[0094]

[0095] where \(\sigma\) 2 represents the variance of the Gaussian white noise.

[0096] Similarly, when the primary user occupies the frequency band, the \(n\)-th secondary user, SU n is interfered by the primary user, and the received signal \(y\) n1 is expressed as:

[0097]

[0098] where \(P\) t represents the transmit power of the primary transmitter, and \(h\) pn represents the channel coefficient between the primary transmitter and the \(n\)-th secondary user, SU n .

[0099] In this case, the SINR corresponding to the \(n\)-th secondary user, SU n is given by:

[0100]

[0101] Based on the SINR when the primary user occupies the frequency band and the SINR when the primary user does not occupy the frequency band, the signal reception rate \(R\) n of the secondary user, SU n is given by:

[0102] \(R\) n =(1 - P f )P r (H0)\(\log_2(1 + \text{SINR}\) n0 )+(1 - P d )P r (H1)\(\log_2(1 + \text{SINR}\) n1 )(5)

[0103] where \(P\) r (H0) is the probability that the primary user does not occupy the frequency band, and \(P\) r (H1) is the probability that the primary user occupies the frequency band.

[0104] To ensure a relatively fair quality of communication service, in this embodiment, the optimization objective is set to maximize the minimum user rate, and the optimization problem is modeled as follows:

[0105]

[0106] where h mp ={h mp1 ,...,h mpk ,...,h mpK} represents the channel coefficient between the primary user and the m-th secondary transmitter, and P Imax represents the maximum interference threshold of the primary user. The first constraint condition indicates that the interference from the secondary user to the primary user should be lower than the maximum interference threshold of the primary user, and the second constraint condition represents the transmission power limit of each secondary transmitter.

[0107] In the optimization problem considered in this embodiment, R n is a sum containing two logarithmic forms and cannot be converted into a convex optimization problem and searched in the form of the bisection method. And fixing one term and optimizing the other term in an alternating optimization form requires a relatively high time complexity. Therefore, in this embodiment, the original problem is converted into a Markov problem (Markov Decision Process, MDP), and a multi-agent reinforcement learning algorithm is used for solution.

[0108] S3: Create M agents according to M secondary transmitters, initialize the parameters of the policy network and the evaluation network of each agent, and initialize the power allocation coefficient of each agent.

[0109] In this embodiment, step S3 specifically includes:

[0110] S3.1: Create M agents from M secondary transmitters, and construct the cognitive cell-free system model into (S, O, A, R, P);

[0111] In the mean-field multi-agent system, the MDP problem is extended to a partially observable MDP (Partially Observable MDP, POMDP). This embodiment considers a multi-agent environment composed of M secondary transmitter agents, and the POMDP is composed of a tuple containing the following elements (S, O, A, R, P), where S = {s1,..., s m ,..., s M} represents the state set of all agents, O = {o1,..., o m ,..., o M} represents the local observation set of all agents, A = {a1,..., am ,..., a M}, represents the set of actions of all agents, \(R = \{r_1,..., r m ,..., r M}\) represents the set of return values of all agents, and \(P\) represents the state transition probability. Each agent uses its policy to select the action \(a m \). After the environment undergoes the joint action \(A\), it transfers to the next state \(S t+1 \), and the state transition function is \(\Gamma: S t \times A t \to S t+1 \). The state transition probability is represented by \(P(S t+1 |S t , A)\). After the agent takes an action, it obtains the return value \(r m \), and transfers to the next observable state \(o m : S \to o m \). The goal of each agent is to maximize its own cumulative expected return where \(\gamma\) represents the discount factor and \(T\) represents the total running time of the system.

[0112] S3.2: Construct the state, action, and return value of each agent according to the power coefficient between the secondary transmitter and the secondary user, the downlink channel coefficient, and the signal reception rate of the secondary user.

[0113] (1) State: In the POMDP problem, each agent selects an action based on its own partially observable state. In the actual cell-free scenario, due to the difference in the deployment location of the secondary transmitter ST, it is considered that each secondary transmitter can only obtain the channel state information between itself and its served user and the power allocation coefficient at the previous moment, and it is difficult to obtain the information of other secondary transmitters and their served users. Therefore, the local observation of the \(m\)-th secondary transmitter ST (agent) can be expressed as:

[0114] o m =\{\eta m1 ,..., \eta mn ,..., \eta mN , h m1 ,..., h mn ,..., h mN}\} (7)

[0115] where \(\eta mn represents the power coefficient allocated by the \(m\)-th secondary transmitter ST m (m\in[1, M]) to the \(n\in[1, N]\)-th secondary user SU n , and \(h mn represents ST mand SU n The downlink channel coefficients between them. It should be noted that each secondary transmitter can only see the channel coefficients of itself and all secondary users, as well as the power allocation coefficients it assigns to all secondary users.

[0116] And the state set of all agents is the set of the transposes of the observations of each secondary transmitter ST:

[0117]

[0118] (2) Action: In the cell-free power allocation problem of multi-agent non-ideal spectrum sensing, each agent is responsible for the power allocation scheme between the secondary transmitter and its served users. Therefore, the action of the m-th secondary transmitter ST can be expressed as:

[0119] a m ={η m1 ,...,η mn ,...,η mN} (9)

[0120] It is worth mentioning that since the secondary transmitter ST and the secondary users are not fully connected, each secondary transmitter ST does not necessarily serve all N secondary users. Therefore, when the secondary transmitter ST does not serve this secondary user, the power allocation coefficient of the secondary transmitter ST for this user is always 0.

[0121] (3) Reward value: The design of the reward value should fully meet the requirements of the optimization problem. In this embodiment, the reward value is designed into two parts:

[0122] r(t)=min{R1,...,R n ,...,R N}+c(t) (10)

[0123] First of all, in order to achieve the optimization goal of maximizing the minimum user rate, the first part of the reward value should be positively correlated with the objective function. Therefore, we use the minimum secondary user signal reception rate when taking action a t at state s t as the first part of the reward value: min{R1,...,R n ,...,R N}.

[0124] After that, considering the interference constraint and power constraint in the optimization problem, a cost function with a penalty term is subtracted from the reward value:

[0125] c(t)=-ap1-bp2 (11)

[0126] Among them, p1 is a fixed value, representing the penalty given when the interference constraint is not satisfied. a is the penalty factor for this item, used to control the size of the penalty; p2 is the number of secondary emitters that do not satisfy the power constraint. The more secondary emitters that do not satisfy the constraint, the greater the penalty given. b is the penalty factor for this item.

[0127] S4: Construct the weight coefficient of the weighted mean field of the MAMFSAC (Multi Agent Mean Field Soft Actor Critic) algorithm.

[0128] Considering that the number of nodes (secondary transmitters, i.e., agents) in the actual cell-free scenario is very large, and the traditional multi-agent algorithm will cause the variance of the evaluation gradient to increase when facing a large number of agents, making it extremely difficult for the evaluation network to converge. This is because the evaluation network of the traditional multi-agent algorithm needs to input the information of other agents, so the scale of data interaction will expand significantly with the increase of agents, making the training of the network very difficult.

[0129] The Mean Field Theory (MFT) is a method that uniformly processes the multiple action effects of the environment on an individual, and unifies the accumulation of multiple action effects into a combined effect. It is equivalent to replacing the influence of the environment on the research object with an equivalent field, which is called the mean field. The mean field can decouple the mutual influence of multiple bodies, separating it into the influence of the mean field on multiple single-body problems without interacting with other individuals. This method of MFT is used to handle some complex problems, and it can turn high-dimensional and intractable complex tasks into low-dimensional and tractable tasks.

[0130] The key of the mean field theory lies in its assumption that the behavior of any single agent will not significantly affect the average behavior of the multi-agent system. Based on this assumption, the mean field idea can use a unified effect to fit the influence of other surrounding agents on a single agent. By assigning weights to the mean field through the influence of relevant factors, when applied to multi-agent deep reinforcement learning, it can better reflect the clustering effect of the mean field idea and the consistency of the multi-agent system.

[0131] Based on the definition of the POMDP problem model, the process of the power allocation method based on MAMFSAC is introduced next. In the reinforcement learning algorithm, it is hoped that the policy can explore the environment as much as possible to obtain the optimal policy. However, if the policy output is a probability distribution with low entropy, it may greedily sample certain values and get stuck. The maximum entropy reinforcement learning model can well solve this problem. Under the condition of meeting the limited conditions (such as obtaining enough rewards), it randomly explores the unknown state space with equal probability. Entropy is defined as the expectation of the amount of information and is a measure of the uncertainty of a random variable. Its calculation formula is as follows:

[0132]

[0133] where X represents a random variable, and its possible values are X = {x1, x2,...}, P(x i ) = P(X = x i ) represents the probability distribution that the random variable X takes the value x i .

[0134] More intuitively, when the uncertainty of a random event (variable) is greater, the entropy is greater; on the contrary, if the random event is a deterministic event, its entropy is zero. When each possible probability of occurrence in a random event is equal, that is, the random event follows a uniform distribution, the entropy is the largest. Essentially, the meaning of the maximum entropy model is: under the condition of meeting the known knowledge or limited conditions, the best inference for the unknown is random uncertainty (equal probability of each random variable). Different from the traditional reinforcement learning that maximizes the expectation of the cumulative return value, the policy J(π) of SAC (Soft Actor Critic) includes the regularization terms of the return and entropy. In this way, on the basis of successfully completing the established task, the intelligent agent randomly explores more action spaces as much as possible, and increasing the exploration ability is more conducive to finding the global optimal solution. The expression of the policy is:

[0135]

[0136] where α is the temperature coefficient, which is a regularization coefficient used to control the importance of entropy. Entropy regularization increases the exploration degree of the reinforcement learning algorithm. The larger α is, the stronger the exploration ability is, which helps to accelerate the subsequent policy learning and reduce the possibility that the policy falls into a poor local optimum. r(S t , A t ) represents the reward obtained by executing action A t in state S t , and H(π(·|S t )) is the regularization term of entropy, representing the randomness of the policy π in state S t . Its calculation method is as follows:

[0137] H(π(·|S t )) = -log(π(·|S t )) (14)

[0138] where π(·|S t ) is the probability of taking action · in state S t .

[0139] The loss function of the entropy regularization term is as follows:

[0140]

[0141] where E represents taking the average, and (S t , A t , r t , S t+1 ) represents one piece of experience information, is the replay experience pool used to store experience information, means that (S t , A t , r t , S t+1 ) is taken from the replay experience pool

[0142] Since the optimal action of each agent cannot be obtained through the state value function V(S t ), in this embodiment, the state-action value function, that is, the Q function, is used to help each agent effectively obtain its optimal policy. After adding the local joint action of the mean field, the expression of the Q function is reconstructed as follows:

[0143]

[0144] where N(i) is the set of the remaining agents except agent i, is the weight coefficient of agent j with respect to agent i, and Q i (S, a i , a j ) represents the Q value when taking actions a i and a j in state S, and Q i (S, A) represents the Q value given by the evaluation network of agent i when taking the global action A in state S.

[0145] Generally speaking, the greater the influence of agent j on agent i, the higher its corresponding weight. And the mutual influence relationship between agents is largely related to the distance between agents. Therefore, the reciprocal of the distance between agents is often used as its weight coefficient. Assume the position of agent i is (x i , yi ), the position of agent j is (x i , y i ), then the weight coefficient of agent j with respect to agent i is expressed as:

[0146]

[0147] For the above expression of the Q function, in this embodiment, approximation processing is carried out through the mean field theory. The action of agent i is a i , and according to N(i), the average weighted action of other agents except agent i is calculated and the action a j of agent j is converted into and a margin δα i,j to represent.

[0148]

[0149] where δα i,j represents the margin, which is a very small value and can be ignored.

[0150] From this, the expression of the Q function after adding the mean field can be deduced:

[0151]

[0152] where, Since agent j is distributed around agent i, W is actually a bounded fluctuating value. When there are enough agents, agent j can be regarded as uniformly distributed, and at this time W can be approximated as zero. Under the approximation of the weighted mean field, the pairwise interaction ∑Q i (S, a i , a j ) between agent j and each agent i is simplified to In this way, only the weighted average action needs to be calculated, thus simplifying the interaction between agents and further simplifying the dimension of the evaluation network and the computational complexity.

[0153] Next, the design of the weight coefficient will be discussed. Considering the cell-free scenario characteristics of this embodiment of the invention, if only the distance is simply used as the weight coefficient between two agents, in fact, if there is no common service user between two agents, then the strategies output by these two agents have almost negligible influence on the service users of each other, and then these two agents can be regarded as two independent agents. Therefore, using the distance between two agents as the weight coefficient in the current scenario is not fine enough.

[0154] From the previous SINRn0 (Formula 2) and SINR n1 As can be seen from expression (Formula 4), the channel gain determines the size of the signal-to-interference-plus-noise ratio. Therefore, the rate of each secondary transmitter ST affecting the secondary user is largely related to its channel state information. Thus, in this embodiment, the channel gain is selected as the weight of each agent for this secondary user. Next, the specific formula derivation with the secondary user as the center and the channel gain as the weight will be discussed.

[0155] Since the cell-free scenario in this embodiment is partially connected, the secondary transmitter ST does not necessarily serve all secondary users. For the convenience of unified calculation, this embodiment assumes that all secondary transmitters ST have an N-dimensional pre-connection vector pre j .

[0156]

[0157] Among them, pre j (n) represents the connection situation between the j-th secondary transmitter ST and the n-th secondary user, L f represents the path loss, Threshold represents the path loss threshold, and pre j = {pre j (1),..., pre j (n),..., pre j (N)}. When the path loss is lower than the path loss threshold Threshold, it is stipulated that the secondary transmitter ST is connected to this secondary user, otherwise it is not connected. To introduce the secondary user as the center, is redefined as an N-dimensional vector, where is the weight coefficient of agent j for agent i on the n-th secondary user. Thus, a new weight coefficient can be obtained:

[0158]

[0159] Among them, h jnk represents the channel coefficient of the j-th secondary transmitter ST for the n-th secondary user on the k-th antenna. Correspondingly, w i also becomes an N-dimensional vector. Therefore, the weight coefficient w i (n) of agent i on the n-th secondary user is obtained:

[0160] S5: Use the multi-agent mean-field soft actor-critic algorithm to solve the POMDP problem model to obtain the optimal power allocation coefficient.

[0161] Specifically, step S5 of this embodiment specifically includes:

[0162] S5.1: Create M agents, each agent includes a policy network and an evaluation network, initialize the policy network parameters θ and the agent power allocation coefficients {η1,..., η N}, initialize the network parameters ω1, ω2, The policy network inputs the local observation o of the current agent m , and the evaluation network inputs the state set S of all agents, the action a of the current agent i and the weighted average action

[0163] S5.2: Obtain the experience information (S t , O t , A t , R t , S t+1 , O t+1 ) of all agents and store it in the experience pool , where S t , O t , A t , R t respectively represent the state set, partial observation set, action set and return value set of all agents at the current moment t, and S t+1 , O t+1 respectively represent the state set and partial observation set of all agents at the next moment t + 1. The experience pool stores the experience information of agents at different times, that is, the experience pool stores the state set, partial observation set, action set and return value set of agents at different times.

[0164] Specifically, S5.21: Each agent obtains the local observation o at the current moment t t , and inputs the local observation o t into the policy network to obtain the action a corresponding to the current moment t t ;

[0165] S5.22: Use the action a corresponding to each agent at the current moment t t to obtain the current return value r of each agent t and the observation o at the next moment t+1 ;

[0166] S5.23: Obtain the experience information (S t , O t , A t , R t , S t+1 , O t+1 ) of all agents, and store it in the experience pool In.

[0167] As mentioned above, in the definition of the weighted average action , w i and both appear as a scalar. In fact, the action of each agent is also an N-dimensional vector. After rewriting the weight coefficient as a vector centered on the secondary user, each element of a i and can be controlled by the newly defined w j , specifically expressed as:

[0168]

[0169] where represents the weighted average action of the remaining agents except agent i on the nth secondary user, and a j (n) represents the action of the jth agent on the nth secondary user, that is, the power allocation coefficient of the jth secondary transmitter to the nth secondary user.

[0170] According to the mean field theory, this embodiment modifies the input of the action part in the evaluation network of the agent. Specifically, the input change of the evaluation network is as follows. The dimension of the state remains (K + 1)MN unchanged, and the dimension of the action changes from MN to 2N. In the actual scenario, the number of secondary transmitters is generally more than three times the number of users. Therefore, through the compression of the action by the mean field, the dimension of the evaluation network can be significantly reduced, improving the performance of the network.

[0171]

[0172] In terms of the computational complexity, since the evaluation network is a fully connected layer structure and the number of neurons in the middle two hidden layers is 256 each, taking backpropagation as an example, for the non-mean field, the input of its evaluation network is (K + 2)MN. Then, one calculation requires ((K + 2)MN * 256 + 256 * 256 + 256 * 1) operations, while for the mean field, one calculation requires (((K + 1)MN + 2N) * 256 + 256 * 256 + 256 * 1) operations. For the problem scale, only the three variables K, M, and N in the input are variables, and the number of neurons in the hidden layer remains 256 unchanged, which can be regarded as a constant. Therefore, the computational complexity of the non-mean field is O((K + 2)MN), while the computational complexity of the mean field is O((K + 1)MN + 2N). When the value of M is quite large, the computational complexity of the mean field is significantly reduced.

[0173] According to the mean field theory, the action value function at time t is updated as follows:

[0174]

[0175] Among them, β represents the learning rate, is the weighted average action of the actions of other agents except agent i, and the policy is obtained from the Boltzmann distribution:

[0176]

[0177] Among them, λ represents the Boltzmann constant.

[0178] From this, the state value function of agent i at time t can be obtained:

[0179]

[0180] And the loss function of the Q function is as follows:

[0181]

[0182] Among them, represents the data collected by the policy in the past. Since SAC is an off-policy algorithm. To make the training more stable, the evaluation network in this embodiment uses a target Q network There are also two target Q networks, which correspond to the two Q networks one by one. The update method of the target network in SAC is: initially use the parameters ω of the Q network, and then use the delayed soft update method during network update, that is, after several rounds, assign the parameters ω of the Q network to the parameters ω of the target Q network with a certain weight - . Please refer to Figure 3 , Figure 3 is a schematic structural diagram of an evaluation network provided by an embodiment of the present invention. The input layer consists of the global state S, the action a of the agent, and the average action of other agents It is composed of 256 neurons in both hidden layers, and finally the Q value is output.

[0183] Similarly, the loss function of the policy network is:

[0184]

[0185] Please refer to Figure 4 , Figure 4It is a schematic structural diagram of a policy network provided by an embodiment of the present invention. The data of the input layer is the local observation o of each agent, including the channel state information at the current moment and the power allocation coefficient at the previous moment. The number of neurons in the hidden layer is the same as that of the evaluation network. It is worth mentioning that the output layer of the SAC policy network is not like that in traditional reinforcement learning. The number of neurons in its output layer is twice the dimension of the action. This is because for a continuous action space, the policy network of SAC outputs the probability distribution of each action, and then samples from this distribution to obtain the action.

[0186] S5.3: Train the agents according to the experience information of the agents at different times, update the network parameters of the policy network and the evaluation network, and obtain the trained multi-agent system.

[0187] Step S5.3 of this embodiment specifically includes:

[0188] S5.31: Select the experience information of the agents at different times in the experience pool as training data;

[0189] S5.32: Input the training data into the policy network and the evaluation network, calculate the weighted weight coefficient of the current action, and calculate the weighted average action according to the weighted weight coefficient;

[0190] S5.33: According to calculate the loss function of entropy, according to obtain the loss function of the evaluation network, according to obtain the loss function of the policy network, where E represents taking the average, represents (S t , A t , r t , S t+1 ) is taken from the replay experience pool R, α is the temperature coefficient, H(π(·|S t )) is the entropy regularization term, π(A t |S t ) represents the probability of executing action A t in state S t , represents the Q value calculated by the evaluation network with parameters ω when the global state is S t , the agent's own action is a t , and the weighted average action of the remaining agents is ; represents the state value calculated by the evaluation network with parameters ω t+1 when the global state is S - , γ represents the discount factor, π θ (At |S t ) represents the probability that when in state S t , the policy network with parameter θ outputs action A t .

[0191] S5.34: Perform backpropagation operations on the loss function of the entropy, the loss function of the evaluation network, and the loss function of the policy network to obtain the corresponding gradients and According to and update the evaluation network parameters ω i and According to update the policy network θ, and according to update the parameters of the maximum entropy network, where λ Q and τ represent the evaluation network update coefficients, and λ π and λ α represent the policy network update coefficient and the regularization term update coefficient respectively.

[0192] S5.4: Input the local observation data of the secondary user into the trained multi-agent system to obtain the optimal power allocation coefficient of the secondary user.

[0193] Please refer to Figure 5 , Figure 5 which is a schematic diagram of a cognitive cell-free power allocation method based on the MAMFSAC algorithm provided by an embodiment of the present invention. Figure 5 Details show the data exchange process between each agent and the internal network of MAMFSAC. Each agent includes 1 policy network (Actor) and 1 evaluation network (Critic). Each agent can only obtain the local observation o. The policy network of each agent outputs the corresponding action a according to the local observation o. The actions of all agents are jointly input to the environment to obtain the current reward r, the state s' and the observation o' at the next moment. These data form a batch of data and are stored in the experience pool. In the network update stage, each agent obtains a batch of data from the experience pool. The global action information A in the data is split into its own action a and the weighted average action the global state S, its own action a and the average action are sent to the evaluation network and generate the corresponding Q value, and the Q value can guide the update of the policy network. The following is the specific process of the power allocation algorithm for the cognitive cell-free system based on MAMFSAC.

[0194]

[0195]

[0196] To verify the correctness of the cognitive cell-free power allocation method based on the MAMFSAC algorithm of the present invention, it is necessary to compare it with traditional optimization algorithms. However, when traditional optimization algorithms solve problems such as max min, they need to convert the solution expression into a convex function and then use the method of convex optimization to solve it. According to the method proposed in the literature (J. Qiu, K. Xu, X. Xia, Z. Shen and W. Xie, "Downlink Power Optimization for Cell-Free Massive MIMO Over Spatially Correlated Rayleigh Fading Channels," in IEEE Access, vol. 8, pp. 56214-56227, 2020, doi: 10.1109 / ACCESS.2020.2981967.), the logarithmic function in the optimization problem is converted into an exponential function through inequality transformation. However, there are two logarithmic functions in the optimization problem of this embodiment, and due to the detection probability P d and the false alarm probability P f , these two logarithmic functions cannot be combined into one logarithmic function, so the method proposed in the above literature cannot be used for solution. Therefore, an extreme case is considered, that is, let P d = 1, so that the two logarithmic functions in the optimization objective become one logarithmic function.

[0197] Please refer to Figure 6 , Figure 6 which is a comparison chart of the secondary transmitter rates obtained by different power allocation methods. The y-axis on the left represents the rate of the secondary transmitter, and the y-axis on the right represents the ratio of the results obtained by the reinforcement learning (RL) proposed in the present invention to the results obtained by the traditional iterative method (trad). The comparison simulation results show that under the same scenario and channel state information conditions, the maximum-minimum rates obtained by using the method of the present invention can all reach more than 99.3% of the optimal rate. And as the number of secondary transmitters increases, it gets closer and closer to the optimal rate, indicating that the method proposed in the present invention has stronger scalability and robustness. In addition, this embodiment also counts the time taken for the method of the present invention to execute 100 times and the traditional optimization algorithm to execute 100 times after training as shown in Table 1. It can be seen that whether in small-scale or large-scale scenarios, the running time of the power allocation method of the present invention is much less than that of the traditional optimization method. Applied to the actual scenario, the time required for the power allocation method of the present invention to run once is basically in the millisecond level, and the time required for the traditional optimization method to run once is in the second level. Therefore, the power allocation method of the present invention can meet the needs of the actual scenario, while the traditional optimization method at the second level cannot meet the needs of the actual scenario.

[0198] Table 1. Comparison of running times of different power allocation methods

[0199] Number of agents Running time of the method of the present invention (s) Running time of the traditional optimization algorithm (s) 6 0.317352057 476.530507 9 0.424709082 554.374361 12 0.59490633 633.694727 15 0.753329992 709.478022 18 1.164143085 770.36361 21 1.349216223 813.401351

[0200] The present invention provides a cognitive cell-free power allocation method based on the MAMFSAC algorithm. By using a data-driven algorithm of multi-agent reinforcement learning, after the agents are trained, the original non-convex problem can be solved in a very short time. In large-scale scenarios, the present invention utilizes the mean-field theory to replace the action information of other agents in the evaluation network with an equivalent mean action, thereby reducing the original action dimension of O(MN) to O(N), greatly improving the performance and scalability of the algorithm. According to the characteristics of the cell-free network scenario, the present invention designs the channel coefficient as the weight of the weighted mean field. This weight design method fully considers the coupling relationship between the secondary transmitter and the secondary user and the small-scale influence. Compared with the traditional method using distance as the weight, the equivalent mean action calculated by the weight proposed in the present invention is more accurate.

[0201] Another embodiment of the present invention provides a storage medium in which a computer program is stored, and the computer program is used to execute the steps of the cognitive cell-free power allocation method based on the MAMFSAC algorithm described in the above embodiment. Another aspect of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the cognitive cell-free power allocation method based on the MAMFSAC algorithm described in the above embodiment are implemented. Specifically, the above-mentioned integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0202] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A cognitive cell-free power allocation method based on multi-agent mean field soft actor critic MAMFSAC algorithm, characterized in that: include: S1: constructing a cognitive cell-free system model, wherein the cognitive cell-free system model includes a primary transmitter, a primary user, M secondary transmitters and N secondary users; S2: Obtaining an optimal target model of the signal reception rate of each secondary user according to the cognitive cell-free system model; comprising: S2.1: Get the n∈[1,N]th secondary user SU n The signal y received when the primary user does not occupy the frequency band n0 : in, represents the expected signal, η mn Denoted as the m∈[1,M]th secondary transmitter ST m Assigned to the n∈[1,N]th secondary user SU n The power coefficient, h mn Indicated as ST m with SU n The downlink channel coefficient between mn ={h mn1 ,...,h mnk ,...,h mnK }, h mnk Indicates ST m with SU n The downlink channel coefficient of the kth antenna between mn For ST m with SU n The conjugate beamforming precoding matrix between n Represents the data sent to the nth secondary user, represents the interference between secondary users, q i represents the data sent to the i-th secondary user, w n The mean is 0 and the variance is σ 2 Gaussian white noise; S2.2: Get the nth secondary user SU n Signal-to-interference-noise ratio when the primary user does not occupy the frequency band: Among them, P max Indicates the maximum transmission power of each secondary emitter; S2.3: Get the nth secondary user SU n The signal y received when the primary user occupies the frequency band n1 : Among them, P t Indicates the transmission power of the primary transmitter, h pn Represents the primary transmitter and the nth secondary user SU n The channel coefficient between ; S2.4: Get the nth secondary user SU n The corresponding signal-to-interference-noise ratio when the primary user occupies the frequency band is: S2.5: According to the signal-to-interference-noise ratio when the primary user occupies the frequency band and the signal-to-interference-noise ratio when the primary user does not occupy the frequency band, the secondary user SU is obtained. n The signal receiving rate R n for: R n =(1-P f )P r (H0)log2(1+SINR n0 )+(1-P d )P r (H1)log2(1+SINR n1 ) Among them, P r (H0) is the probability that the primary user does not occupy the frequency band, P r (H1) is the probability of primary users occupying the frequency band, H1 is channel occupancy, H0 is channel idle; P f =P r {The judgment is H1|H0} represents the false alarm probability, P d =P r {The judgment is H1|H1} represents the detection probability; S2.6: Construct the optimal target model of the signal reception rate of the secondary user: Among them, h mp represents the channel coefficient between the primary user and the mth secondary transmitter, Indicates the maximum interference threshold of the primary user; S3: creating M agents according to the M secondary transmitters to form a multi-agent system, and modeling the optimal target model as a partially observable Markov POMDP problem model; including: S3.1: Create the M secondary transmitters into M intelligent agents, and construct the cognitive cell-free system model into (S, O, A, R, P), where S = {s1, ..., s m ,...,s M } represents the state set of all agents, O = {o1,...,o m ,...,o M } represents the local observation set of all agents, A={a1,...,a m ,...,a M } represents the action set of all agents, R = {r1,...,r m ,...,r M } represents the set of reward values ​​of all agents, and P represents the state transition probability; S3.2: construct the state, action and reward value of each intelligent agent according to the power coefficient between the secondary transmitter and the secondary user, the downlink channel coefficient and the signal reception rate of the secondary user; S4: Construct the weight coefficients of the weighted mean field of the multi-agent mean field soft actor-critic algorithm; including: S4.1: Assume that all secondary transmitters include an N-dimensional pre-connection vector pre j : Among them, pre j (n) represents the connection between the jth secondary transmitter and the nth secondary user, L f Indicates the path loss, Threshold indicates the path loss threshold, j ={pre j (1),...,pre j (n),...,pre j (N)}; S4.2: Obtain the weight coefficient of agent j on agent i on the nth secondary user: Where K represents the number of antennas of the secondary transmitter, h jnk represents the channel coefficient of the jth secondary transmitter to the nth secondary user on the kth antenna; S4.3: Obtain the weight coefficient w of agent i on the nth secondary user i (n): S5: solving the POMDP problem model using the multi-agent mean field soft actor critic algorithm to obtain the optimal power allocation coefficient; including: S5.1: Create M agents, each of which includes a policy network and an evaluation network, initialize the policy network parameters θ and the agent power allocation coefficients {η1,...,η N }, initialize the network parameters ω1 of the evaluation network, ω2, S5.2: Obtain the experience information of all agents at different times and store it in the experience pool , wherein the experience information at different times includes a state set, a partial observation set, an action set, and a reward value set at different times; S5.3: training the agent according to the experience information of the agent at different times, updating the network parameters of the strategy network and the evaluation network, and obtaining a trained multi-agent system; S5.4: Inputting the local observation data of the secondary user into the trained multi-agent system to obtain the optimal power allocation coefficient of the secondary user.

2. The cognitive non-cellular power allocation method based on the MAMFSAC algorithm according to claim 1, characterized in that: The S3.2 includes: S3.21: Get the local observation o of the mth agent m : the m ={η m1 ,...,or mn ,...,or mN ,h m1 ,...,h mn ,...,h mN } Among them, η mn Denoted as the mth secondary transmitter ST m Assigned to the nth secondary user SU n The power coefficient, h mn Indicated as ST m and SU n Downlink channel coefficient between ; S3.22: According to the local observation o m Get the state set of all agents: S3.23: Get the action of the mth agent: a m ={η m1 ,...,or mn ,...,or mN }; S3.24: Design the return value at time t as: r(t)=min{R1,...,R n ,...,R N }+c(t), Among them, min{R1,...,R n ,...,R N } indicates state s t Take action a t The minimum secondary user signal receiving rate is c(t), and c(t) represents the cost function with penalty term.

3. The cognitive non-cellular power allocation method based on the MAMFSAC algorithm according to claim 1, characterized in that: The S5.2 includes: S5.21: Each agent obtains the local observation o at the current time t t , the local observation o t Input into the strategy network to obtain the action a corresponding to the current time t t ; S5.22: Use the action a corresponding to each agent at the current time t t Get the current reward value r of each agent t and the observation o at the next moment t+1 ; S5.23: Get the experience information of all agents (S t ,O t ,A t ,R t ,S t+1 ,O t+1 ) and store it in the experience pool middle.

4. The cognitive non-cellular power allocation method based on the MAMFSAC algorithm according to claim 1, characterized in that: The S5.3 includes: S5.31: In the experience pool Select the agent's experience information at different times as training data; S5.32: Input the training data into the strategy network and the evaluation network, calculate the weighted weight coefficient of the current action, and calculate the weighted average action according to the weighted weight coefficient; S5.33: According to Calculate the entropy loss function according to Obtain the loss function of the evaluation network according to Obtain the loss function of the policy network, where E represents the average, Indicates (S t ,A t ,r t ,S t+1 ) is taken from the replay experience pool α is the temperature coefficient, H(π(·|S t )) is the regularization term of entropy, π(A t |S t ) indicates that in state S t When , execute action A t The probability of Indicates that the global state is S t , the agent's own action is a t , the weighted average action of the remaining agents is The Q value calculated by the evaluation network with parameter ω, Indicates that the global state is S t+1 When the parameter is ω - The state value calculated by the evaluation network, γ represents the discount factor, π θ (A t |S t ) indicates that in state S t When θ is the parameter of the policy network, the output action is A. t probability; S5.34: Perform back propagation operations on the entropy loss function, the evaluation network loss function, and the policy network loss function to obtain corresponding gradients and Update the parameters ω of the evaluation network i and The parameters θ of the policy network and the parameter α of the entropy regularization term.

5. A storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the cognitive non-cellular power allocation method based on the MAMFSAC algorithm as claimed in any one of claims 1 to 4.

6. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor calls the computer program in the memory, implements the steps of the cognitive non-cellular power allocation method based on the MAMFSAC algorithm as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Intelligent beam tracking method based on beam image stacking under cellular-free millimeter waves

    CN116346185A

  • Multi-agent learning method for non-cellular network user scheduling and resource configuration

    CN117221925A