6G network-oriented potential learning intelligent resource allocation method

By introducing a multi-agent reinforcement learning algorithm based on channel prediction and hybrid action representation in 6G networks, the problem of unutilized time-varying channel characteristics and dependencies is solved, the efficiency and speed of resource allocation are improved, and higher transmission rates and spectrum efficiency are achieved.

CN120633723APending Publication Date: 2025-09-12NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510600838.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing resource allocation technologies cannot effectively utilize the time-varying characteristics of channels and the dependency between channel allocation and transmit power in 6G networks, resulting in poor performance and an inability to meet the real-time requirements and high efficiency demands of spectrum-sharing networks.

Method used

A channel prediction network and a hybrid action representation network are used. Channel prediction is performed through a GRU network, and a hybrid action representation is constructed in combination with a variational autoencoder. A multi-agent reinforcement learning algorithm is used to optimize channel selection and power allocation, establish a dependency relationship between channel selection and power allocation, and improve the calculation speed and convergence speed of resource allocation.

Benefits of technology

It improves the resource allocation performance of spectrum sharing networks, increases the system's transmission rate and spectrum efficiency, and achieves faster network convergence speed, making it suitable for future actual wireless communication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633723A_ABST
    Figure CN120633723A_ABST
Patent Text Reader

Abstract

The invention discloses a potential learning intelligent resource allocation method for a 6G network, and the method comprises the steps: 1, collecting 6G network channel data, carrying out the normalization, and constructing a data set; step 2, training a channel prediction network by using the data set and obtaining a channel prediction result; 3, constructing an input state vector of each secondary user agent according to the channel data and the channel prediction result, and constructing a reward function; 4, building a hybrid action representation network by adopting an embedded table and a variational auto-encoder, and carrying out pre-training; and 5, training the multi-agent reinforcement learning network according to the input state vector of each secondary user agent, the reward function and the pre-trained mixed action representation network so as to perform intelligent resource allocation. According to the method disclosed by the invention, the spectrum sharing system can obtain the maximum transmission rate by utilizing two kinds of potential knowledge of the dependency relationship between the channel time-varying characteristic and the channel allocation and the transmitting power, the system transmission rate higher than that of the existing method is realized, and the network convergence speed is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communication technology, and in particular relates to a potential learning intelligent resource allocation method for 6G networks. Background Art

[0002] Spectrum sharing, an emerging technology, can alleviate spectrum shortages and facilitate high-speed transmission for sixth-generation (6G) communications. Specifically, this technology allows secondary users (SUs) to reuse frequency bands authorized to primary users (PUs) while maintaining the quality of service (QoS) of the PUs. However, protecting PUs requires constraining SU performance through resource management. Therefore, in spectrum-sharing networks, resource allocation that improves SU performance while maintaining PU QoS is crucial.

[0003] Existing resource allocation solutions fall into two main categories: traditional optimization schemes based on optimization theory and those based on machine learning. Specifically, traditional optimization schemes typically utilize several common methods, such as approximation-based methods and heuristics-based methods. Due to the large-scale nature of spectrum sharing networks, resource allocation problems are often non-deterministic and polynomially hard. Due to the complexity and non-convexity of such problems, the solutions obtained by traditional optimization schemes cannot guarantee any global optimality. Furthermore, these traditional optimization schemes may incur high computational complexity and are unable to adapt to the high dynamics of spectrum sharing networks.

[0004] In recent years, in order to solve these challenging resource allocation problems, machine learning-based solutions, especially deep reinforcement learning (DRL) solutions, have been proposed. In their article "Energy-efficient mode selection and resource allocation for D2D-enabled heterogeneous networks: A deep reinforcement learning approach" (IEEE Trans. Wireless Commun., 2021), Tao Zhang, Kun Zhu et al. proposed a deep reinforcement learning algorithm based on deep deterministic policy gradient (DDPG) to address the energy efficiency problem of supporting end-to-end heterogeneous cellular networks, to select communication modes and allocate resources including bandwidth, power, and channels. The patent with publication number CN117880843A proposes an intelligent resource allocation algorithm based on dual deep Q networks and deep deterministic gradients to improve the secondary network data transmission rate. Although the performance of DRL methods is generally better than traditional RL methods, the efficiency of existing DRL algorithms in improving system performance in highly dynamic wireless communication environments is still very low. This is because the highly dynamic characteristics of the actual network environment cause the wireless channel to exhibit time-varying characteristics and non-stationary statistical characteristics. However, algorithms based on the stationary assumption can suffer significant performance degradation in such scenarios, especially when ignoring the impact of channel quality fluctuations. Furthermore, in dynamic spectrum sharing networks, channel allocation represents a discrete action, while transmit power corresponds to a continuous parameter associated with that action in the hybrid action space. Furthermore, since SUs and PUs share the same licensed frequency band, the SU's channel allocation is closely related to transmit power. This is because when selecting a channel that a PU has already accessed, the SU must appropriately control its transmit power to ensure a satisfactory data rate while avoiding damage to the PU. Most existing DRL algorithms discretize the continuous power values ​​in the hybrid action space involving channel selection and power allocation to transform the heterogeneous space into a homogeneous one. However, this leads to scalability issues, as the number of discrete actions to be explored grows exponentially with the untrustworthy continuous parameter space. Furthermore, they ignore the dependencies between continuous and discrete actions in the hybrid action space. Therefore, the existing resource allocation technology has poor performance in actual communication scenarios and cannot meet real-time requirements. It also ignores the potential dependency between each secondary user's channel allocation and transmit power, resulting in performance limitations. There is an urgent need to propose a new intelligent resource allocation method. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned existing technologies and provide a potential learning intelligent resource allocation method for 6G networks. By utilizing two potential knowledges, namely the time-varying characteristics of the channel and the dependency between channel allocation and transmission power, the method improves the calculation speed and convergence speed, improves the resource allocation performance of the dynamic spectrum sharing network, and improves the transmission rate of the system.

[0006] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:

[0007] A potential learning intelligent resource allocation method for 6G networks, comprising:

[0008] Step 1: Collect and normalize 6G network channel data to construct a dataset;

[0009] Step 2: Use the data set to train the channel prediction network and obtain the channel prediction results;

[0010] Step 3: Construct the input state vector of each secondary user agent based on the channel data and channel prediction results, and construct the reward function;

[0011] Step 4: Use the embedding table and variational autoencoder to build a hybrid action representation network and perform pre-training;

[0012] Step 5: Train a multi-agent reinforcement learning network based on the input state vector, reward function, and pre-trained hybrid action representation network of each secondary user agent for intelligent resource allocation.

[0013] To optimize the above technical solutions, specific measures taken also include:

[0014] The channel data described in the above step 1 includes the channel gain from the primary base station to the user and the channel gain from the secondary base station to the user, wherein the channel gain from the primary base station to the user includes the channel gain from the primary base station to each primary user and the channel gain from the primary base station to each secondary user on each sub-frequency band; the channel gain from the secondary base station to the user includes the channel gain from the secondary base station to each secondary user and the channel gain from the secondary base station to each primary user on each sub-frequency band.

[0015] The dataset described in step 1 above is constructed as follows: the normalized channel data is constructed into a multidimensional channel data matrix, where the ordinate is the number of channels and the abscissa is the time scale. The multidimensional channel data of the first five moments are taken as input, and the multidimensional channel data of the last moment is used as the label value of the channel prediction data. The sliding window size is set to the length of the prediction data, and the sliding window is used to construct the dataset.

[0016] The channel prediction network in step 2 is a GRU network. At time t, it resets the gate R t and update gate Zt Respectively expressed as:

[0017] R t =σ(X t W xr +H t-1 W hr +b r )

[0018] Z t =σ(X t W xz +H t-1 W hz +b z )

[0019] in, is the input feature sequence of d rows and n columns, is a weight matrix with n rows and h columns, is the h-order weight matrix, is a 1-row, h-column bias matrix, with variable j∈{r,z};

[0020] The candidate hidden states are:

[0021] in, is the weight parameter, is the bias parameter;

[0022] The effect of the hidden state combined with the update gate is:

[0023] Among them, H t-1 is the hidden state at the previous moment.

[0024] In step 3 above, at time t, the input state vector of the mth secondary user agent is constructed as:

[0025]

[0026] The channel data s at time t m (t) is: Channel data prediction result s at time t+i pre,m (t+i) is:

[0027]

[0028] Among them, h m,k (t) represents the channel data of the mth secondary user in the kth channel at time t; h pre,m,k (t+i) represents the channel data prediction result of the mth secondary user in the kth channel at time t+i.

[0029] In step 3 above, construct the reward function where r m,t is the reward of the mth secondary user agent at time t;

[0030] When the mth secondary user agent does not access any channel, or the transmission power P k,m When r is 0, m,t =0;

[0031] When the mth secondary user agent chooses to access the kth channel, and the interference to the primary user of the channel is ρ k,m |h S,k,m (t)| 2 P k,m Less than or equal to the threshold Γ k When r m,t =R k,m -μ(Γ k -ρ k,m |h S,k,m (t)| 2 P k,m );

[0032] When the mth secondary user agent cannot access the kth channel, and the interference to the primary user of the channel ρ k,m |h S,k,m (t)| 2 P k,m Exceeding the threshold Γ k When r m,t =μ(Γ k -(1-ρ k,m )|h S,k,m (t)| 2 P k,m );

[0033] Where μ is the interference penalty factor, h S,k,m (t) is the channel gain from the secondary base station to the primary user of the kth channel, and the binary spectrum allocation index ρ k,m Assign strategy to the mth secondary user on the kth channel, R k,m is the available rate of the mth secondary user on the kth channel.

[0034] The hybrid action representation network constructed in step 4 above uses an embedding table and a variational autoencoder to construct a potential hybrid space and embed the dependency between discrete actions and continuous actions; wherein, the embedding table express discrete actions, ζ is a learnable parameter, and each row of the embedding table e ζ,a =E ζ (a) A d1-dimensional continuous vector corresponding to a discrete action a; in the variational autoencoder, the encoder Embed the row vectors in the table and state s t As a condition, the continuous parameter Mapping to a low-dimensional latent representation of d2 dimensions encoder The encoding process is expressed as: Under the same conditions, the decoder From z t Reconstructed in is the fully connected layer for reconstruction, is a conversion network, and the decoding process is expressed as: is any potential vector of dimension d1, a t Select the result for the channel.

[0035] The above hybrid action representation network also introduces the dynamic changes of the state dynamic predictor learning environment to optimize the hybrid action representation. Specifically, for any sample The state residual is expressed as The state dynamics predictor is expressed as: in, Indicates that the state residual is predicted using a cascade structure.

[0036] The above step 4 uses a random strategy to collect samples To the experience pool Then randomly select from the experience pool Samples of set batch size are extracted from the network for pre-training, and the mixed action representation loss function used in the pre-training process is:

[0037]

[0038] in, is the power allocation result, e at Encode the result for the embedding table, represents the embedding table and the loss function of the variational autoencoder, Represents the state L2 norm squared prediction error.

[0039] The above multi-agent reinforcement learning network includes actor network, critic network, actor target network and critic target network. During the training process of the multi-agent reinforcement learning network, according to the input state s at the current moment, t , obtain action embedding representation With z t , using the pre-trained mixed action representation network to obtain the channel selection result a t And power allocation results Then calculate the reward r(t) according to the reward function, complete a single environment interaction, and switch to the state s at the next moment t+1, and so on;

[0040] During the training of the multi-agent reinforcement learning network, the target network parameters are soft-updated. The update process of the actor target network and the critic target network is: Among them, τ is a hyperparameter between 0 and 1, and τ<<1; k m and η m are the weight parameters of the actor network and the critic network respectively; and are the weight parameters of the actor target network and the critic target network respectively.

[0041] The present invention has the following beneficial effects:

[0042] The present invention introduces channel prediction results as the potential knowledge of a new resource allocation framework. Compared with the traditional reinforcement learning framework, it learns the inter-channel dependency and time-varying characteristics.

[0043] This invention uses a hybrid action representation network to establish a dependency between channel selection and power allocation. Compared with the random resource allocation method, it accelerates network convergence, increases the overall transmission rate of the system, and better improves the spectrum efficiency of actual communication scenarios.

[0044] The hybrid action representation-based intelligent resource allocation framework combined with channel prediction proposed in the present invention does not restrict multi-agent reinforcement learning algorithms and application scenarios, making the framework highly generalizable and applicable to future practical wireless communication systems.

[0045] The channel prediction enhanced intelligent resource allocation method based on hybrid action representation proposed in the present invention can enable the spectrum sharing system to obtain the maximum transmission rate, achieve a higher system transmission rate than the existing method, and achieve a faster network convergence speed, which makes the present invention promising for application in actual communication scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a flow chart of the multi-agent reinforcement learning algorithm based on hybrid action representation proposed in this invention;

[0047] Figure 2 It is the prediction network structure diagram proposed by the present invention;

[0048] Figure 3 It is a hybrid action representation model diagram proposed by the present invention;

[0049] Figure 4 This is a diagram of the intelligent resource allocation framework based on the prediction network proposed by the present invention;

[0050] Figure 5This is a framework diagram of the multi-agent reinforcement learning algorithm based on hybrid action representation proposed in this invention;

[0051] Figure 6 This is a comparison chart of the total transmission rate of the present invention and other existing technologies under different conditions;

[0052] Figure 7 This is a comparison chart of the total transmission rate of the present invention and other existing technologies at different experience pool capacities. DETAILED DESCRIPTION

[0053] The embodiments of the present invention are described in further detail below with reference to the accompanying drawings.

[0054] The present invention provides a potential learning intelligent resource allocation method for 6G networks, which solves the problems of existing resource allocation technology in actual communication scenarios, such as poor performance and inability to meet real-time requirements, and performance limitations caused by ignoring the potential dependency between each secondary user channel allocation and transmission power. Figure 1 , the method specifically comprises the following steps:

[0055] Step 1: Collect and normalize 6G network channel data to construct a dataset;

[0056] Step 2: Use the data set to train the channel prediction network and obtain the channel prediction results;

[0057] Step 3: Construct the input state vector of each secondary user agent based on the channel data and channel prediction results, and construct the reward function;

[0058] Step 4: Use the embedding table and variational autoencoder to build a hybrid action representation network and perform pre-training;

[0059] Step 5: Train a multi-agent reinforcement learning network based on the input state vector, reward function, and pre-trained hybrid action representation network of each secondary user agent for intelligent resource allocation.

[0060] In the embodiment, the specific steps of step 1 channel data processing are as follows:

[0061] Step 1-1: Channel data collection. Channel data collection mainly includes the channel gain from the primary base station to the user and the channel gain from the secondary base station to the user. The channel gain from the primary base station to the user includes the channel gain h from the primary base station to the primary user on the kth sub-band. P,k (t) and the channel gain h from the primary base station to the mth secondary user P,k,m (t); The channel gain from the secondary base station to the user includes the channel gain h from the secondary base station to the mth secondary user on the kth sub-band k,m (t) and the channel gain h from the secondary base station to the kth primary user S,k,m(t). In this invention, the channel gain data is provided by Huawei.

[0062] Step 1-2: Channel data preprocessing. In this embodiment of the present invention, the channel data collected based on step 1-1 is preprocessed. First, all data is divided into 80% for training and 20% for testing; and all data is normalized to be scaled to the range [0, 1]:

[0063]

[0064] Among them, x represents the channel data to be processed, x min Indicates the minimum value of each column of channel data, x max Indicates the maximum value of each column of channel data.

[0065] Steps 1-3, channel data division. In order to enable the training network to adapt better and faster to time-varying channel data, the normalized data is converted into a matrix, where the ordinate is the number of channels and the abscissa is the time scale, that is, how much acquired data is used to predict the next data. In the present invention, the multidimensional channel data of the first five moments is taken as input, and the data of the last moment is used as the label value of the channel prediction data. The sliding window size is set to the length of the predicted data, and a data set is constructed using the sliding window. The data set is randomly divided into a training set and a test set, where the training set is used to generate a prediction model and optimize its parameters, and the test set is used to verify the accuracy of the prediction algorithm.

[0066] In the embodiment, the specific steps of step 2 prediction network training are as follows:

[0067] Step 2-1: Prediction Network Construction. This paper uses a gated recurrent unit (GRU) network to implement time series prediction. GRUs have been shown to have similar performance to long short-term memory (LSTM) networks, but with fewer network parameters, a simpler structure, and lower computational complexity.

[0068] At time t, the reset gate and update gate are expressed as

[0069] R t =σ(X t W xr +H t-1 W hr +b r )

[0070] Z t =σ(X t W xz +H t-1 W hz +b z )

[0071] in, is the input feature sequence of d rows and n columns, is a weight matrix with n rows and h columns, is the h-order weight matrix, is a 1-row, h-column bias matrix, with variable j∈{r,z};

[0072] The candidate hidden states are:

[0073]

[0074] in, is the weight parameter, Is the bias parameter. Hidden state combined with update gate Z t The influence of the previous hidden state H t-1 and candidate hidden states Determined together, it can be expressed as

[0075]

[0076] GRU is used to build a prediction network based on the time domain to improve the processing efficiency and accuracy of the prediction network. In this invention, the above-mentioned prediction network based on the time domain is mainly used for multi-channel prediction. Figure 2 .

[0077] Step 2-2: Use the training data to train the channel prediction network. The trainable parameters of the channel prediction network are randomly initialized, the maximum number of iterations of the prediction network training is set to 20, the initial learning rate is set to 0.001, and the training data is input into the network in batches, with the batch size set to 64. During the training process, the mean square error (MSE), one of the commonly used evaluation indicators, is used as the loss function, and the calculation formula is: in is the prediction result of the i-th group of samples, y i is the true value of the i-th sample, and N is the total number of samples. During network training, the Adam optimization algorithm is used as the optimizer, with a weight decay of 0.0001. The prediction network uses the directional propagation of the training error of each batch to optimize the trainable parameters of the entire channel prediction network. When all batches of training data have completed backpropagation, it is considered one iteration. This process is repeated until the maximum number of training iterations is reached, completing the channel prediction network training.

[0078] Step 2-3: Input the test set data into the channel prediction network to obtain the channel prediction results.

[0079] In the embodiment, the specific steps of step 3 for establishing a dynamic spectrum sharing network environment are as follows:

[0080] Step 3-1: Construct the input state vector of each secondary user agent. At time t, the input state vector of the mth secondary user agent is expressed as

[0081]

[0082] It mainly consists of two parts. One part is the channel data at time t, which is expressed as

[0083]

[0084] The other part is the channel data prediction result at time t+i, which is expressed as

[0085]

[0086] Among them, h m,k (t) represents the channel data of the mth secondary user in the kth channel at time t; h pre,m,k (t+i) represents the channel data prediction result of the mth secondary user in the kth channel at time t+i.

[0087] Step 3-2: Construct a reward function. The reward function can be divided into three cases.

[0088] In the first case, when the mth secondary user agent does not access any channel or the transmission power is 0, r m,t =0.

[0089] In the second case, when the mth secondary user agent chooses to access the kth channel, the transmission power is P k,m , and the interference to the primary user of the kth channel is within the threshold, that is, ρ k,m |h S,k,m (t)| 2 P k,m ≤Γ k , then the reward function is

[0090] r m,t =R k,m -μ(Γ k -ρ k,m |h S,k,m (t)| 2 P k,m )

[0091] Among them, μ is the interference penalty factor, h S,k,m (t) is the channel gain from the secondary base station to the primary user of the kth channel, and the binary spectrum allocation index ρ k,m Assign strategy to the mth secondary user on the kth channel, R k,m is the available rate of the mth secondary user on the kth channel.

[0092] In the third case, the interference to the kth channel primary user exceeds the threshold, and when the mth secondary user agent cannot access the channel, the reward function is

[0093] r m,t =μ(Γ k -(1-ρ k,m )|h S,k,m (t)| 2 P k,m )

[0094] At this point, the total accumulated rewards for all users is

[0095] In the embodiment, the mixing action in step 4 represents the following specific steps:

[0096] Step 4-1: Build a mixed action representation model. Figure 3 ,The hybrid action representation model mainly consists of an embedding table and a variational autoencoder, which constructs a latent ,mixture space and embeds the dependency between discrete and continuous actions.

[0097] Embedded Table express discrete actions, where ζ is a learnable parameter, and each row e of the embedding table ζ,a =E ζ (a) A d1-dimensional continuous vector corresponding to a discrete action a. For continuous parameters, a variational autoencoder is used to construct a latent representation space. The row vectors in the embedding table and state s t As a condition, the encoder The continuous parameters Mapping to low-dimensional latent representation encoder The output adopts Gaussian distribution where μ x and σ x Represent the mean and standard deviation respectively. The encoding process is expressed as

[0098]

[0099] Under the same conditions, the decoder From z t Reconstructed For low-dimensional latent representation, in is the fully connected layer for reconstruction, is a conversion network. The decoding process is expressed as

[0100]

[0101] in, is any potential vector of dimension d1;

[0102] The state dynamic predictor is introduced to learn the dynamic changes of the environment, solving the problem that the potential representation space cannot distinguish the different effects of mixed actions on the environment, so as to further optimize the mixed action representation. The state residual is expressed as The predictor is represented as

[0103]

[0104] in, A cascade structure is used to predict the state residual.

[0105] Step 4-2: Pre-training of mixed action representations. Use a random strategy to collect samples into the experience pool, and then randomly extract samples of a set batch size from the experience pool for network training.

[0106] To train the hybrid action representation model, the trainable parameters of the hybrid action representation network are randomly initialized, and samples are collected using a random strategy before starting the hybrid action representation network pre-training. To the experience pool

[0107] Setting the mixed action representation network training maximum number of iterations is 2000, the initial learning rate is set to 0.0001, the batch size is set to 64, and samples of the set batch size are randomly drawn from the experience pool.

[0108] During training, the loss function of the mixed action representation is calculated and expressed as

[0109]

[0110] in, Encode the result for the embedding table, represents the embedding table and the loss function of the variational autoencoder, Represents the state L2 norm squared prediction error. The first term represents the L2 norm squared reconstruction error, and the second term represents the potential representation z t The Kullback-Leibler divergence between the variational posterior distribution of the model and the standard Gaussian prior is calculated. The hybrid action representation network is trained using the Adam optimization algorithm with a weight decay of 0.0001. The hybrid action representation network optimizes the trainable parameters of the entire network using the directional propagation of training errors from each batch. One training round consists of a batch of samples randomly drawn from the experience pool. This process is repeated until the maximum number of training iterations is reached, completing the pre-training of the hybrid action representation network.

[0111] In the embodiment, the specific steps of step 5 of training the intelligent resource allocation network based on potential learning are as follows:

[0112] Intelligent resource allocation network training based on potential learning. According to the input state s at the current moment t , obtain action embedding representation z t , using the pre-trained mixed action representation network to obtain the channel selection result a t And power allocation results Then calculate the reward r(t) according to the reward function and switch to the state s at the next moment t+1 , and repeat this cycle.

[0113] like Figure 4 As shown, this patent proposes an intelligent resource allocation framework based on a prediction network, which mainly introduces the channel data prediction result at time t+i into the input state vector of the mth secondary user agent. Figure 5 , jointly train an intelligent resource allocation network based on hybrid action representation. In principle, the multi-agent reinforcement learning framework based on hybrid action representation does not rely on a specific algorithm. Any deep reinforcement learning algorithm designed for continuous control can be used based on hybrid action representation.

[0114] This patent proposes a new multi-agent deep deterministic policy gradient (MDDPG) method, which combines hybrid action representation to realize intelligent resource allocation network training. The MDDPG trainable parameters are randomly initialized, and the maximum number of iterations of the network is 50. For any m-th secondary user agent, DDPG consists of an actor network μ(s m,t κ m ), critic network μ(s m,t ;k m ), actor target network and critic target network During the training process, MDDPG is based on the current input state s t , obtain action embedding representation With z t , using the pre-trained mixed action representation network to obtain the channel selection result a t And power allocation results Then calculate the reward r(t) according to the reward function, complete a single environment interaction, and switch to the state s at the next moment t+1 , and the samples To the experience pool MSE is used as the loss function, and the actor network and critic network use the Adam optimization algorithm as the optimizer, with a weight decay of 0.0003. To improve the stability of training, the target network parameters are soft-updated so that the target network slowly approaches the learned network. The update process of the actor target network and the critic target network is as follows:

[0115]

[0116] Among them, τ is a hyperparameter between 0 and 1, and τ<<. κ m and η m are the weight parameters of the actor network and the critic network respectively. and are the weight parameters of the actor target network and the critic target network respectively.

[0117] The effects of the present invention are further described below in conjunction with simulation experiments:

[0118] 1. Simulation conditions and parameter settings:

[0119] The simulation experiments of the present invention were implemented on the Python 3.8, Pytorch 2.2.1 simulation platform. The computer CPU model was Intel i5-13600KF, equipped with an independent graphics card NVIDIA GeForce GTX 1660. In the channel prediction network, 8000 time slot channel data were used for multidimensional channel prediction. The maximum number of training iterations was 10, the initial learning rate was 0.005, the weight decay was 0.0001, the batch size was 64, and the Adam optimization algorithm was selected as the network training optimizer. In the intelligent resource allocation network, the number of hidden layers of the hybrid action representation was 256, the batch size was 64, and the Adam optimization algorithm was selected as the network training optimizer.

[0120] 2. Simulation content:

[0121] Figure 6 This is a comparison chart of the total transmission rate between the present invention and the prior art when the number of users is 2 and under different conditions. Figure 6The paper includes three methods: a random resource allocation method (i.e., the Random method), a multi-agent reinforcement learning method based on hybrid action representation (i.e., the HAR-MDDPG method), and a multi-agent reinforcement learning method based on hybrid action representation combined with a channel prediction network (i.e., the HAR-MCP-MDDPG method). The horizontal axis in the figure represents different episodes, and the vertical axis represents the total transmission rate of the system. It can be seen that compared with the baseline Random method, the HAR-MDDPG method and the HAR-MCP-MDDPG method proposed in this paper achieve the highest sum rate. This shows that the hybrid action representation model can fully learn the close relationship between channel selection and transmit power. Moreover, the HAR-based method takes advantage of the hybrid action representation and allows the underlying strategy to be learned from compact semantics. In addition, the proposed HAR-MDDPG method without updating the hybrid action representation has greater fluctuations than the HAR-MCP-MDDPG method without updating the hybrid action representation. This is because during the hybrid action representation initialization process, the prediction results are beneficial to the HAR-MCP-MDDPG method, which enables the method to effectively learn the underlying channel characteristics.

[0122] Figure 7 This is a comparison chart of the cumulative rewards of the present invention and other existing technologies when the number of users is 2 and the capacity of the experience pool is different. Figure 7 (a) and (b) are the comparison diagrams of the cumulative rewards of the HAR-MDDPG method and the HAR-MCP-MDDPG method when the number of users is 2 and the capacity of the experience pool is different. The size of the experience pool is directly related to the calculation speed and convergence speed of the method. When , the proposed method converges quickly. Smaller memory may speed up the calculation, but it will reduce the performance, such as Figure 7 As shown by the dotted line in (b), it may even prevent the algorithm from converging, such as Figure 7 As shown by the dotted line in (a). Larger memory does not necessarily lead to better performance, it may cause the reward to drop sharply and lead to slower convergence, such as Figure 7 Therefore, setting the appropriate memory is also very important. Figure 7 As can be seen in (a) and (b), the HAR-MCP-MDDPG method achieves a higher cumulative reward than the HAR-MDDPG method. This is because the HAR-based model establishes a dependency between discrete actions (channel selection) and continuous actions (power allocation) in the hybrid action, which is crucial for determining the optimal resource allocation method. At the same time, the HAR-MCP-MDDPG method achieves the highest cumulative reward because it can exploit the dependency and time-varying characteristics of the channel.

[0123] Based on the above simulation results and analysis, the channel prediction enhanced intelligent resource allocation method based on hybrid action representation proposed in the present invention can enable the spectrum sharing system to obtain the maximum transmission rate, achieve a higher system transmission rate than the existing method, and achieve faster network convergence speed, which makes the present invention promising for application in actual communication scenarios.

[0124] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A potential learning intelligent resource allocation method for 6G networks, characterized in that: include: Step 1: Collect and normalize 6G network channel data to construct a dataset; Step 2: Use the data set to train the channel prediction network and obtain the channel prediction results; Step 3: Construct the input state vector of each secondary user agent based on the channel data and channel prediction results, and construct the reward function; Step 4: Use the embedding table and variational autoencoder to build a hybrid action representation network and perform pre-training; Step 5: Train a multi-agent reinforcement learning network based on the input state vector, reward function, and pre-trained hybrid action representation network of each secondary user agent for intelligent resource allocation.

2. The method for allocating potential learning intelligent resources for 6G networks according to claim 1, characterized in that: The channel data in step 1 includes the channel gain from the primary base station to the user and the channel gain from the secondary base station to the user, wherein the channel gain from the primary base station to the user includes the channel gain from the primary base station to each primary user and the channel gain from the primary base station to each secondary user on each sub-frequency band; the channel gain from the secondary base station to the user includes the channel gain from the secondary base station to each secondary user and the channel gain from the secondary base station to each primary user on each sub-frequency band.

3. The method for allocating potential learning intelligent resources for 6G networks according to claim 1, characterized in that: The dataset described in step 1 is constructed as follows: the normalized channel data is constructed into a multidimensional channel data matrix, where the ordinate is the number of channels and the abscissa is the time scale. The multidimensional channel data of the first five moments are taken as input, and the multidimensional channel data of the last moment is used as the label value of the channel prediction data. The sliding window size is set to the length of the prediction data, and the sliding window is used to construct the dataset.

4. The method for allocating potential learning intelligent resources for 6G networks according to claim 1, wherein: The channel prediction network in step 2 is a GRU network, which resets the gate R at time t. t and update gate Z t Respectively expressed as: R t =σ(X t W xr +H t-1 W hr +b r ) Z t =σ(X t W xz +H t-1 W hz +b z ) in, is the input feature sequence of d rows and n columns, is a weight matrix with n rows and h columns, is the h-order weight matrix, is a 1-row, h-column bias matrix, with variable j∈{r,z}; The candidate hidden states are: in, is the weight parameter, is the bias parameter; The effect of the hidden state combined with the update gate is: Among them, H t-1 is the hidden state at the previous moment.

5. The method for allocating potential learning intelligent resources for 6G networks according to claim 1, characterized in that: In step 3, at time t, the input state vector of the mth secondary user agent is constructed as: The channel data s at time t m (t) is: Channel data prediction result s at time t+i pre,m (t+i) is: Among them, h m,k (t) represents the channel data of the mth secondary user in the kth channel at time t; h pre,m,k (t+i) represents the channel data prediction result of the mth secondary user in the kth channel at time t+i.

6. The method for allocating potential learning intelligent resources for 6G networks according to claim 1, characterized in that: In step 3, construct the reward function where r m,t is the reward of the mth secondary user agent at time t; When the mth secondary user agent does not access any channel, or the transmission power P k,m When r is 0, m,t =0; When the mth secondary user agent chooses to access the kth channel, and the interference to the primary user of the channel is ρ k,m |h S,k,m (t)| 2 P k,m Less than or equal to the threshold Γ k When r m,t =R k,m -μ(Γ k -ρ k,m |h S,k,m (t)| 2 P k,m ); When the mth secondary user agent cannot access the channel and the interference to the kth channel primary user is ρ k,m| h S,k,m (t)| 2 P k,m Exceeding the threshold Γ k When r m,t =μ(Γ k -(1-ρ k,m )|h S,k,m (t)| 2 P k,m ); Where μ is the interference penalty factor, h S,k,m (t) is the channel gain from the secondary base station to the primary user of the kth channel, and the binary spectrum allocation index ρ k,m Assign strategy to the mth secondary user on the kth channel, R k,m is the available rate of the mth secondary user on the kth channel.

7. The method for allocating potential learning intelligent resources for 6G networks according to claim 1, characterized in that: The hybrid action representation network constructed in step 4 uses an embedding table and a variational autoencoder to construct a potential hybrid space and embed the dependency between discrete actions and continuous actions; wherein, the embedding table express discrete actions, ζ is a learnable parameter, and each row of the embedding table e ζ,a =E ζ (a) A d1-dimensional continuous vector corresponding to a discrete action a; in the variational autoencoder, the encoder Embed the row vectors in the table and state s t As a condition, the continuous parameter Mapping to a low-dimensional latent representation of d2 dimensions encoder The encoding process is expressed as: Under the same conditions, the decoder From z t Reconstructed in is the fully connected layer for reconstruction, is a conversion network, and the decoding process is expressed as: is any potential vector of dimension d1, a t Select the result for the channel.

8. The method for allocating potential learning intelligent resources for 6G networks according to claim 7, characterized in that: The hybrid action representation network also introduces the dynamic changes of the state dynamic predictor learning environment to optimize the hybrid action representation. Specifically, for any sample The state residual is expressed as The state dynamics predictor is expressed as: in, Indicates that the state residual is predicted using a cascade structure.

9. The method for allocating potential learning intelligent resources for 6G networks according to claim 8, characterized in that: Step 4 uses a random strategy to collect samples To the experience pool Then randomly select from the experience pool Samples of set batch size are extracted from the network for pre-training, and the mixed action representation loss function used in the pre-training process is: in, is the power allocation result, e at Encode the result for the embedding table, represents the embedding table and the loss function of the variational autoencoder, Represents the state L2 norm squared prediction error.

10. The method for allocating potential learning intelligent resources for 6G networks according to claim 1, characterized in that: The multi-agent reinforcement learning network includes an actor network, a critic network, an actor target network and a critic target network. During the training process of the multi-agent reinforcement learning network, according to the input state s at the current moment, t , obtain action embedding representation With z t , using the pre-trained mixed action representation network to obtain the channel selection result a t And power allocation results Then calculate the reward r(t) according to the reward function, complete a single environment interaction, and switch to the state s at the next moment t+1 , and so on; During the training of the multi-agent reinforcement learning network, the target network parameters are soft-updated. The update process of the actor target network and the critic target network is: Among them, τ is a hyperparameter between 0 and 1, and τ<<1; k m and η m are the weight parameters of the actor network and the critic network respectively; and are the weight parameters of the actor target network and the critic target network respectively.

Citation Information

Patent Citations

  • Intelligent resource allocation algorithm for spectrum sharing security communication gateway under assistance of intelligent reflecting surface

    CN117880843A