Hybrid Access Cognitive Radio Network Slice Resource Allocation Method and Device
By building an agent in a cognitive radio network and combining a graph convolutional neural network and reinforcement learning algorithm, the transmission power and channel occupation of cognitive users are optimized, and the problem of low spectrum resource utilization efficiency is solved, achieving more efficient resource allocation and network energy efficiency.
Patent Information
- Application Number
- CN202111365101.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-11-17
AI Technical Summary
The prior art is difficult to effectively optimize the efficiency of spectrum resource usage, especially in complex cognitive radio network environments, and traditional optimization methods are difficult to cope with complex network environments and high-dimensional optimization parameters.
The hybrid access cognitive wireless network slice resource allocation method is adopted, and by constructing an agent and combining graph convolution neural networks and reinforcement learning algorithms, the transmission power and channel occupation of cognitive users are optimized, which meets the communication delay and transmission rate requirements of different service slices, and at the same time comply with the interference temperature threshold constraints.
This improves spectrum usage efficiency, realizes better resource allocation strategies, and enhances the energy efficiency and stability of the network.
Smart Images

Figure CN114095940B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and in particular, to a method and device for resource allocation of a hybrid access cognitive radio network slice. Background Art
[0002] With the development of information technology, the industrial mode has also changed, and spectrum resources have become increasingly precious. On the one hand, with the continuous development of emerging wireless communication services, spectrum resources are becoming increasingly tense; on the other hand, the spectrum resource utilization efficiency of existing inefficient legacy systems is low and it is difficult to transform. In the past decade, the use of wireless devices (such as vehicles, mobile phones, tablets, and various wireless sensors) has increased rapidly, which has promoted the development of the fifth-generation wireless communication (5th generation mobile networks, 5G). In the 5G wireless network, it is expected that the data rate will be 10 times the current data rate, and it has stronger connectivity and 100% coverage, and is expected to provide better quality of service and user experience. At present, the relevant LTE-G230 (wireless private network technology) standard has been released, and the complex spectrum characteristics of the 230MHz band therein bring new challenges to the sharing and common use of multiple industries and services. Although the state has authorized frequency bands for various departments such as State Grid and water conservancy to carry out their respective businesses, there are still some frequency bands that are used on demand by various regions and departments, which will cause most of the spectrum to be in an idle state and the frequency utilization rate is very low. Optimizing the spectrum usage rules to improve the spectrum usage efficiency has become an urgent problem to be solved.
[0003] Network slicing technology is one of the key technologies of 5G. It can uniformly allocate a large number of dedicated underlying physical network resources, virtualize network functions, provide efficient services for users, and specifically meet the diverse needs of users. Along with the development of the times and the progress of technology, many new and high-standard service requirements have emerged. Therefore, integrating existing physical network resources and allocating them to different communication services on demand not only reduces the construction cost but also can reasonably allocate limited resources to multiple service slices. Based on this, 5G proposes typical application scenarios, namely enhanced mobile broadband (eMBB), ultra-reliable low-latency communication (URLLC), and massive Internet of Things communication.
[0004] In addition, cognitive radio (CR) is a technology used to optimize spectrum usage rules. In CR, if it does not affect the normal communication of authorized users, unauthorized users can be allowed to access the authorized spectrum area for communication. Nowadays, the structure of wireless networks is becoming increasingly complex, and the interference brought by dense network coverage cannot be underestimated. In cognitive radio networks, complex network environments, infinite state spaces, and high-dimensional optimization parameters pose a challenge to traditional optimization methods.
[0005] As an important branch of machine learning, reinforcement learning plays a huge role in the decision-making and optimization of complex problems, such as defeating top masters in chess games, allocating and scheduling network resources, and making intelligent recommendations based on user interests. With the development of artificial intelligence technology, various problems can be solved driven by data and algorithms. Compared with traditional algorithms that manually select features, deep learning has greater potential in wireless networks. Deep learning can automatically extract features from data, without much manual intervention, and directly use an end-to-end approach for training, thus reducing the model complexity. Reinforcement learning can continuously interact with the environment, adopt a "trial and error" method, and continuously accumulate experience to adjust strategies. By combining deep learning and reinforcement learning algorithms, historical experience data can be accumulated as training data for the neural network during the learning process, leveraging the advantages of deep learning to better train the model and optimize decisions. The graph convolutional neural network has stronger information extraction ability than ordinary convolutional neural networks in complex topology graph scenarios. Combining the graph convolutional neural network with the reinforcement learning algorithm can effectively improve the information extraction ability of the agent in complex topology scenarios and obtain a better resource allocation strategy.
[0006] However, the complexity of real-world application problems is very high and the amount of information is also large. It may not be possible to grasp the global information at the same time. Therefore, the single-agent reinforcement learning algorithm is no longer applicable. Summary of the Invention
[0007] In view of this, the present invention provides a method and device for resource allocation of hybrid access cognitive radio network slices to implement a reinforcement learning method applicable to the resource allocation scenario of hybrid access cognitive radio network slices, thereby improving the spectrum utilization efficiency.
[0008] To achieve the above object, the present invention is implemented by the following solutions:
[0009] According to one aspect of the embodiments of the present invention, a method for resource allocation of hybrid access cognitive radio network slices is provided, including:
[0010] Construct an agent for each cognitive user in the hybrid access cognitive radio network, where the state of the agent corresponds to the information that the current cognitive user needs to interact with other cognitive users in the hybrid access cognitive radio network and the information that does not need to be interacted, and the action of the agent corresponds to the transmission power of the current cognitive user and the channel occupied by the current cognitive user;
[0011] Determine the state of the agent, and input the determined state of the agent into the graph convolutional-based neural network model in the hybrid access cognitive radio network scenario to calculate the attention head for each neighboring cognitive user in the state of the agent and output the action-related information of the agent via the convolutional layer in the neural network model after connecting all the attention heads;
[0012] Based on the action-related information of the agent, obtain the predicted results of the transmission power and the occupied channel of the cognitive user recognized by the agent, and based on the predicted results of the transmission power and the occupied channel of the cognitive user, determine whether the current action of the agent meets the communication delay requirements for ultra-reliable low-latency communication slice users and the transmission rate requirements for enhanced mobile broadband slice users, and whether the value of the interference temperature threshold constraint function meets the interference temperature threshold requirements;
[0013] If the communication delay requirements, transmission rate requirements, and interference temperature threshold requirements are not met, calculate the value of the reward function, and use the value of the reward function to reward or punish the current action of the agent to train the neural network model and the agent; wherein, the reward function is determined based on the interference temperature threshold constraint function and the cognitive user energy efficiency function, and the interference temperature threshold constraint function is a function of the transmission power and the channel determined based on the ratio of the interference power to the channel bandwidth;
[0014] Use the trained neural network model and agent to allocate network slice resources for cognitive users in the hybrid access cognitive radio network.
[0015] In some embodiments, the information that the current cognitive user needs to interact with other cognitive users in the hybrid access cognitive radio network includes: the binary association coefficient between the cognitive user and the cognitive base station, the binary association coefficient between the cognitive user and the channel, the transmission power of the cognitive user, and the binary coefficient indicating whether the cognitive user meets the communication requirements;
[0016] The information that the current cognitive user does not need to interact with other cognitive users in the hybrid access cognitive radio network includes: the signal-to-interference-plus-noise ratio of the cognitive user and the binary association coefficient between the primary user and the channel.
[0017] In some embodiments, determine the state of the agent, and input the determined state of the agent into the graph convolutional neural network model in the hybrid access cognitive radio network scenario, so as to calculate the attention head for each neighbor cognitive user in the state of the agent, and after connecting all the attention heads, output the action-related information of the agent via the convolutional layer in the neural network model, including:
[0018] Determine the state of the agent, and input the determined state of the agent into the graph convolutional neural network model in the hybrid access cognitive radio network scenario, so as to calculate the attention head for each neighbor cognitive user in the state of the agent, and after connecting all the attention heads, output the action-related information of the agent after passing through the non-linear activation function and the convolutional layer in the neural network model in sequence.
[0019] In some embodiments, the attention head calculated for each neighbor cognitive user in the state of the agent is expressed as:
[0020]
[0021] Among them, represents the attention head m of the neighboring cognitive user j of the agent i, and W m represents the weight matrix of the attention head m, and h i represents the output of the neuron corresponding to the cognitive user of the agent i, and h j represents the output of the neuron corresponding to the neighboring cognitive user j, and X +i represents the set composed of the cognitive user of the agent i and its neighboring cognitive users, and k ∈ X +i represents the set X +i the cognitive user k in, and h k represents the output of the neuron corresponding to the neighboring cognitive user k, τ represents the scaling factor, (W m h i ) T represents the transposed matrix of W m h k and (W m h j ) T represents the transposed matrix of W m h j ;
[0022] The output of the convolutional layer in the neural network model is expressed as:
[0023]
[0024] Among them, h i ' represents the output of the convolutional layer, σ represents the non-linear activation function, concatenate[·] represents the concatenation operation, and X +i represents the set composed of the cognitive user of the agent i and its neighboring cognitive users, and j ∈ X +i represents the set X +i the cognitive user j in, and h j represents the set X +i the output of the neuron corresponding to the cognitive user in, and W m represents the weight matrix of the attention head m, and m ∈ M represents the attention head m in the set M of all attention heads of neighboring cognitive users.
[0025] In some embodiments, the current action of the agent is rewarded or punished using the value of the reward function to train the neural network model and the agent, including:
[0026] Calculate the value of the loss function using the value of the reward function, and return the value of the loss function to the neural network model to reward and punish the current action of the agent and train the neural network model and the agent; wherein, the loss function includes a KL gradient regularization term of the attention weight distribution, and the KL gradient regularization term is used to measure the difference between the attention weight distribution corresponding to the connection results of all current attention heads and the target attention weight distribution.
[0027] In some embodiments, the loss function is expressed as:
[0028]
[0029] wherein, L(θ) represents the value of the loss function, θ represents the parameters of the neural network, BS represents the mini-batch number (the number of a batch of data drawn from the data pool), r b represents the value of the reward function in the b-th mini-batch (the b-th data extraction), γ represents the signal-to-interference-plus-noise ratio, Q(s b ', a'; θ) represents the Q value at the next action a', the next state s b ' and the neural network parameters are θ, Q(s, a; θ) represents the Q value at the action a, the state s and the neural network parameters are θ, λ represents the coefficient of the regularization loss, M represents the number of attention heads, represents the weight distribution of the current state and the weight distribution of the next state of the KL divergence value, and represent the attention weight distribution of the attention head m of the agent in the convolutional layer k.
[0030] In some embodiments, the reward function is expressed as:
[0031]
[0032]
[0033]
[0034] wherein, r i represents the reward value for the current action of the agent i, represents the interference temperature threshold constraint function, represents the transmission power when the cognitive user of the agent i is connected to the cognitive base station a and selects the channel k, represents the association coefficient between the cognitive user of the agent i and the cognitive base station a, η i represents the energy efficiency of the cognitive user of the agent i, k ∈ C represents the channel k in the channel set C, S represents the Sigmoid function, ITmax denotes the interference temperature threshold, \(n\in CU\) represents the cognitive user \(n\) in the cognitive user set \(CU\), \(a\in CBS\) represents the cognitive base station \(a\) in the cognitive base station set \(CBS\), represents the association coefficient between the cognitive user of agent \(i\) and the cognitive base station \(a\), represents the association coefficient between the cognitive user of agent \(i\) and the channel \(k\), represents the gain between the cognitive user of agent \(i\) and the channel \(k\), \(k\) cons is the Boltzmann constant, \(B\) represents the channel bandwidth, \(R\) i represents the transmission rate;
[0035] In some embodiments, the signal-to-interference-plus-noise ratio is expressed as:
[0036]
[0037] where, \(\gamma\) n represents the signal-to-interference-plus-noise ratio of the cognitive user \(n\), \(a\in CBS\) represents the cognitive base station \(a\) in the cognitive base station set \(CBS\), \(k\in C\) represents the channel \(k\) in the channel set \(C\), represents the association coefficient between the cognitive base station \(a\) and the cognitive user \(n\), represents the association coefficient between the channel \(k\) and the cognitive user \(n\), represents the channel gain between the cognitive user \(n\) and the cognitive base station \(a\), represents the transmission power when the cognitive user \(n\) and the cognitive base station \(a\) select the channel \(k\), \(n'\in SU\) represents the cognitive user \(n'\) in the cognitive user set \(SU\), represents the association coefficient between the cognitive base station \(a\) and the cognitive user \(n'\), represents the association coefficient between the channel \(k\) and the cognitive user \(n'\), represents the transmission power when the cognitive user \(n'\) and the cognitive base station \(a\) select the channel \(k\), \(m\in PU\) represents the primary user \(m\) in the primary user set \(PU\), represents the association coefficient between the channel \(k\) and the primary user \(m\), \(g\) n represents the channel gain between the cognitive user \(n\) and the primary base station, represents the transmission power when the primary user \(m\) selects the channel \(k\), \(\sigma\) 2 represents the Gaussian white noise;
[0038] The interference temperature threshold constraint function is expressed as:
[0039]
[0040] where, \(n\in CU\) represents the cognitive user \(n\) in the cognitive user set \(CU\), \(a\in CBS\) represents the cognitive base station \(a\) in the cognitive base station set \(CBS\), represents the association coefficient between the cognitive base station \(a\) and the cognitive user \(n\), Denotes the correlation coefficient between channel k and cognitive user n. Denotes the channel gain between cognitive user n and cognitive base station a. Denotes the transmission power when cognitive user n and cognitive base station a select channel k.
[0041] The transmission rate requirement is expressed as:
[0042]
[0043] Where R n Denotes the transmission rate of cognitive user n, and R min Denotes the minimum transmission rate threshold, n ∈ CU eMBB Denotes the set of cognitive users CU for enhanced mobile broadband slice users eMBB The cognitive user n in it;
[0044] The communication delay requirement is expressed as:
[0045]
[0046] Where d denotes the parameter of the Poisson distribution followed by the arrival rate, ξ is a set value, and D max Is the maximum transmission delay, n ∈ CU URLLC Denotes the set of cognitive users CU for ultra-reliable low-latency communication slice users URLLC The cognitive user n in it.
[0047] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method described in any of the above embodiments.
[0048] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in any of the above embodiments.
[0049] The hybrid access cognitive radio network slice resource allocation method, electronic device, and computer-readable storage medium according to the embodiments of the present invention regard each cognitive user as an agent, establish a mapping relationship between reinforcement learning and the hybrid spectrum access cognitive radio network slice scenario, and implement the application of multi-attention mechanism and loss design, realizing a multi-agent reinforcement learning method applicable to the hybrid spectrum access cognitive radio network slice scenario. By designing the reward function based on interference and efficiency, it distinguishes whether interaction is needed and can obtain higher rewards, showing better performance in stability and convergence, and improving the spectrum utilization efficiency. Description of the Drawings
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0051] Figure 1 is a schematic flowchart of a method for resource allocation of a hybrid access cognitive radio network slice according to an embodiment of the present invention;
[0052] Figure 2 is a schematic structural diagram of a resource allocation model for a hybrid access mode cognitive radio network slice according to an embodiment of the present invention;
[0053] Figure 3 is a schematic structural diagram of a GQN algorithm neural network according to an embodiment of the present invention. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following further elaborates on the embodiments of the present invention with reference to the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.
[0055] Existing single-agent reinforcement learning algorithms are not suitable for solving resource allocation problems in complex networks. In multi-agent reinforcement learning, each agent can only observe the local state and does not understand the global information. In this case, learning and training models are closer to reality. For example, in multi-player online games and robot collaborative production, etc., the relationships between agents are diverse, including simple cooperation or simple competition, as well as relationships that involve both cooperation and competition.
[0056] In order to further improve the energy efficiency of cognitive radio networks by controlling the channel association and power allocation of secondary users while ensuring the service requirements of primary users and secondary users. At the same time, a secondary user hybrid spectrum access mechanism is introduced in the cognitive radio scenario, and secondary users can choose to adopt overlay or underlay access modes according to the state of the access channel. The present invention provides a method for resource allocation of a hybrid access cognitive radio network slice to implement multi-agent reinforcement learning applicable to the resource allocation scenario of cognitive network slices in the hybrid access mode, and combines a graph convolutional neural network and a traditional DQN algorithm to solve complex optimization problems and improve spectrum utilization efficiency.
[0057] Figure 1 is a schematic flowchart of a method for resource allocation of a hybrid access cognitive radio network slice according to an embodiment of the present invention. Refer to Figure 1, the network slice resource allocation method of this embodiment may include the following steps:
[0058] Step S110: Construct an agent for each cognitive user in the hybrid access cognitive radio network. Among them, the state of the agent corresponds to the information that the current cognitive user needs to interact with other cognitive users in the hybrid access cognitive radio network and the information that does not need to interact, and the action of the agent corresponds to the transmission power of the current cognitive user and the channel occupied by the current cognitive user;
[0059] Step S120: Determine the state of the agent, and input the determined state of the agent into the graph convolutional neural network model under the hybrid access cognitive radio network scenario to calculate the attention head for each neighbor cognitive user in the state of the agent and output the information related to the action of the agent via the convolutional layer in the neural network model after connecting all the attention heads;
[0060] Step S130: Obtain the transmission power prediction result and the occupied channel prediction result of the cognitive user of the agent according to the information related to the action of the agent, and based on the transmission power prediction result and the occupied channel prediction result of the cognitive user, judge whether the current action of the agent meets the communication delay requirement when it is a user of the ultra-reliable low-latency communication slice and the transmission rate requirement when it is a user of the enhanced mobile broadband slice, and whether the value of the interference temperature threshold constraint function meets the interference temperature threshold requirement;
[0061] Step S140: If the communication delay requirement, the transmission rate requirement, and the interference temperature threshold requirement are not met, calculate the value of the reward function, and use the value of the reward function to reward or punish the current action of the agent to train the neural network model and the agent; among them, the reward function is determined based on the interference temperature threshold constraint function and the cognitive user energy efficiency function, and the interference temperature threshold constraint function is a function of the transmission power and the channel determined based on the ratio of the interference power to the channel bandwidth;
[0062] Step S150: Use the trained neural network model and agent to allocate network slice resources for the cognitive users in the hybrid access cognitive radio network.
[0063] In the above step S110, the information that the current cognitive user needs to interact with other cognitive users in the hybrid access cognitive radio network may include: the binary correlation coefficient between the cognitive user and the cognitive base station, the binary correlation coefficient between the cognitive user and the channel, the transmission power of the cognitive user, and the binary coefficient indicating whether the cognitive user meets the communication requirement; the information that the current cognitive user does not need to interact with other cognitive users in the hybrid access cognitive radio network may include: the signal-to-interference-plus-noise ratio of the cognitive user and the binary correlation coefficient between the primary user and the channel. Among them, the binary correlation coefficient can be represented by a value of 0 or 1 for association and the other value for non-association. The transmission power of the cognitive user may be different or the same according to the different channels selected by the cognitive user and the different cognitive base stations to which it is connected.
[0064] For example, the state of agent i at time t is represented as where represents the state information that the agent needs to interact with, specifically, it can be represented as In, each symbol in turn represents the transmission power of the cognitive user, the binary correlation coefficient between the cognitive user and the cognitive base station, the binary correlation coefficient between the cognitive user and the channel, and the binary coefficient indicating whether the cognitive user meets the communication requirement; the power, correlation coefficient, and whether the service requirement is met among these parameters are the state information that needs to be interacted with. represents the state information that the agent does not need to interact with, specifically, it can be represented as where {γ m} 1*m represents the signal-to-interference-plus-noise ratio matrix of cognitive users with a size of 1*m, represents the binary correlation coefficient matrix between the primary user and the channel with a size of k*m. The primary user SINR and channel occupancy situation can be the state information that does not need to be interacted with.
[0065] In the above step S120, in the neural network model, the discretized transmission power can be used for calculation. In this way, a discretized power prediction result can be output. For example, if the power is divided into multiple levels, the predicted output transmission power can be represented by the corresponding power level. According to the power level, the corresponding value of the transmission power can be obtained.
[0066] In the above step S120, the state of the agent is determined, and the determined state of the agent is input into the graph convolutional neural network model in the hybrid access cognitive radio network scenario, so as to calculate the attention head for each neighbor cognitive user in the state of the agent, and after connecting all the attention heads, output the action-related information of the agent through the convolutional layer in the neural network model. Specifically, it may include the steps: S121, determine the state of the agent, and input the determined state of the agent into the graph convolutional neural network model in the hybrid access cognitive radio network scenario, so as to calculate the attention head for each neighbor cognitive user in the state of the agent, and after connecting all the attention heads, output the action-related information of the agent through the non-linear activation function and the convolutional layer in the neural network model in sequence.
[0067] In this embodiment, the non-linear activation function can be, for example, an MLP layer or a Relu layer. The neural network model may include an input layer (for inputting feature values, that is, the state of the agent), a multi-layer perceptron layer (MLP layer) / Relu non-linear activation function (Relu layer), a convolutional layer, an output layer, etc. The output action-related information of the agent may include power level, channel, etc. The corresponding power can be obtained according to the power level. Multiple parameters in the hybrid access cognitive radio network scenario can be calculated according to the power and the channel, so as to be used for reasonable constraints.
[0068] In the above step S120, for each agent, only the cognitive users (neighbor cognitive users) that can interact with its cognitive users will affect its resource allocation. Therefore, the attention head is only calculated for the neighbors of the agent's cognitive users. Specifically, when implemented, the attention head calculated for each neighbor cognitive user in the state of the agent can be expressed as:
[0069]
[0070] Among them, represents the attention head m of neighbor cognitive user j of agent i, W m represents the weight matrix of attention head m, h i represents the output of the neuron corresponding to the cognitive user of agent i, h j represents the output of the neuron corresponding to neighbor cognitive user j, X +i represents the set composed of the cognitive user of agent i and its neighbor cognitive users, k ∈ X +i represents the set X +i in the cognitive user k, h k represents the output of the neuron corresponding to neighbor cognitive user k, τ represents the scaling factor, (W m h k ) T represents W m h kThe transpose matrix of, (W m h j ) T represents the transpose matrix of W m h j ;
[0071] The output of the convolutional layer in the neural network model can be expressed as:
[0072]
[0073] where h i ′ represents the output of the convolutional layer, σ represents the non-linear activation function, concatenate[·] represents the concatenation operation, X +i represents the set of cognitive user i and its neighbors, W m represents the weight matrix, m ∈ M represents the attention head m in the attention heads of all neighbor cognitive users.
[0074] In these embodiments, the multi-head attention mechanism can be implemented through the graph convolutional layer.
[0075] In the above step S140, the value of the reward function is used to reward and punish the current action of the agent to train the neural network model and the agent. Specifically, it may include the steps of calculating the value of the loss function using the value of the reward function and returning the value of the loss function to the neural network model to reward and punish the current action of the agent and train the neural network model and the agent; where the loss function may include the KL gradient regularization term of the attention weight distribution, where the KL gradient regularization term can be used to measure the difference between the attention weight distribution corresponding to the connection result of all current attention heads (the distribution of the weights used when connecting the attention heads) and the target attention weight distribution. The weight distribution of the attention heads can be optimized through the KL gradient regularization term in the loss function. In addition, the loss function may also include a regular term to optimize the neural network. Furthermore, if the communication delay requirement, transmission rate requirement, and interference temperature threshold requirement are satisfied, a building function can also be calculated and the neural network can be updated. The loss function of this embodiment can be obtained by adding a regularization term to the conventional loss function. By adding time regularization, the cooperation and competition relationship between agents can be stabilized.
[0076] Specifically, when implemented, the loss function can be expressed as:
[0077]
[0078] where L(θ) represents the value of the loss function, θ represents the parameters of the neural network, BS represents the mini-batch number, r b represents the value of the reward function in the b-th mini-batch (mini-batch b), γ represents the signal-to-interference-plus-noise ratio, Q(sb Q(s', a'; θ) represents the Q-value when the next action is a', the next state is s b ', and the neural network parameter is θ. Q(s, a; θ) represents the Q-value when the action is a, the state is s, and the neural network parameter is θ. λ represents the coefficient of the regularization loss, and M represents the number of attention heads. represents the weight distribution of the current state and the weight distribution of the next state of the KL divergence value and represents the attention weight distribution of the m-th attention head of the agent in the k-th convolutional layer.
[0079] In this embodiment, by adding a temporal regularization term to the loss function, the cooperation and competition relationship between agents can be stabilized.
[0080] In the above step S130, the cognitive user needs to use the channel for communication within the interference range acceptable to the primary user. Therefore, the embodiment of the present invention designs an interference temperature threshold function to measure the interference of the cognitive user to the primary user and limit the degree of interference. In addition, the training process of the agent is also a process of optimizing the action output, so that some constraint conditions can be satisfied under the output transmission power and channel, including the constraint of the interference temperature threshold function. Then, the actions of the agent can be rewarded or punished according to the set constraint conditions. For example, if the interference temperature exceeds a certain threshold, a penalty can be given to this action, and vice versa for a reward. In addition, the cognitive user can pursue the maximization of energy efficiency, so the actions can be rewarded or punished towards the goal of maximizing energy efficiency.
[0081] Specifically, in the above step S130, the reward function can be expressed as:
[0082]
[0083]
[0084]
[0085] where r i represents the reward value for the current action of agent i, represents the interference temperature threshold constraint function, represents the transmission power when the cognitive user of agent i is connected to the cognitive base station a and selects channel k, represents the association coefficient between the cognitive user of agent i and the cognitive base station a, η i represents the energy efficiency of the cognitive user of agent i, k∈C represents the channel k in the channel set C, S represents the Sigmoid function, IT maxdenotes the interference temperature threshold, n∈CU denotes cognitive user n in the cognitive user set CU, a∈CBS denotes the cognitive base station in the cognitive base station set CBS, denotes the association coefficient between the cognitive user of agent i and cognitive base station a, denotes the association coefficient between the cognitive user of agent i and channel k, denotes the gain of the cognitive user of agent i and channel k, k cons is the Boltzmann constant, B represents the channel bandwidth, R i denotes the transmission rate.
[0086] In the specific implementation manner, the signal-to-interference-plus-noise ratio can be expressed as:
[0087]
[0088] where, γ n denotes the signal-to-interference-plus-noise ratio of cognitive user n, a ∈ CBS denotes cognitive base station a in the cognitive base station set CBS, k ∈ C denotes channel k in the channel set C, denotes the association coefficient between cognitive base station a and cognitive user n, denotes the association coefficient between channel k and cognitive user n, denotes the channel gain between cognitive user n and cognitive base station a, denotes the transmission power when cognitive user n and cognitive base station a select channel k, n' ∈ SU denotes cognitive user n' in the cognitive user set SU, denotes the association coefficient between cognitive base station a and cognitive user n’, denotes the association coefficient between channel k and cognitive user n’, denotes the transmission power when cognitive user n’ and cognitive base station a select channel k, m ∈ PU denotes primary user m in the primary user set PU, denotes the association coefficient between channel k and primary user m, g n denotes the channel gain between cognitive user n and the primary base station, denotes the transmission power when primary user m selects channel k, σ denotes Gaussian white noise;
[0089] The interference temperature threshold constraint function can be expressed as:
[0090]
[0091] where, n∈CU denotes cognitive user n in the cognitive user set CU, a∈CBS denotes cognitive base station a in the cognitive base station set CBS, denotes the association coefficient between cognitive base station a and cognitive user n, Denote the correlation coefficient between channel k and cognitive user n. Denote the channel gain between cognitive user n and cognitive base station a. Denote the transmission power when cognitive user n and cognitive base station a select channel k.
[0092] The transmission rate requirement can be expressed as:
[0093]
[0094] where R n Denote the transmission rate of cognitive user n, and R min Denote the minimum transmission rate threshold, and n ∈ CU eMBB Denote the set CU of cognitive users that are users of the enhanced mobile broadband slice eMBB for cognitive user n in it.
[0095] The communication delay requirement is expressed as:.
[0096]
[0097] where d denotes the parameter of the Poisson distribution followed by the arrival rate, ξ is a set value, and D max is the maximum transmission delay, and n ∈ CU URLLC Denote the set CU of cognitive users that are users of the ultra-reliable low-latency communication slice URLLC for cognitive user n in it.
[0098] In the above embodiments, the M / M / 1 queuing model (using queuing theory to calculate the user's computing delay) is used to add the transmission delay constraint of the user. Through the probability formula constrain the communication delay, and thus the formula for the above communication delay requirement can be obtained.
[0099] Furthermore, in the above step S150, after training the agent and the neural network, when applying them to network slice resource allocation, the parameters in the neural network, such as the number of cognitive users, can be set according to the actual scenario to utilize the agent to give the optimal network resource allocation result.
[0100] In the embodiments of the present invention, the cognitive radio network scenario includes a primary base station and multiple secondary base stations; primary users are connected to the primary base station, and secondary users are connected to the secondary base stations. According to different user communication services, primary users and secondary users are divided into enhanced mobile broadband (eMBB) slice users (with the lowest communication rate requirement) and ultra-reliable low-latency communication (URLLC) slice users (with the maximum communication latency requirement). Under the condition of ensuring the normal communication of primary users, secondary users select the overlay network or underlay network access mode to share the spectrum with primary users according to the state of the access channel. The spectrum resources are divided into a set number of channels. It can be set that when there are no primary users in the channel accessed by cognitive users, cognitive users access the spectrum in the overlay mode; when accessing the spectrum in the underlay mode, cognitive users need to use the channel for communication under the condition of meeting the interference temperature constraint, so that cognitive users can use the channel for communication within the interference range acceptable to primary users. In the hybrid access cognitive radio network, the parameters involved in cognitive users include: transmission power, channel gain with the cognitive base station, channel gain with the primary base station, association coefficient with the cognitive base station, and association coefficient with the channel. Cognitive users can occupy only one channel. The parameters involved in primary users include: transmission power, channel gain with the cognitive base station, channel gain with the primary base station, and association coefficient with the channel. Primary users can occupy two channels (randomly or continuously for a certain period of time). The signal-to-interference-plus-noise ratio (SINR) of cognitive users can be calculated based on the above parameters. According to the Shannon channel formula, the transmission rate of cognitive users can be calculated based on the transmission power of cognitive users. For eMBB slice users, a minimum transmission rate threshold can be set, and for URLLC slice users, a maximum transmission latency threshold can be set. For calculating the transmission latency of users, the M / M / 1 queuing model can be used. The communication requirements of URLLC slice users can be represented by probability. The interference temperature is defined and can be obtained based on the ratio of interference power to channel bandwidth. The cognitive users in the underlay access mode are correspondingly set with a maximum interference temperature threshold. Various associations can be represented by binary association constraints. The maximum transmission power limit of cognitive users can be set. The optimization of network slice resource allocation can be attributed to the optimization under various constraints, for example, the service demand constraints of eMBB slice users and URLLC slice users, and the interference temperature constraint. Finally, the training of agents, etc., can be carried out based on the graph convolutional reinforcement learning algorithm, in which the multi-head attention mechanism and temporal relation regularization are considered in the neural network. When the model is trained and applied, the input features can be weighted and summed, and then the multiple attention heads of the agent are connected, and after passing through a multi-layer perceptron or a non-linear activation function, the output of the convolutional layer is obtained; among them, in the convolutional layer, the agent is divided into those whose state information needs to interact and those that do not need to interact.The output may include the power of the cognitive user and the selected channel; based on the power and the channel, it can be determined whether the power, the channel, and the interference temperature threshold, etc. meet the constraints. If they do not meet the constraints, the action can be penalized; if they meet the constraints, the action can be rewarded. Other parameters regarding the primary user can be preset in advance.
[0101] In addition, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method described in any of the above embodiments are implemented.
[0102] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any of the above embodiments are implemented.
[0103] The above method will be described below in conjunction with a specific embodiment. However, it should be noted that this specific embodiment is only for better illustrating the present application and does not constitute an improper limitation to the present application.
[0104] It is considered to divide the whole process into two parts: (1) Establish a cognitive radio network slice resource allocation model in the hybrid access mode; (2) Map the cognitive network slice resource allocation problem to reinforcement learning to establish a model, and implement a cognitive network slice resource allocation algorithm based on multi-agent graph convolutional reinforcement learning. The following will explain the two parts separately.
[0105] (1) Establish a cognitive radio network slice resource allocation model in the hybrid access mode
[0106] In order to further improve the energy efficiency of the cognitive network, this embodiment introduces a hybrid spectrum access mechanism, and the secondary user can choose to adopt the overlay or underlay access mode according to the state of the access channel. Figure 2 It is a schematic structural diagram of a cognitive radio network slice resource allocation model in the hybrid access mode according to an embodiment of the present invention. As Figure 2 shown, there is a primary base station and multiple secondary base stations in the cognitive radio network scenario. The primary user is connected to the primary base station, and the secondary user is connected to the secondary base station. According to different communication services of the users, the primary user and the secondary user can be divided into eMBB (enhanced mobile broadband) slice users and URLLC (ultra-reliable low-latency communication) slice users. The eMBB slice users have the lowest communication rate requirement, and the URLLC slice users have the maximum communication latency requirement. On the basis of ensuring the normal communication of the primary user, the secondary user chooses to adopt the overlay (overlay network / coverage network) or underlay (underlying network) access mode according to the state of the access channel to share the spectrum with the primary user, so as to make more efficient use of the spectrum resources.
[0107] Refer again to Figure 2 In the cognitive radio network scenario, there is one primary base station PBS and A (for example, three) cognitive base stations CBS. The primary base station PBS and the cognitive base stations CBS share the same spectrum resources. M (for example, three) primary users PU are connected to the primary base station PBS, and N cognitive users CU are connected to the cognitive base stations CBS. For example, each cognitive base station CBS is connected to three cognitive users CU. The spectrum resources are divided into K channels, K = {1, 2,... K}. The cognitive users CU adopt a hybrid access mode to share the spectrum resources with the primary users PU. When there is a primary user PU in the channel accessed by the cognitive user CU, the cognitive user CU accesses the spectrum in the underlay mode. When there is no primary user PU in the channel accessed by the cognitive user CU, the cognitive user CU accesses the spectrum in the overlay mode. When accessing the spectrum in the underlay mode, the cognitive user CU needs to use the channel for communication within the interference range acceptable to the primary user PU. In this embodiment, the concept of interference temperature is introduced to quantify the interference generated by the cognitive user CU to the primary user PU. The cognitive user CU needs to use the channel for communication on the premise of satisfying the interference temperature constraint.
[0108] The set of cognitive base stations is denoted as CBS = {1, 2,..., A}, and the cognitive users associated with the eMBB slice are denoted as CU eMBB = {1, 2,..., N1}, and the cognitive users associated with the URLLC slice are denoted as CU URLLC = {1, 2,..., N2}, N = N1 + N2, and the set of cognitive users is denoted as CU = {1, 2,..., N}. The primary users associated with the eMBB slice are denoted as PU eMBB = {1, 2,..., M1}, and the primary users associated with the URLLC slice are denoted as PU URLLC = {1, 2,..., M2}, M = M1 + M2, and the set of primary users is denoted as PU = {1, 2,..., M}. The set of channels is denoted as C = {1, 2,..., K}, the bandwidth of each channel is B, and the total bandwidth is W = K * B.
[0109] The transmission power of cognitive user n is denoted as The channel gain with cognitive base station a is The channel gain with the primary base station is g n . is the association coefficient between the cognitive base station and the cognitive user, indicates that cognitive user n and cognitive base station a are associated, otherwise indicates not associated. is the association coefficient between the channel and the cognitive user, indicates that cognitive user n and channel k are associated, otherwise indicates not associated. The transmission power of primary user m is denoted as The channel gain with the cognitive base station a is The channel gain with the primary base station is g m . is the channel and primary user correlation coefficient, represents that the primary user m is associated with channel k, otherwise In addition, it is assumed that the primary user can occupy two channels, and the cognitive user can only occupy one channel. The communication behavior of the primary user is simplified to that the primary user will randomly and continuously occupy a certain two channels for a certain period of time.
[0110] The signal-to-interference-plus-noise ratio (SINR) of the cognitive user n can be calculated according to the definition of SINR, as shown in Equation (1):
[0111]
[0112] where γ n represents the signal-to-interference-plus-noise ratio, a ∈ CBS represents the cognitive base station a in the cognitive base station CBS, k ∈ C represents the channel k in the channel set C, represents the correlation coefficient between the cognitive base station and the cognitive user, represents the channel and cognitive user correlation coefficient, represents the channel gain between the cognitive user and the cognitive base station, represents the transmission power of the cognitive user, n' ∈ SU represents the secondary user n' in the secondary user set SU, represents the correlation coefficient between the cognitive base station a and the cognitive user n’, represents the correlation coefficient between the channel k and the cognitive user n’, represents the transmission power when the cognitive user n’ selects the channel k with the cognitive base station a, m ∈ PU represents the primary user m in the primary user set, represents the channel and primary user correlation coefficient, g n represents the channel gain between the cognitive user n and the primary base station, represents the transmission power of the primary user, and σ represents the Gaussian white noise.
[0113] According to the Shannon channel formula R = B·log2(1 + γ), the transmission rate R of the cognitive user n can be obtained n , and then the energy efficiency of the cognitive user can be obtained, as shown in Equation (2):
[0114]
[0115] where η n represents the energy efficiency of the cognitive user n, R n represents the transmission rate R of the cognitive user n n , R n= B·log2(1 + γ), where B represents, and γ represents.
[0116] For rate-sensitive users (eMBB slice users), set the minimum transmission rate threshold R min . For latency-sensitive users (URLLC slice users), set the maximum transmission latency threshold D max .
[0117] To calculate the transmission latency of users, use the M / M / 1 queuing model. Assume that the arrival rate follows a Poisson distribution with parameter d, and the transmission latency of users follows an exponential distribution with parameter R n -d. Use probability to represent the communication requirements of URLLC slice users, where ξ is a very small number. Therefore, the communication requirements of URLLC slice users are D n represents the transmission latency of cognitive user n, and P represents probability.
[0118] The interference temperature is defined as the ratio of interference power to channel bandwidth, denoted as where k cons is the Boltzmann constant, P i is the interference power, and BW is the channel bandwidth. For cognitive users adopting the underlay access mode, set the maximum interference temperature threshold IT max .
[0119] The hybrid access cognitive network slice resource allocation optimization problem can be expressed as shown in Equation (3):
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128] where η n represents the energy efficiency of cognitive users, R n the transmission rate of cognitive users, represents the transmission power of cognitive users; R minDenotes the minimum transmission rate, \(n\in CU\) eMBB Indicates that the cognitive user \(n\) belongs to the eMBB slice user (sensitive to rate); \(d\) represents the parameter of the Poisson distribution followed by the arrival rate, and the transmission delay of the cognitive user follows an exponential distribution with parameter \(R\) n \(-d\) Indicates that the cognitive user \(n\) belongs to the URLLC slice user (sensitive to delay), \(D\) max Denotes the maximum transmission delay, and \(\xi\) represents a parameter. Represents the interference power of each channel, \(B\) represents the bandwidth of each channel, \(k\) cons Is the Boltzmann constant, \(I_T\) max The maximum interference temperature Denotes the channel \(k\) in the channel set \(C\). Represents the maximum transmission power of the cognitive user \(n\) Represents the transmission power of the cognitive user \(n\) corresponding to the cognitive base station \(a\) on the channel \(k\). Represents the correlation coefficient between the channel \(k\) and the primary user \(m\) Denotes the primary user \(m\) in the primary user set \(PU\). \(a\in CBS\) represents the cognitive base station \(a\) in the cognitive base station set \(CBS\) Represents the correlation coefficient between the cognitive base station \(a\) and the cognitive user \(n\)
[0129] Formulas (C1) to (C3) are binary correlation coefficient constraints. Formula (C4) represents the maximum transmission power limit of the cognitive user. Formulas (C5) and (C6) are service demand constraints corresponding to eMBB slice users and URLLC slice users. Formula (C7) is the interference temperature constraint, that is, the tolerance of the primary user to the interference generated by the cognitive user.
[0130] (II) Map the cognitive network slice resource allocation problem to reinforcement learning and establish a model
[0131] The graph convolutional reinforcement learning algorithm has two key technologies, namely the multi-head attention mechanism and the temporal relation regularization. The convolutional kernel adopts the multi-head attention mechanism to capture high-order effective information, so as to better learn the interaction between agents and make the training more stable. For agent \(i\), its neighbors are represented as \(X\) i . For the attention head \(m\), it can be expressed as shown in Equation (4):
[0132]
[0133] Among them, Represents the attention head \(m\) of the neighbor cognitive user \(j\) of agent \(i\), \(W\) m Represents the weight matrix, \(h\) i Represents the output of the neuron, \(X\) +idenotes the set of cognitive user i and its neighbors, τ denotes the scaling factor, (W m h i ) T denotes the transpose matrix of W m h i . The weighted sum of the input eigenvalues is calculated, then the M attention heads of agent i are concatenated, and then passed through a layer of MLP (Multi-Layer Perceptron) or a non-linear ReLU (Rectified Linear Unit / Activation Function) layer to obtain the output of the final convolutional layer, as shown in Equation (5);
[0134]
[0135] where h′ i denotes the output of the convolutional layer, σ denotes the non-linear activation function, concatenate[·] denotes the concatenation operation, X +i denotes the set of cognitive user i and its neighbors, W m denotes the weight matrix, m ∈ M denotes the attention head m among the attention heads of all neighbor cognitive users.
[0136] In addition, since information interaction is required between different agents, in order to reduce the transmission of unnecessary information and increase the proportion of effective information interaction, the state of the agent is divided into two categories (state information that needs to be interacted and state information that does not need to be interacted ). The neural network structure of the graph convolutional reinforcement learning algorithm GQN is as Figure 3 shown. Among different agents Agent1, Agent2,..., AgentN, some need to interact and some do not.
[0137] The second main technology is to propose temporal relation regularization to promote stable cooperation among agents within a certain time. In practical application scenarios, stable cooperation among agents often leads to long-term maximum benefits. Therefore, the KL divergence is used to measure the difference between the current attention weight distribution and the target attention weight distribution to strengthen the stable cooperation among agents. The KL divergence is added as a regularization term to the loss function, as shown in Equation (6):
[0138]
[0139] where L(θ) represents the value of the loss function, θ represents the parameters of the neural network, BS represents the number of mini-batches in the reinforcement learning algorithm, r b$r_b$ represents the reward function value in mini - batch $b$ (the $b$-th mini - batch), $Q(s, a;\theta)$ represents the $Q$ value when the action is $a$, the state is $s$, and the neural network parameters are $\theta$, $s$ represents the state, $s'$ represents the next state, $a$ represents the action, $\lambda$ represents the coefficient of the regularization loss, and $M$ represents the number of attention heads. $D_{KL}(p(s), q(s'))$ represents the KL - divergence value between the weight distribution of the current state and the weight distribution of the next state. $\alpha_{i,k,m}$ represents the attention weight distribution of agent $i$ in attention head $m$ of convolutional layer $k$.
[0140] The smaller the Kullback - Leibler Divergence (KL - divergence), the more likely it is for agents to achieve long - term consistent cooperation, which is very helpful for capturing the characteristics of the cooperation relationship between agents.
[0141] In the scenario of hybrid access cognitive network slicing, the resource allocation problem of all cognitive users is theoretically a complex non - convex optimization problem. To solve this problem, this embodiment proposes a multi - agent reinforcement learning method (CRNGQN algorithm) for the scenario of hybrid access cognitive radio network slicing. The multi - agent reinforcement learning algorithm is based on the GQN algorithm and uses a graph structure to represent the cooperation relationship between agents. In the CRNGQN algorithm, the basic elements of reinforcement learning are set as follows.
[0142] Assume that at time $t$, the states of all agents do not change. First, design the state of agent $i$ as $s_i^t=\left[x_i^t, c_i^t, p_i^t, \mathbf{h}_i^t, \mathbf{G}_i^t, \mathbf{S}_i^t\right]$ is the state information that the agent needs to interact with, including $x_i^t$ is a binary coefficient indicating whether the cognitive user meets the communication requirement, where $x_i^t = 1$ means the cognitive user meets the communication requirement, otherwise $x_i^t = 0$. $c_i^t$ is the state information that the agent does not need to interact with, including
[0143] where, $p_i^t$ represents the transmission power of agent (cognitive user) $i$. $\mathbf{h}_i^t$ represents the association coefficient between agent (cognitive user) $i$ and the cognitive base station. $\mathbf{G}_i^t$ represents the association coefficient between agent (primary user) $i$ and the channel. $\{\gamma m \}$ 1*m $\mathbf{S}_i^t$ represents the SINR matrix (signal - to - interference - plus - noise ratio matrix) of the primary user. $\mathbf{O}_i^t$ represents the channel occupancy matrix of the primary user.
[0144] Secondly, design the action of agent $i$ as the optimization variable of this scenario $a_i^t=\left[\mathbf{h}_i^t, p_i^t\right]$, where $\mathbf{h}_i^t$ represents channel association, which is a discrete variable, while the power The variable is a continuous variable. Similar to the processing method of the traditional DQN algorithm, it is necessary to discretize the power in the CRNGQN algorithm. The value range of the power is shown in Equation (7):
[0145]
[0146] where represents the maximum power of occupying Channel a.
[0147] Finally, design the reward function. Considering that the goal of the system is to maximize η i , and it is necessary to meet the interference temperature constraint requirements. Therefore, in the design of the reward function, actions that do not meet the constraint conditions should be punished, and actions that increase the target value should be rewarded. Considering the above factors, the reward function is designed as shown in Equation (8):
[0148]
[0149] where r i represents the reward value of Agent i, represents the interference temperature threshold constraint function, represents the transmission power of Agent (cognitive user) i, represents the association coefficient between Agent (primary user) i and the channel, and η i represents the energy efficiency of Agent (primary user) i.
[0150] where is the interference temperature threshold constraint function, and the expression is shown in Equation (9):
[0151]
[0152] where is the Sigmoid function, k ∈ C represents Channel k in the channel set C, S represents the Sigmoid function, and IT max represents the interference temperature threshold, n ∈ CU represents Cognitive User n in the cognitive user set CU, a ∈ CBS represents Cognitive Base Station a in the cognitive base station set CBS, represents the association coefficient between the cognitive user of Agent i and the cognitive base station a, represents the association coefficient between the cognitive user of Agent i and Channel k, represents the gain between the cognitive user of Agent i and Channel k, and k cons is the Boltzmann constant, B represents the channel bandwidth, and R i represents the transmission rate.
[0153] The slice resource allocation technology for hybrid access cognitive radio networks in this embodiment combines deep reinforcement learning technology in a complex cognitive network to efficiently select the optimal strategy. The main innovations are as follows: (1) In the 5G network slice scenario, a multi-slice cognitive radio network model based on a hybrid underlay-overlay spectrum access mode is proposed. This model introduces network slicing technology into the multi-service cognitive radio network scenario, where cognitive users can access channels without primary users using the overlay spectrum access mode and access channels with primary users using the underlay spectrum access mode. (2) A multi-objective adaptive deep reinforcement learning framework is proposed for the resource allocation problem in cognitive radio networks. This framework can be applied to various objective functions, such as network throughput, spectrum efficiency, and energy efficiency. (3) A deep reinforcement learning algorithm combined with a graph convolutional neural network is proposed. This algorithm uses GAT to help the DQN agent extract environmental information and introduces a temporal relationship regularization mechanism to improve the stability of the agent relationship. The algorithm classifies the state of the agent according to whether the agent needs to communicate, reducing unnecessary state information interaction and improving the efficiency of information interaction and the algorithm convergence speed.
[0154] In this embodiment, a multi-agent reinforcement learning method for the slice resource allocation scenario of a hybrid access mode cognitive network is implemented. Based on cognitive radio technology, a hybrid spectrum access mechanism is introduced, which can further improve the spectrum utilization rate. For this scenario, a CRNGQN (Cognitive Relay Networks, cognitive relay network, CRN) algorithm based on graph convolutional reinforcement learning is proposed, and state, action, and reward functions are designed specifically, a graph structure is established, and the model is set as three modules: an encoding layer, a graph convolutional layer, and a DQN (deep reinforcement learning) layer, so as to strengthen the cooperation and communication among multiple agents and improve the convergence speed of the model. The influence of the learning rate, the number of graph convolutional layers, and the number of neighbors of the agent on the experimental results is explored through experimental simulations, which proves the feasibility of the algorithm. Compared with the DGN algorithm and the DQN algorithm that do not classify states according to whether interaction is required, the proposed algorithm can obtain higher rewards, convergence speed, and stability. Comparing the hybrid spectrum access mode with the overlay and underlay access modes, the energy efficiency under the proposed spectrum hybrid access mode is better than that of the single overlay and underlay spectrum access modes.
[0155] In the description of this specification, the descriptions with reference to the terms "one embodiment", "a specific embodiment", "some embodiments", "for example", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. The order of steps involved in each embodiment is used to schematically illustrate the implementation of the present invention, and the order of steps is not limited and can be adjusted appropriately as needed.
[0156] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0157] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0158] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps of the process Figure 1 in one or more processes and / or blocks Figure 1 specified functions in one or more blocks.
[0160] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A resource allocation method for a hybrid access cognitive radio network slice, characterized in that, Including: Construct an agent for each cognitive user in the hybrid access cognitive radio network, where the state of the agent corresponds to the information that the current cognitive user needs to interact with other cognitive users in the hybrid access cognitive radio network and the information that does not need to be interacted, and the action of the agent corresponds to the transmission power of the current cognitive user and the channel occupied by the current cognitive user; Determine the state of the agent, and input the determined state of the agent into the graph convolutional neural network model under the hybrid access cognitive radio network scenario to calculate the attention head for each neighbor cognitive user in the state of the agent and output the information related to the action of the agent via the convolutional layer in the neural network model after connecting all the attention heads; Obtain the transmission power prediction result and the occupied channel prediction result of the cognitive user of the agent according to the information related to the action of the agent, and based on the transmission power prediction result and the occupied channel prediction result of the cognitive user, judge whether the current action of the agent meets the communication delay requirement when it is a user of the ultra-reliable low-latency communication slice and the transmission rate requirement when it is a user of the enhanced mobile broadband slice, and whether the value of the interference temperature threshold constraint function meets the interference temperature threshold requirement; If the communication delay requirement, the transmission rate requirement, and the interference temperature threshold requirement are not met, punish the reward function, and use the value of the reward function to reward and punish the current action of the agent to train the neural network model and the agent; where the reward function is determined based on the interference temperature threshold constraint function and the cognitive user energy efficiency function, and the interference temperature threshold constraint function is a function of the transmission power and the channel determined based on the ratio of the interference power to the channel bandwidth; Use the trained neural network model and agent to allocate network slice resources for the cognitive users in the hybrid access cognitive radio network; Among them, using the value of the reward function to reward and punish the current action of the agent to train the neural network model and the agent includes: Calculate the value of the loss function using the value of the reward function, and return the value of the loss function to the neural network model to reward and punish the current action of the agent and train the neural network model and the agent; where the loss function includes the KL gradient regularization term of the attention weight distribution, and the KL gradient regularization term is used to measure the difference between the attention weight distribution corresponding to the connection result of all current attention heads and the target attention weight distribution.
2. The resource allocation method for a hybrid access cognitive radio network slice according to claim 1, characterized in that, The information that the current cognitive user needs to interact with other cognitive users in the hybrid access cognitive radio network includes: the binary correlation coefficient between the cognitive user and the cognitive base station, the binary correlation coefficient between the cognitive user and the channel, the transmission power of the cognitive user, and the binary coefficient indicating whether the cognitive user meets the communication requirement; The information that the current cognitive user does not need to interact with other cognitive users in the hybrid access cognitive radio network includes: the signal-to-interference-plus-noise ratio of the cognitive user and the binary correlation coefficient between the primary user and the channel.
3. The resource allocation method for a hybrid access cognitive radio network slice according to claim 1, characterized in that, Determine the state of the agent, and input the determined state of the agent into the graph convolutional neural network model in the hybrid access cognitive radio network scenario, so as to calculate the attention heads for each neighbor cognitive user in the state of the agent, and after connecting all the attention heads, output the action-related information of the agent via the convolutional layer in the neural network model, including: Determine the state of the agent, and input the determined state of the agent into the graph convolutional neural network model in the hybrid access cognitive radio network scenario, so as to calculate the attention heads for each neighbor cognitive user in the state of the agent, and after connecting all the attention heads, output the action-related information of the agent through the non-linear activation function and the convolutional layer in the neural network model in sequence.
4. The resource allocation method for a hybrid access cognitive radio network slice according to claim 3, characterized in that, The attention head calculated for each neighbor cognitive user in the state of the agent is expressed as: Among them, represents the attention head m of the neighboring cognitive user j of the agent i, W m represents the weight matrix of the attention head m, h i represents the output of the neuron corresponding to the cognitive user of the agent i, h j represents the output of the neuron corresponding to the neighboring cognitive user j, X +i represents the set composed of the cognitive user of the agent i and its neighboring cognitive users, k ∈ X +i represents the set X +i the cognitive user k in, h k represents the output of the neuron corresponding to the neighboring cognitive user k, τ represents the scaling factor, (W m h k ) T represents the transposed matrix of W m h k the transposed matrix of (W m h j ) T represents the transposed matrix of W m h j ; The output of the convolutional layer in the neural network model is expressed as: Among them, h i ′ represents the output of the convolutional layer, σ represents the non-linear activation function, concatenate[·] represents the concatenation operation, and X +i represents the set composed of the cognitive users of agent i and their neighboring cognitive users. j ∈ X +i represents the set X +i represents the cognitive user j in the set X j represents the set X +i represents the output of the neuron corresponding to the cognitive user in the set X m represents the weight matrix of the attention head m, and m ∈ M represents the attention head m in the set M of all the attention heads of neighboring cognitive users.
5. The resource allocation method for a hybrid access cognitive radio network slice according to claim 1, characterized in that, The loss function is expressed as: Among them, L(θ) represents the value of the loss function, θ represents the parameters of the neural network, BS represents the number of mini - batches, and r b represents the value of the reward function in the b - th mini - batch, γ represents the signal - to - interference - plus - noise ratio, Q(s b ', a'; θ) represents the Q - value when the next action is a', the next state is s b ' and the neural network parameters are θ, Q(s, a; θ) represents the Q - value when the action is a, the state is s and the neural network parameters are θ, λ represents the coefficient of the regularization loss, M represents the number of attention heads, represents the weight distribution of the current state and the weight distribution of the next state of the KL divergence value, and represents the attention weight distribution of the m - th attention head of the agent in the k - th convolutional layer.
6. The resource allocation method for a hybrid access cognitive radio network slice according to claim 1, characterized in that, The reward function is expressed as: where r i represents the reward function value for the current action of agent i, represents the interference temperature threshold constraint function, represents the transmission power when the cognitive user of agent i is connected to cognitive base station a and selects channel k, represents the association coefficient between the cognitive user of agent i and cognitive base station a, η i represents the energy efficiency of the cognitive user of agent i, k∈C represents channel k in channel set C, S represents the Sigmoid function, IT max represents the interference temperature threshold, n∈CU represents cognitive user n in cognitive user set CU, a∈CBS represents cognitive base station a in cognitive base station set CBS, represents the association coefficient between the cognitive user of agent i and cognitive base station a, represents the association coefficient between the cognitive user of agent i and channel k, represents the gain between the cognitive user of agent i and channel k, k cons is the Boltzmann constant, B represents the channel bandwidth, R i represents the transmission rate.
7. The resource allocation method for a hybrid access cognitive radio network slice according to claim 1, characterized in that, The signal-to-interference-plus-noise ratio is expressed as: Among them, γ n represents the signal-to-interference-plus-noise ratio of cognitive user n, a ∈ CBS represents cognitive base station a in cognitive base station set CBS, k ∈ C represents channel k in channel set C, represents the correlation coefficient between cognitive base station a and cognitive user n, represents the correlation coefficient between channel k and cognitive user n, represents the channel gain between cognitive user n and cognitive base station a, represents the transmission power when cognitive user n and cognitive base station a select channel k, n' ∈ SU represents cognitive user n' in cognitive user set SU, represents the correlation coefficient between cognitive base station a and cognitive user n', represents the correlation coefficient between channel k and cognitive user n', represents the transmission power when cognitive user n' and cognitive base station a select channel k, m ∈ PU represents primary user m in primary user set PU, represents the correlation coefficient between channel k and primary user m, g n represents the channel gain between cognitive user n and the primary base station, represents the transmission power when primary user m selects channel k, σ 2 represents additive white Gaussian noise; The interference temperature threshold constraint function is expressed as: Among them, n∈CU represents cognitive user n in the cognitive user set CU, a∈CBS represents cognitive base station a in the cognitive base station set CBS, represents the association coefficient between cognitive base station a and cognitive user n, represents the association coefficient between channel k and cognitive user n, represents the channel gain between cognitive user n and cognitive base station a, represents the transmission power when cognitive user n and cognitive base station a select channel k; The transmission rate requirement is expressed as: Among them, R n represents the transmission rate of cognitive user n, and R min represents the minimum transmission rate threshold, where n ∈ CU eMBB denotes the set of cognitive users CU eMBB for the enhanced mobile broadband slice users, and cognitive user n in The communication delay requirement is expressed as: where d represents the parameter of the Poisson distribution followed by the arrival rate, ξ is a set value, D max is the maximum transmission delay, and n ∈ CU URLLC represents the set of cognitive users CU for ultra-reliable and low-latency communication slice users URLLC where n is the cognitive user in CU 8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, on which a computer program is stored, wherein, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Centralized cognitive radio spectrum allocation method based on improved reinforcement learning
CN108809456A
Predictive anti-interference method for industrial wireless network and gateway equipment
CN111446985A