A radio transmission method and apparatus based on deep reinforcement learning
By optimizing channel selection and power allocation using deep reinforcement learning, the problem of resource allocation in radio transmission systems is solved, thereby improving system stability, transmission rate, and user fairness.
Patent Information
- Application Number
- CN202310027753.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-01-09
AI Technical Summary
In existing technologies for radio transmission systems, users have difficulty obtaining full knowledge of channel state information, leading to difficulties in resource allocation. Furthermore, reinforcement learning methods suffer from slow convergence, which affects system performance.
A joint optimization model for channel selection and power allocation is established using a deep reinforcement learning-based approach. By initializing the neural network and memory pool, the neural network parameters are optimized using a greedy strategy and backpropagation algorithm to achieve optimal channel access and power allocation.
This improved the overall stability and transmission rate of the radio transmission system while ensuring fairness in transmission for secondary users.
Smart Images

Figure CN116229693B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, and in particular to a radio transmission method and apparatus based on deep reinforcement learning. Background Technology
[0002] In real life, people's demand for spectrum resources is increasing. Using cognitive radio technology to access users to idle spectrum can improve spectrum utilization. Resource allocation is one of the key technologies of cognitive radio, which improves the overall performance of the system by allocating the best channels and optimizing transmission power.
[0003] Currently, most solutions to resource allocation problems are based on optimal control or game theory, while some utilize model-free strategies in reinforcement learning. However, the prerequisite for using optimal control or game theory is that all users in the wireless network know the state information of all channels, which is difficult to achieve in practical applications. Furthermore, model-free strategies in reinforcement learning also encounter slow convergence, random noise, and measurement errors. Therefore, proposing a radio transmission method based on deep reinforcement learning to improve the overall performance of radio transmission systems is of paramount importance. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a radio transmission method and apparatus based on deep reinforcement learning, which can help to make optimal channel access and power allocation strategies for each user, not only ensuring the fairness of secondary user transmission, but also improving the overall stability and transmission rate of the transmission system.
[0005] To address the aforementioned technical problems, the first aspect of this invention discloses a radio transmission method based on deep reinforcement learning, the method comprising:
[0006] A joint optimization model for channel selection and power allocation is established, and the number of training rounds, memory pool, deep neural network, and parameter set of the deep neural network are initialized. The parameter set includes the initial network parameters of the deep neural network.
[0007] For the current round of training of the joint optimization model, initialize the state of the first agent corresponding to the channel selection;
[0008] The action of the first agent is determined according to a greedy strategy. The state of the second agent corresponding to the power allocation is determined according to the action and state of the first agent. The action of the second agent is then determined according to the greedy strategy.
[0009] The actions of the first agent and the actions of the second agent are input into the deep neural network for analysis, and the reward content returned by the deep neural network is obtained.
[0010] Update the state of the agent, and generate a state transition based on the state of the agent, the action of the agent, the report content, and the updated state of the agent, and store the state transition in the memory pool;
[0011] A preset number of data sets are randomly sampled from the memory pool, and a loss function is calculated based on the data sets. The initial network parameters of the deep neural network are updated based on the loss function and the backpropagation algorithm to obtain the current neural network parameters.
[0012] The current neural network parameters are determined as the initial network parameters of the deep neural network when training the joint optimization model for the next time, and the training operation on the joint optimization model continues until the number of training rounds of the joint optimization model reaches the number of training rounds, and the joint optimization model obtained from the last training is determined as the target joint optimization model;
[0013] Channel and power selection are performed using the aforementioned joint optimization model.
[0014] As an optional implementation, in the first aspect of the present invention, the method further includes:
[0015] The neural network parameters are updated using a soft update method to obtain the optimal network parameters for the current iteration.
[0016] The optimal network parameters for the current iteration are determined as the initial network parameters of the deep neural network for the next training of the joint optimization model. The training operation of the joint optimization model is continued until the number of training iterations of the joint optimization model reaches the number of training rounds. The joint optimization model obtained from the last training iteration is then determined as the target joint optimization model.
[0017] As an optional implementation, in the first aspect of the present invention, the initialization formula for the state of the first agent corresponding to the channel selection is:
[0018] ;
[0019] in, This represents the initial state of the first agent corresponding to the channel selection in the t-th time slot. This indicates the channel occupancy status of the t-th time slot. This represents the channel gain from the second user in the t-th time slot to the cognitive base station. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. Let t represent the channel gain from the primary user to the cognitive base station in the t-th time slot. When t=0, it represents the first initialization of the state of the first agent corresponding to the channel selection. At the beginning of each time slot, the state of the first agent corresponding to the channel selection is initialized, and the state of the first agent after initialization in each time slot is used for training the joint optimization model in that time.
[0020] As an optional implementation, in the first aspect of the present invention, the parameter set further includes a threshold for the greedy strategy;
[0021] Determining the action of the first agent according to a greedy strategy includes:
[0022] The state of the first agent is input into the deep neural network to obtain the first return value of the deep neural network;
[0023] A first probability is randomly generated. When the first probability is less than or equal to the threshold of the greedy policy, the action of the first agent is randomly selected. When the first probability is greater than the threshold of the greedy policy, the action of the first agent is selected according to the first action selection formula.
[0024] The step of determining the state of the second agent corresponding to the power allocation based on the action and state of the first agent, and determining the action of the second agent according to a greedy strategy, includes:
[0025] The state of the second agent corresponding to the power allocation is determined based on the action and state of the first agent.
[0026] The state of the second agent is input into the deep neural network to obtain the second return value of the deep neural network;
[0027] A second probability is randomly generated. When the second probability is less than or equal to the threshold of the greedy policy, the action of the second agent is randomly selected. When the second probability is greater than the threshold of the greedy policy, the action of the second agent is selected according to the second action selection formula.
[0028] As an optional implementation, in the first aspect of the present invention, the formula for selecting the first action is:
[0029] ,
[0030] In the formula, This represents the action of the first agent in the t-th time slot. This represents the set of actions of the first agent in the t-th time slot. This represents the first return value. This represents the state of the first agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the first intelligent agent;
[0031] The formula for selecting the second action is:
[0032] ,
[0033] In the formula, This represents the action of the second agent in the t-th time slot. This represents the set of actions of the second agent in the t-th time slot. This indicates the second return value. This represents the state of the second agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the second agent.
[0034] As an optional implementation, in the first aspect of the present invention, the calculation formula for the reward content is:
[0035] ,
[0036] in, This indicates the access status of the nth secondary user to channel m in the t-th time slot. This represents the channel dryness ratio of the nth user in the t-th time slot on channel m. This represents the linear fairness index for the t-th time slot. This represents the reachable rate of the nth user in the tth time slot. This represents the transmit power of the nth secondary user in the tth time slot. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. This is the interference threshold.
[0037] As an optional implementation, in a first aspect of the invention, updating the state of the agent, generating a state transition based on the agent's state, the agent's action, the report content, and the updated state of the agent, and storing the state transition in the memory pool, includes:
[0038] Update the state of the first agent and the state of the second agent;
[0039] A first state transition is generated based on the state of the first agent, the action of the first agent, the report content, and the updated state of the first agent, and the first state transition is stored in the first memory pool corresponding to the channel selection.
[0040] A second state transition is generated based on the state of the second agent, the action of the second agent, the report content, and the updated state of the second agent. When the time slot is greater than or equal to a preset time slot threshold, the second state transition is stored in the second memory pool corresponding to the power allocation.
[0041] As an optional implementation, in the first aspect of the present invention, the step of randomly sampling a preset number of data sets from the memory pool includes:
[0042] From the first memory pool, a first data set corresponding to the first agent is randomly sampled; from the second memory pool, a second data set corresponding to the second agent is randomly sampled.
[0043] The step of calculating the loss function based on the dataset includes:
[0044] Calculate the first loss function corresponding to the first agent based on the first data set, and calculate the second loss function corresponding to the second agent based on the second data set.
[0045] As an optional implementation, in the first aspect of the present invention, the parameter set further includes the learning rate of the deep neural network;
[0046] The formula for calculating the first loss function is:
[0047] ,
[0048] The formula for calculating the second function is:
[0049] ,
[0050] The formula for the backpropagation algorithm is:
[0051] , ,
[0052] in, This refers to the first data set. This represents the number of samples in the first dataset. This refers to the second data set. This indicates the number of samples in the second dataset. This represents the first output value of the deep neural network corresponding to the j-th sample in the first data set. This represents the second output value output by the deep neural network corresponding to the j-th sample in the second data set. and These represent the states of the first and second intelligent agents, respectively. Take action below The expected cumulative discount benefit obtained, j represents the j-th sample in the first sample set and the second sample set, This represents the state of the j-th sample. This represents the action of the j-th sample. This represents the current neural network parameters corresponding to the first agent. This represents the initial network parameters corresponding to the first agent. This represents the current neural network parameters corresponding to the second agent. This represents the initial network parameters corresponding to the second agent. This represents the learning rate corresponding to the first intelligent agent. This represents the learning rate corresponding to the second intelligent agent.
[0053] As an optional implementation, in the first aspect of the present invention,
[0054] A second aspect of the present invention discloses a radio transmission device based on deep reinforcement learning, the device comprising:
[0055] An initialization module is used to establish a joint optimization model for channel selection and power allocation, and to initialize the number of training rounds, memory pool, deep neural network, and parameter set of the deep neural network for the joint optimization model. The parameter set includes the initial network parameters of the deep neural network.
[0056] The initialization module is also used to initialize the state of the first agent corresponding to the channel selection for the current round of training the joint optimization model.
[0057] The first determining module is configured to determine the action of the first agent according to a greedy strategy, determine the state of the second agent corresponding to the power allocation according to the action and state of the first agent, and determine the action of the second agent according to the greedy strategy.
[0058] An input module is used to input the actions of the first agent and the actions of the second agent into the deep neural network for analysis, and to obtain the reward content returned by the deep neural network;
[0059] An update module is used to update the state of the agent, generate a state transition based on the state of the agent, the action of the agent, the report content, and the updated state of the agent, and store the state transition in the memory pool;
[0060] The sampling module is used to randomly sample a preset number of data sets from the memory pool, calculate a loss function based on the data sets, update the initial network parameters of the deep neural network based on the loss function and the backpropagation algorithm, and obtain the current neural network parameters.
[0061] The first determining module is further configured to determine the current neural network parameters as the initial network parameters of the deep neural network when training the joint optimization model for the next time, and continue to perform training operations on the joint optimization model until the number of training rounds of the joint optimization model reaches the number of training rounds, and determine the joint optimization model obtained from the last training as the target joint optimization model;
[0062] The selection module is used to select the channel and power through the target joint optimization model.
[0063] As an optional implementation, in a second aspect of the invention, the apparatus further includes:
[0064] The update module is also used to update the current neural network parameters according to the soft update method to obtain the current optimal network parameters;
[0065] The second determining module is used to determine the current optimal network parameters as the initial network parameters of the deep neural network when training the joint optimization model for the next time, and to trigger the first determining module to perform the operation of continuing to train the joint optimization model until the number of training rounds of the joint optimization model reaches the number of training rounds, and to determine the joint optimization model obtained from the last training as the target joint optimization model.
[0066] As an optional implementation, in the second aspect of the invention, the initialization formula for the state of the first agent corresponding to the channel selection is:
[0067] ;
[0068] in, This represents the initial state of the first agent corresponding to the channel selection in the t-th time slot. This indicates the channel occupancy status of the t-th time slot. This represents the channel gain from the second user in the t-th time slot to the cognitive base station. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. Let t represent the channel gain from the primary user to the cognitive base station in the t-th time slot. When t=0, it represents the first initialization of the state of the first agent corresponding to the channel selection. At the beginning of each time slot, the state of the first agent corresponding to the channel selection is initialized, and the state of the first agent after initialization in each time slot is used for training the joint optimization model in that time.
[0069] As an optional implementation, in a second aspect of the invention, the parameter set further includes a threshold for a greedy strategy;
[0070] The specific implementation method of the first determining module determining the action of the first agent according to the greedy strategy is as follows:
[0071] The state of the first agent is input into the deep neural network to obtain the first return value of the deep neural network;
[0072] A first probability is randomly generated. When the first probability is less than or equal to the threshold of the greedy policy, the action of the first agent is randomly selected. When the first probability is greater than the threshold of the greedy policy, the action of the first agent is selected according to the first action selection formula.
[0073] The first determining module determines the state of the second agent corresponding to the power allocation based on the action and state of the first agent, and determines the action of the second agent according to the greedy strategy. The specific implementation of this method is as follows:
[0074] The state of the second agent corresponding to the power allocation is determined based on the action and state of the first agent.
[0075] The state of the second agent is input into the deep neural network to obtain the second return value of the deep neural network;
[0076] A second probability is randomly generated. When the second probability is less than or equal to the threshold of the greedy policy, the action of the second agent is randomly selected. When the second probability is greater than the threshold of the greedy policy, the action of the second agent is selected according to the second action selection formula.
[0077] As an optional implementation, in a second aspect of the present invention, the formula for selecting the first action is:
[0078] ,
[0079] In the formula, This represents the action of the first agent in the t-th time slot. This represents the set of actions of the first agent in the t-th time slot. This represents the first return value. This represents the state of the first agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the first intelligent agent;
[0080] The formula for selecting the second action is:
[0081] ,
[0082] In the formula, This represents the action of the second agent in the t-th time slot. This represents the set of actions of the second agent in the t-th time slot. This indicates the second return value. This represents the state of the second agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the second agent.
[0083] As an optional implementation, in a second aspect of the invention, the formula for calculating the reward content is:
[0084] ,
[0085] in, This indicates the access status of the nth secondary user to channel m in the t-th time slot. This represents the channel dryness ratio of the nth user in the t-th time slot on channel m. This represents the linear fairness index for the t-th time slot. This represents the reachable rate of the nth user in the tth time slot. This represents the transmit power of the nth secondary user in the tth time slot. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. This is the interference threshold.
[0086] As an optional implementation, in a second aspect of the invention, the update module updates the state of the agent and generates a state transition based on the state of the agent, the action of the agent, the report content, and the updated state of the agent, and stores the state transition in the memory pool. The specific implementation of this method is as follows:
[0087] Update the state of the first agent and the state of the second agent;
[0088] A first state transition is generated based on the state of the first agent, the action of the first agent, the report content, and the updated state of the first agent, and the first state transition is stored in the first memory pool corresponding to the channel selection.
[0089] A second state transition is generated based on the state of the second agent, the action of the second agent, the report content, and the updated state of the second agent. When the time slot is greater than or equal to a preset time slot threshold, the second state transition is stored in the second memory pool corresponding to the power allocation.
[0090] As an optional implementation, in the second aspect of the present invention, the specific implementation of the sampling module randomly sampling a preset number of data sets from the memory pool is as follows:
[0091] From the first memory pool, a first data set corresponding to the first agent is randomly sampled; from the second memory pool, a second data set corresponding to the second agent is randomly sampled.
[0092] The specific implementation method of the sampling module calculating the loss function based on the data set is as follows:
[0093] Calculate the first loss function corresponding to the first agent based on the first data set, and calculate the second loss function corresponding to the second agent based on the second data set.
[0094] As an optional implementation, in a second aspect of the present invention, the parameter set further includes the learning rate of the deep neural network;
[0095] The formula for calculating the first loss function is:
[0096] ,
[0097] The formula for calculating the second function is:
[0098] ,
[0099] The formula for the backpropagation algorithm is:
[0100] , ,
[0101] in, This refers to the first data set. This represents the number of samples in the first dataset. This refers to the second data set. This indicates the number of samples in the second dataset. This represents the first output value of the deep neural network corresponding to the j-th sample in the first data set. This represents the second output value output by the deep neural network corresponding to the j-th sample in the second data set. and These represent the states of the first and second intelligent agents, respectively. Take action below The expected cumulative discount benefit obtained, j represents the j-th sample in the first sample set and the second sample set, This represents the state of the j-th sample. This represents the action of the j-th sample. This represents the current neural network parameters corresponding to the first agent. This represents the initial network parameters corresponding to the first agent. This represents the current neural network parameters corresponding to the second agent. This represents the initial network parameters corresponding to the second agent. This represents the learning rate corresponding to the first intelligent agent. This represents the learning rate corresponding to the second intelligent agent.
[0102] A third aspect of the present invention discloses another radio transmission device based on deep reinforcement learning, the device comprising:
[0103] Memory containing executable program code;
[0104] A processor coupled to the memory;
[0105] The processor calls the executable program code stored in the memory to execute the radio transmission method based on deep reinforcement learning disclosed in the first aspect of the present invention.
[0106] The fourth aspect of the present invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute the radio transmission method based on deep reinforcement learning disclosed in the first aspect of the present invention.
[0107] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0108] In this embodiment of the invention, the neural network parameters in the established joint optimization model for channel selection and power allocation are optimized and trained. A greedy strategy is used to select the agent's actions. The loss function is calculated based on the state transitions stored in the memory pool. Optimal network parameters are obtained through backpropagation and soft updates. These optimal network parameters are then used to iteratively optimize the joint optimization model for channel selection and power allocation. Finally, the joint optimization model is used to select the channel and power. Therefore, implementing this invention can provide optimal channel access and power allocation strategies for each user, ensuring fairness in secondary user transmission and improving the overall stability and transmission rate of the transmission system. Attached Figure Description
[0109] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0110] Figure 1 This is a flowchart illustrating a radio transmission method based on deep reinforcement learning disclosed in an embodiment of the present invention.
[0111] Figure 2 This is a flowchart illustrating another radio transmission method based on deep reinforcement learning disclosed in an embodiment of the present invention.
[0112] Figure 3 This is a schematic diagram of the structure of a radio transmission device based on deep reinforcement learning disclosed in an embodiment of the present invention;
[0113] Figure 4 This is a schematic diagram of another radio transmission device based on deep reinforcement learning disclosed in an embodiment of the present invention;
[0114] Figure 5 This is a schematic diagram of the structure of another radio transmission device based on deep reinforcement learning disclosed in an embodiment of the present invention;
[0115] Figure 6 This is a schematic diagram of a cognitive radio system model disclosed in an embodiment of the present invention;
[0116] Figure 7 This is a schematic diagram illustrating the impact of different strategies on the average utility function value r under a joint optimization model of channel selection and power allocation, as disclosed in an embodiment of the present invention.
[0117] Figure 8 This invention discloses an interference threshold at the power utilities (PU) under a joint optimization model for channel selection and power allocation. A schematic diagram illustrating the effect on the average utility function value r. Detailed Implementation
[0118] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0119] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or end that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or ends.
[0120] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0121] This invention discloses a radio transmission method and apparatus based on deep reinforcement learning, which can optimize and train the neural network parameters in the established joint optimization model of channel selection and power allocation, and then optimize and iterate the joint optimization model of channel selection and power allocation through the optimal network parameters. Then, the joint optimization model is used to select channels and power, which is beneficial to make the optimal channel access and power allocation strategy for each user. This not only ensures the fairness of secondary user transmission, but also improves the overall stability and transmission rate of the transmission system.
[0122] To better understand the deep reinforcement learning-based radio transmission method and apparatus described in this invention, the cognitive radio system model applicable to the deep reinforcement learning-based radio transmission method and apparatus is first described. Specifically, the cognitive radio system model can be as follows: Figure 6 As shown, Figure 6 This is a schematic diagram of a cognitive radio system model according to an embodiment of the present invention. Figure 6As shown, this cognitive radio system model is a multi-user cognitive radio model, consisting of two parts: the first part is a main network (PUN) composed of M primary users (PUs) and 1 primary station (PBS), where PU-m, m=1,...,M, represents the m-th primary user; the second part is a cognitive radio network (CRN) composed of 1 cognitive base station (CBS) and N secondary users (SUs), where SU-n, n=1,...,N, represents the n-th secondary user. In this model, the PUN covers M orthogonal channels, and each PU corresponds to one orthogonal channel. The CRN network is located within the coverage area of the PUN and operates synchronously with the PUN in time slots.
[0123] like Figure 6 As shown in the cognitive radio system model, the SU accesses the licensed channel via underlay mode. A single SU can only access one channel per time slot to complete its data transmission. When a SU accesses a channel occupied by a PU, the SU's transmit power must be limited to keep the interference from the SU to the PU within interference tolerance. Within the time slot, the CBS performs spectrum sensing at the beginning of each time slot to obtain instantaneous channel state information (CSI), and determines the access channel and transmit power of each SU based on the channel state information. The decision is then broadcast to all SUs through a specific control channel.
[0124] It should be noted that, Figure 6 The cognitive radio system model shown is only intended to illustrate a cognitive radio system model applicable to deep reinforcement learning-based radio transmission methods. The primary user (PU), main station (PBS), secondary user (SU), and cognitive base station (CBS) involved are also merely illustrative. Figure 6 The cognitive radio system model shown is not limited in this respect. Furthermore, the cognitive radio system model applicable to the deep reinforcement learning-based radio transmission method has been described above. The following section provides a detailed explanation of the deep reinforcement learning-based radio transmission method and apparatus.
[0125] Example 1
[0126] Please see Figure 1 , Figure 1 This is a flowchart illustrating a radio transmission method based on deep reinforcement learning disclosed in an embodiment of the present invention. Wherein, Figure 1 The described deep reinforcement learning-based radio transmission method can be applied to radio transmission systems, such as cognitive radio systems, and the embodiments of this invention are not limited thereto. Figure 1 As shown, this deep reinforcement learning-based radio transmission method may include the following operations:
[0127] 101. Establish a joint optimization model for channel selection and power allocation, and initialize the number of training rounds, memory pool, deep neural network, and parameter set of the deep neural network for the joint optimization model. The parameter set includes the initial network parameters of the deep neural network.
[0128] In this embodiment of the invention, optionally, the number of training rounds can be represented as N, and the memory pool corresponding to channel selection can be represented as... The memory pool corresponding to power allocation can be represented as A deep neural network can be represented as a Q-network, and its parameter set can include the initial network parameters of the deep neural network. The initial network parameters corresponding to channel selection can be represented as follows: The initial network parameters corresponding to power allocation can be expressed as: Optionally, the parameter set may also include discount factors, learning rates, thresholds for the greedy strategy, and / or soft substitution coefficients, where the discount factors corresponding to channel selection and power allocation can be expressed as follows: and The learning rates corresponding to channel selection and power allocation can be expressed as follows: and The threshold of the greedy strategy can be expressed as The soft substitution coefficients corresponding to channel selection and power allocation can be expressed as follows: and This embodiment is not limited.
[0129] 102. For the current round of training the joint optimization model, initialize the state of the first agent corresponding to the channel selection.
[0130] In this embodiment of the invention, optionally, the state of the first agent corresponding to channel selection can be represented as: ,in, This indicates the channel occupancy status of the t-th time slot. This represents the channel gain from the second user in the t-th time slot to the cognitive base station. This represents the channel gain from the secondary user to the primary user base station in the t-th time slot. This represents the channel gain from the primary user to the cognitive base station in the t-th time slot, which is not limited in this embodiment.
[0131] 103. Determine the action of the first agent according to the greedy strategy, determine the state of the second agent corresponding to the power allocation according to the action and state of the first agent, and determine the action of the second agent according to the greedy strategy.
[0132] In this embodiment of the invention, optionally, the action of the first intelligent agent can be represented as follows: The state of the second agent can be represented as The actions of the second agent can be represented as .
[0133] 104. Input the actions of the first agent and the actions of the second agent into the deep neural network for analysis, and obtain the reward content returned by the deep neural network.
[0134] In this embodiment of the invention, optionally, the reward content returned by the deep neural network can be represented as... .
[0135] 105. Update the agent's state, and generate a state transition based on the agent's state, actions, feedback content, and the updated state. Store the state transition in the memory pool.
[0136] In this embodiment of the invention, optionally, updating the state of the agent may include changing the state of the first agent from state [state name missing]. Updated to The state of the second agent is changed from state Updated to The state transition can include the state transition corresponding to the first agent. State transitions corresponding to the second agent Storing state transitions in the memory pool includes storing the state transitions corresponding to the first agent in the memory pool. And store the state transition corresponding to the second agent in the memory pool. .
[0137] 106. Randomly sample a preset number of data sets from the memory pool, calculate the loss function based on the data sets, update the initial network parameters of the deep neural network based on the loss function and the backpropagation algorithm, and obtain the current neural network parameters.
[0138] In this embodiment of the invention, optionally, the data set randomly sampled from the memory pool of a preset number includes the data set corresponding to the first agent. and the data set corresponding to the second intelligent agent. The loss function formula for the first agent is: The loss function formula for the second agent is: The formula for the backpropagation algorithm corresponding to the first intelligent agent is: The formula for the backpropagation algorithm corresponding to the second agent is: .
[0139] 107. Determine the current neural network parameters as the initial network parameters of the deep neural network for the next training of the joint optimization model, and continue to perform training operations on the joint optimization model until the number of training rounds of the joint optimization model reaches the required number of training rounds, and determine the joint optimization model obtained from the last training as the target joint optimization model.
[0140] 108. Channel and power selection is performed using a joint objective optimization model.
[0141] As can be seen, the deep reinforcement learning-based radio transmission method described in this embodiment of the invention can optimize and train the neural network parameters in the established joint optimization model of channel selection and power allocation. It selects the agent's actions through a greedy strategy, calculates the loss function through state transitions stored in the memory pool, and obtains the optimal network parameters through backpropagation and soft update methods. This improves the accuracy and reliability of the determined optimal network parameters. Furthermore, it optimizes and iterates the joint optimization model of channel selection and power allocation through the optimal network parameters, and then selects the channel and power through the joint optimization model. This is beneficial for making the optimal channel access and power allocation strategy for each user, which not only ensures the fairness of secondary user transmission, but also improves the overall stability and transmission rate of the transmission system.
[0142] In an optional embodiment, the parameter set further includes a threshold for the greedy strategy;
[0143] Determining the action of the first agent based on a greedy strategy can include the following operations:
[0144] The state of the first agent is input into the deep neural network to obtain the first return value of the deep neural network;
[0145] A first probability is randomly generated. When the first probability is less than or equal to the threshold of the greedy policy, the action of the first agent is randomly selected. When the first probability is greater than the threshold of the greedy policy, the action of the first agent is selected according to the first action selection formula.
[0146] Determining the state of the second agent corresponding to power allocation based on the actions and states of the first agent, and determining the actions of the second agent according to a greedy strategy, may include the following operations:
[0147] The state of the second agent corresponding to the power allocation is determined based on the actions and state of the first agent.
[0148] The state of the second agent is input into the deep neural network to obtain the second return value of the deep neural network;
[0149] A second probability is randomly generated. When the second probability is less than or equal to the threshold of the greedy policy, the action of the second agent is randomly selected. When the second probability is greater than the threshold of the greedy policy, the action of the second agent is selected according to the second action selection formula.
[0150] In this optional embodiment, the first return value of the deep neural network can be represented as The first return value of a deep neural network can be represented as The first probability can be represented as p, and the formula for choosing the first action can be expressed as... The second probability can be represented as q, and the formula for choosing the second action can be expressed as... .
[0151] As can be seen, implementing this optional embodiment allows the state of the first agent to be input into the deep neural network, obtaining a first return value from the deep neural network. Based on a randomly generated first probability, and according to a greedy strategy and a first action selection formula, the action of the first agent is selected. The state of the second agent corresponding to the power allocation is determined based on the action and state of the first agent. The state of the second agent is then input into the deep neural network, obtaining a second return value from the deep neural network. Based on a randomly generated second probability, and according to a greedy strategy and a second action selection formula, the action of the second agent is selected. This effectively improves the reliability and stability of the determined agent state. Determining the agent's action based on the first and second action selection formulas improves the accuracy and quantifiability of the selected agent action, further ensuring the accuracy and reliability of the joint optimization model's training.
[0152] Example 2
[0153] Please see Figure 2 , Figure 2 This is a flowchart illustrating another radio transmission method based on deep reinforcement learning disclosed in an embodiment of the present invention. Wherein, Figure 2 The described deep reinforcement learning-based radio transmission method can be applied to radio transmission systems, such as cognitive radio systems, and the embodiments of this invention are not limited thereto. Figure 2 As shown, this deep reinforcement learning-based radio transmission method may include the following operations:
[0154] 201. Establish a joint optimization model for channel selection and power allocation, and initialize the number of training rounds, memory pool, deep neural network, and parameter set of the deep neural network for the joint optimization model. The parameter set includes the initial network parameters of the deep neural network.
[0155] 202. For the current round of training the joint optimization model, initialize the state of the first agent corresponding to the channel selection.
[0156] 203. Determine the action of the first agent according to the greedy strategy, determine the state of the second agent corresponding to the power allocation according to the action and state of the first agent, and determine the action of the second agent according to the greedy strategy.
[0157] 204. Input the actions of the first agent and the actions of the second agent into the deep neural network for analysis, and obtain the reward content returned by the deep neural network.
[0158] 205. Update the agent's state, and generate a state transition based on the agent's state, actions, feedback content, and the updated state. Store the state transition in the memory pool.
[0159] 206. Randomly sample a preset number of data sets from the memory pool, calculate the loss function based on the data sets, update the initial network parameters of the deep neural network based on the loss function and the backpropagation algorithm, and obtain the current neural network parameters.
[0160] 207. Update the neural network parameters for the current iteration using the soft update method to obtain the optimal network parameters for the current iteration.
[0161] In this embodiment of the invention, optionally, the soft update formula corresponding to channel selection can be expressed as: The soft update formula corresponding to power allocation can be expressed as: ,in, and These represent the soft substitution coefficients corresponding to channel selection and power allocation, respectively.
[0162] 208. Determine the optimal network parameters for the current training as the initial network parameters for the deep neural network in the next training of the joint optimization model, continue to perform training operations on the joint optimization model until the number of training rounds of the joint optimization model is reached, and determine the joint optimization model obtained from the last training as the target joint optimization model.
[0163] 209. Channel and power selection is performed using a joint objective optimization model.
[0164] In this embodiment of the invention, for other descriptions of steps 201-206 and 209, please refer to the detailed description of steps 101-106 and 108 in Embodiment 1. These descriptions will not be repeated in this embodiment of the invention.
[0165] As can be seen, the deep reinforcement learning-based radio transmission method described in this embodiment of the invention can optimize and train the neural network parameters in the established joint optimization model of channel selection and power allocation. It selects the agent's actions through a greedy strategy, calculates the loss function based on the state transitions stored in the memory pool, and obtains the optimal network parameters through backpropagation and soft updates. This improves the accuracy and reliability of the determined optimal network parameters, which are then used as the initial network parameters for the deep neural network in the next training of the joint optimization model. The training operation continues on the joint optimization model, and the joint optimization model of channel selection and power allocation is iterated and optimized using the optimal network parameters. This ensures the accuracy and reliability of each optimization operation of the joint optimization model. Then, channel and power selection is performed through the joint optimization model, which helps to make optimal channel access and power allocation strategies for each user. This not only ensures the fairness of secondary user transmission but also improves the overall stability and transmission rate of the transmission system.
[0166] In an optional embodiment, the initialization formula for the state of the first agent corresponding to channel selection is:
[0167] ,
[0168] in, This represents the initial state of the first agent corresponding to the selection of the t-th time slot channel. This indicates the channel occupancy status of the t-th time slot. This represents the channel gain from the second user in the t-th time slot to the cognitive base station. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. Let t represent the channel gain from the primary user to the cognitive base station in the t-th time slot. When t=0, it represents the first initialization of the state of the first agent corresponding to channel selection. At the beginning of each time slot, the state of the first agent corresponding to channel selection is initialized, and the state of the first agent after initialization in each time slot is used for training the joint optimization model in this time.
[0169] In this optional embodiment, optionally, at the beginning of each round, the initialization state corresponding to the first agent can be represented as: ,in, This indicates the channel occupancy status of the current time slot. , and This indicates the channel gain of each link.
[0170] As can be seen, implementing this optional embodiment can express the state of the first agent corresponding to the initial channel selection using a formula, thereby improving the accuracy and reliability of subsequently determining the actions of the first agent and the second agent.
[0171] In another alternative embodiment, the formula for selecting the first action is:
[0172] ,
[0173] In the formula, This represents the action of the first agent in the t-th time slot. This represents the set of actions of the first agent in the t-th time slot. Indicates the first return value. This represents the state of the first agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the first agent;
[0174] The formula for selecting the second action is:
[0175] ,
[0176] In the formula, This represents the action of the second agent in the t-th time slot. This represents the set of actions of the second agent in the t-th time slot. This indicates the second return value. This represents the state of the second agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the second agent.
[0177] In this optional embodiment, optionally, at step t of each round, the first agent corresponding to channel selection first determines... Greedy strategy selects action and will and Simultaneously, it serves as the state input for the second agent corresponding to power allocation, and then according to... Greedy strategy selects action .
[0178] It is evident that implementing this optional embodiment can select the agent's actions according to a greedy strategy, thereby improving the accuracy and reliability of the determined agent's actions, and thus enhancing the accuracy of subsequently inputting the agent's actions into the cognitive wireless environment to generate rewards.
[0179] In yet another alternative embodiment, the formula for calculating the reward content is:
[0180] ,
[0181] in, This indicates the access status of the nth secondary user to channel m in the t-th time slot. This represents the channel dryness ratio of the nth user in the t-th time slot on channel m. This represents the linear fairness index for the t-th time slot. This represents the reachable rate of the nth user in the tth time slot. This represents the transmit power of the nth secondary user in the tth time slot. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. This is the interference threshold.
[0182] In this optional embodiment, for the channel model, large-scale path loss attenuation and small-scale sharp fading are considered. It is assumed that the channel follows quasi-static fading, and a block fading coincides with a single time slot. The channel gain of the t-th time slot can be expressed as:
[0183] ,
[0184] in, This represents the sharp fading channel gain in the t-th time slot. This represents the large-scale path loss component. This represents the distance between x and y. Indicates the reference distance. Let represent the path loss exponent; x and y represent the transmitter and receiver, respectively, and xy represent the channel link between x and y. Therefore, This represents the channel gain from the second user in the t-th time slot to the cognitive base station. This represents the channel gain from the secondary user to the primary user base station in the t-th time slot. This represents the channel gain from the primary user to the cognitive base station in the t-th time slot.
[0185] In this optional embodiment, This indicates the channel that SU-n selects to access in the t-th time slot. Let SU-n represent the transmit power selected by SU-n in the t-th time slot, where SU-n is the transmit power selected by SU-n in the t-th time slot. Let be the maximum transmit power of the SU. Therefore, the set of channel access strategies for all SUs in the t-th time slot is represented as: The set of power allocation strategies is represented as .
[0186] In this optional embodiment, the number of idle channels is assumed to be time-varying, and an indicator function is defined. To indicate the channel occupancy status of the current time slot:
[0187] ,
[0188] but Indicate the channel occupancy status of the t-th time slot; define an indicator function. Let this represent the access status of SU-n to channel m in the t-th time slot:
[0189] ,
[0190] In the t-th time slot, the channel dryness ratio (SINR) of SU-n on channel m is expressed as:
[0191] ,
[0192] The reachable rate of SU-n in the t-th time slot is:
[0193] ,
[0194] Introducing a linear fairness index The formula for calculating the fairness index of the system in the t-th time slot is:
[0195] ,
[0196] If the speed of one or more SUs is zero, then This indicates that the system is unfair, and This means that the speed of each unit in the system is exactly the same, therefore, The closer it is to 1, the better the system fairness.
[0197] In this optional embodiment, the system's sum rate calculation formula is:
[0198] ,
[0199] Define utility function Then the utility function in the t-th time slot is:
[0200] .
[0201] In this optional embodiment, such as Figure 7 As shown, Figure 7 This is a schematic diagram illustrating the impact of different strategies on the average utility function value r under a joint optimization model of channel selection and power allocation, as disclosed in an embodiment of the present invention. Figure 7 As shown, the average utility function value using the proposed joint optimization model is... The value increases with the number of training epochs and gradually converges around the 100th epoch. Compared to the myopic greedy strategy and the random selection strategy, the average utility function value of the joint optimization algorithm is [value missing]. Significant improvements have been made; such as Figure 8 As shown, Figure 8 This invention discloses an interference threshold at the power utilities (PU) under a joint optimization model for channel selection and power allocation. A schematic diagram illustrating the effect on the average utility function value r, as shown below. Figure 8 As shown, with Increase This allows for data transmission with higher transmission power, therefore The sum and rate will gradually increase, thus causing the average utility function value of the system to decrease. Increase.
[0202] As can be seen, implementing this optional embodiment can combine channel occupancy, channel access, channel dryness ratio, linear fairness index, and speed calculation formula to calculate the reward content, ensuring the accuracy and reliability of the determined reward content, while also guaranteeing the fairness of the joint optimization model and the channel transmission rate in the joint optimization model.
[0203] In yet another optional embodiment, updating the agent's state and generating a state transition based on the agent's state, the agent's actions, the report content, and the agent's updated state, and storing the state transition in a memory pool, may include the following operations:
[0204] Update the state of the first agent and the state of the second agent;
[0205] A first state transition is generated based on the state of the first agent, the action of the first agent, the report content, and the updated state of the first agent, and the first state transition is stored in the first memory pool corresponding to the channel selection.
[0206] The second state transition is generated based on the state of the second agent, the actions of the second agent, the report content, and the updated state of the second agent. When the time slot is greater than or equal to the preset time slot threshold, the second state transition is stored in the second memory pool corresponding to the power allocation.
[0207] In this optional embodiment, the state of the first agent is updated as follows: Update the state of the second agent to The first state transition corresponding to the first intelligent agent can be represented as: The first memory pool can be represented as The second state transition corresponding to the second agent can be represented as: The second memory pool can be represented as .
[0208] As can be seen, implementing this optional embodiment can store the updated state transitions corresponding to the first and second agents into their respective memory pools, thereby improving the accuracy and reliability of the determined state transitions and ensuring the reliability of the samples subsequently extracted from the memory pools.
[0209] In yet another alternative embodiment, randomly sampling a preset number of data sets from the memory pool may include the following operations:
[0210] From the first memory pool, randomly sample the first data set corresponding to the first agent; from the second memory pool, randomly sample the second data set corresponding to the second agent.
[0211] The loss function is calculated based on the dataset, including:
[0212] Calculate the first loss function corresponding to the first agent based on the first data set, and calculate the second loss function corresponding to the second agent based on the second data set.
[0213] In this optional embodiment, the sampling method can be random sampling, system sampling, or sampling based on preset conditions; this embodiment does not limit the method. The first data set can be represented as follows: The second data set can be represented as , and These represent the number of samples in the first and second datasets, respectively. This represents the first output value of the deep neural network corresponding to the j-th sample in the first dataset. This represents the second output value of the deep neural network corresponding to the j-th sample in the second dataset. and These represent the states of the first and second agents, respectively. Take action below The expected cumulative discount return obtained, j represents the j-th sample in the first sample set and the second sample set, This represents the state of the j-th sample. This represents the action of the j-th sample. This represents the current neural network parameters corresponding to the first agent. This represents the initial network parameters corresponding to the first agent. This represents the current neural network parameters corresponding to the second agent. This represents the initial network parameters corresponding to the second agent. This represents the learning rate corresponding to the first agent. This represents the learning rate corresponding to the second agent.
[0214] As can be seen, implementing this optional embodiment can randomly sample a first data set corresponding to the first agent from the first memory pool, randomly sample a second data set corresponding to the second agent from the second memory pool, calculate a first loss function corresponding to the first agent based on the first data set, and calculate a second loss function corresponding to the second agent based on the second data set. This improves the accuracy and reliability of the determined loss function, thereby ensuring the accuracy of subsequent optimization of network parameters.
[0215] In yet another optional embodiment, the parameter set also includes the learning rate of the deep neural network;
[0216] The formula for calculating the first loss function is:
[0217] ,
[0218] The second calculation formula for the function is:
[0219] ,
[0220] The formula for the backpropagation algorithm is:
[0221] , ,
[0222] in, Represents the first data set. This represents the number of samples in the first dataset. Indicates the second data set, This indicates the number of samples in the second dataset. This represents the first output value of the deep neural network corresponding to the j-th sample in the first dataset. This represents the second output value of the deep neural network corresponding to the j-th sample in the second dataset. and These represent the states of the first and second agents, respectively. Take action below The expected cumulative discount return obtained, j represents the j-th sample in the first sample set and the second sample set, This represents the state of the j-th sample. This represents the action of the j-th sample. This represents the current neural network parameters corresponding to the first agent. This represents the initial network parameters corresponding to the first agent. This represents the current neural network parameters corresponding to the second agent. This represents the initial network parameters corresponding to the second agent. This represents the learning rate corresponding to the first agent. This represents the learning rate corresponding to the second agent.
[0223] As can be seen, implementing this optional embodiment can combine the loss function formula to calculate the first loss function corresponding to the first agent based on the first data set, calculate the second loss function corresponding to the second agent based on the second data set, and update the network parameters through the backpropagation algorithm based on the loss function, thereby improving the accuracy and reliability of the determined loss function and network parameters, and thus ensuring the accuracy of subsequent optimization of the joint optimization model.
[0224] Example 3
[0225] Please see Figure 3 , Figure 3 This is a schematic diagram of a radio transmission device based on deep reinforcement learning disclosed in an embodiment of the present invention. Figure 3 The described deep reinforcement learning-based radio transmission device can be applied to radio transmission systems, such as cognitive radio systems, and the embodiments of this invention are not limited thereto. Figure 3 As shown, the deep reinforcement learning-based radio transmission device may include:
[0226] Initialization module 301 is used to establish a joint optimization model for channel selection and power allocation, and to initialize the number of training rounds, memory pool, deep neural network, and parameter set of the joint optimization model. The parameter set includes the initial network parameters of the deep neural network.
[0227] The initialization module 301 is also used to initialize the state of the first agent corresponding to the channel selection for the current round of training the joint optimization model.
[0228] The first determining module 302 is used to determine the action of the first intelligent agent according to a greedy strategy, determine the state of the second intelligent agent corresponding to the power allocation according to the action and state of the first intelligent agent, and determine the action of the second intelligent agent according to a greedy strategy.
[0229] The input module 303 is used to input the actions of the first agent and the actions of the second agent into the deep neural network for analysis, and to obtain the reward content returned by the deep neural network.
[0230] The update module 304 is used to update the state of the agent, and generate a state transition based on the state of the agent, the action of the agent, the content of the report, and the updated state of the agent, and store the state transition in the memory pool.
[0231] The acquisition module 305 is used to randomly sample a preset number of data sets from the memory pool, calculate the loss function based on the data sets, update the initial network parameters of the deep neural network based on the loss function and the backpropagation algorithm, and obtain the current neural network parameters.
[0232] The first determining module 302 is further configured to determine the current neural network parameters as the initial network parameters of the deep neural network when training the joint optimization model for the next time, and continue to perform training operations on the joint optimization model until the number of training rounds of the joint optimization model reaches the number of training rounds, and determine the joint optimization model obtained from the last training as the target joint optimization model;
[0233] Selection module 306 is used to select the channel and power through the objective joint optimization model.
[0234] It is evident that implementation Figure 3 The described deep reinforcement learning-based radio transmission device can optimize and train the neural network parameters in the established joint optimization model of channel selection and power allocation. It selects the agent's actions through a greedy strategy, calculates the loss function through state transitions stored in the memory pool, and obtains the optimal network parameters through backpropagation and soft update methods. This improves the accuracy and reliability of the determined optimal network parameters. Then, it optimizes and iterates the joint optimization model of channel selection and power allocation through the optimal network parameters, and finally selects the channel and power through the joint optimization model. This helps to make the optimal channel access and power allocation strategy for each user, which not only ensures the fairness of secondary user transmission, but also improves the overall stability and transmission rate of the transmission system.
[0235] In an optional embodiment, such as Figure 4 As shown, the deep reinforcement learning-based radio transmission device may further include:
[0236] The update module 304 is also used to update the current neural network parameters according to the soft update method to obtain the current optimal network parameters;
[0237] The second determining module 307 is used to determine the optimal network parameters of the current training as the initial network parameters of the deep neural network when training the joint optimization model for the next time, and to trigger the first determining module 302 to continue training the joint optimization model until the number of training rounds of the joint optimization model reaches the number of training rounds, and to determine the joint optimization model obtained from the last training as the target joint optimization model.
[0238] It is evident that implementation Figure 4The described deep reinforcement learning-based radio transmission device can optimize and train the neural network parameters in the established joint optimization model of channel selection and power allocation. It selects agent actions through a greedy strategy, calculates the loss function based on state transitions stored in the memory pool, and obtains the optimal network parameters through backpropagation and soft updates. This improves the accuracy and reliability of the determined optimal network parameters, which are then used as the initial network parameters for the next training of the joint optimization model. The training operation continues, and the joint optimization model is iteratively optimized using these optimal network parameters. This ensures the accuracy and reliability of each optimization operation. Channel and power selection is then performed through the joint optimization model, facilitating optimal channel access and power allocation strategies for each user. This not only ensures fairness in secondary user transmission but also improves the overall stability and transmission rate of the transmission system.
[0239] In another alternative embodiment, such as Figure 4 As shown, the initialization formula for the state of the first agent corresponding to channel selection is:
[0240] ;
[0241] in, This represents the initial state of the first agent corresponding to the selection of the t-th time slot channel. This indicates the channel occupancy status of the t-th time slot. This represents the channel gain from the second user in the t-th time slot to the cognitive base station. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. Let t represent the channel gain from the primary user to the cognitive base station in the t-th time slot. When t=0, it represents the first initialization of the state of the first agent corresponding to channel selection. At the beginning of each time slot, the state of the first agent corresponding to channel selection is initialized, and the state of the first agent after initialization in each time slot is used for training the joint optimization model in this time.
[0242] It is evident that implementation Figure 4 The described deep reinforcement learning-based radio transmission device can formulate the state of the first agent corresponding to the initial channel selection, thereby improving the accuracy and reliability of subsequently determining the actions of the first agent and the second agent.
[0243] In yet another alternative embodiment, such as Figure 4 As shown, the parameter set also includes the threshold for the greedy strategy;
[0244] The specific methods by which the first determining module 302 determines the action of the first intelligent agent according to the greedy strategy include:
[0245] The state of the first agent is input into the deep neural network to obtain the first return value of the deep neural network;
[0246] A first probability is randomly generated. When the first probability is less than or equal to the threshold of the greedy policy, the action of the first agent is randomly selected. When the first probability is greater than the threshold of the greedy policy, the action of the first agent is selected according to the first action selection formula.
[0247] The first determining module 302 determines the state of the second agent corresponding to the power allocation based on the actions and states of the first agent, and the specific method for determining the actions of the second agent according to the greedy strategy includes:
[0248] The state of the second agent corresponding to the power allocation is determined based on the actions and state of the first agent.
[0249] The state of the second agent is input into the deep neural network to obtain the second return value of the deep neural network;
[0250] A second probability is randomly generated. When the second probability is less than or equal to the threshold of the greedy policy, the action of the second agent is randomly selected. When the second probability is greater than the threshold of the greedy policy, the action of the second agent is selected according to the second action selection formula.
[0251] It is evident that implementation Figure 4 The described deep reinforcement learning-based radio transmission device can input the state of a first agent into a deep neural network, obtain a first return value from the deep neural network, and select the action of the first agent according to a randomly generated first probability, a greedy strategy, and a first action selection formula. It can also determine the state of a second agent corresponding to power allocation based on the action and state of the first agent, input the state of the second agent into the deep neural network, obtain a second return value from the deep neural network, and select the action of the second agent according to a randomly generated second probability, a greedy strategy, and a second action selection formula. This effectively improves the reliability and stability of the determined agent state. Determining the agent's action according to the first and second action selection formulas improves the accuracy and quantifiability of the selected agent action, further ensuring the accuracy and reliability of optimizing and training the joint optimization model.
[0252] In yet another alternative embodiment, such as Figure 4 As shown, the formula for selecting the first action is:
[0253] ,
[0254] In the formula, This represents the action of the first agent in the t-th time slot. This represents the set of actions of the first agent in the t-th time slot. Indicates the first return value. This represents the state of the first agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the first agent;
[0255] The formula for selecting the second action is:
[0256] ,
[0257] In the formula, This represents the action of the second agent in the t-th time slot. This represents the set of actions of the second agent in the t-th time slot. This indicates the second return value. This represents the state of the second agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the second agent.
[0258] It is evident that implementation Figure 4 The described deep reinforcement learning-based radio transmission device can select the agent's actions according to a greedy policy, which improves the accuracy and reliability of the determined agent's actions, thereby improving the accuracy of the subsequent input of the agent's actions into the cognitive wireless environment to generate rewards.
[0259] In yet another alternative embodiment, such as Figure 4 As shown, the formula for calculating the reward content is:
[0260] ,
[0261] in, This indicates the access status of the nth secondary user to channel m in the t-th time slot. This represents the channel dryness ratio of the nth user in the t-th time slot on channel m. This represents the linear fairness index for the t-th time slot. This represents the reachable rate of the nth user in the tth time slot. This represents the transmit power of the nth secondary user in the tth time slot. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. This is the interference threshold.
[0262] It is evident that implementation Figure 4The described deep reinforcement learning-based radio transmission device can combine channel occupancy, channel access, channel dryness ratio, linear fairness index, and speed calculation formula to calculate the reward content, ensuring the accuracy and reliability of the determined reward content, while also guaranteeing the fairness of the joint optimization model and the channel transmission rate in the joint optimization model.
[0263] In yet another alternative embodiment, such as Figure 4 As shown, the update module 304 updates the agent's state and generates a state transition based on the agent's state, actions, report content, and the updated state. The specific methods for storing the state transition in the memory pool include:
[0264] Update the state of the first agent and the state of the second agent;
[0265] A first state transition is generated based on the state of the first agent, the action of the first agent, the report content, and the updated state of the first agent, and the first state transition is stored in the first memory pool corresponding to the channel selection.
[0266] The second state transition is generated based on the state of the second agent, the actions of the second agent, the report content, and the updated state of the second agent. When the time slot is greater than or equal to the preset time slot threshold, the second state transition is stored in the second memory pool corresponding to the power allocation.
[0267] It is evident that implementation Figure 4 The described deep reinforcement learning-based radio transmission device can store the updated state transitions of the first and second agents into their respective memory pools, improving the accuracy and reliability of the determined state transitions, and thus ensuring the reliability of the samples subsequently extracted from the memory pools.
[0268] In yet another alternative embodiment, such as Figure 4 As shown, the specific method by which the sampling module 305 randomly samples a preset number of data sets from the memory pool includes:
[0269] From the first memory pool, randomly sample the first data set corresponding to the first agent; from the second memory pool, randomly sample the second data set corresponding to the second agent.
[0270] The specific methods by which the sampling module 305 calculates the loss function based on the dataset include:
[0271] Calculate the first loss function corresponding to the first agent based on the first data set, and calculate the second loss function corresponding to the second agent based on the second data set.
[0272] It is evident that implementation Figure 4The described deep reinforcement learning-based radio transmission device can randomly sample a first data set corresponding to a first agent from a first memory pool, randomly sample a second data set corresponding to a second agent from a second memory pool, calculate a first loss function corresponding to the first agent based on the first data set, and calculate a second loss function corresponding to the second agent based on the second data set. This improves the accuracy and reliability of the determined loss function, thereby ensuring the accuracy of subsequent optimization of network parameters.
[0273] In yet another alternative embodiment, such as Figure 4 As shown, the parameter set also includes the learning rate of the deep neural network;
[0274] The formula for calculating the first loss function is:
[0275] ,
[0276] The second calculation formula for the function is:
[0277] ,
[0278] The formula for the backpropagation algorithm is:
[0279] , ,
[0280] in, Represents the first data set. This represents the number of samples in the first dataset. Indicates the second data set, This indicates the number of samples in the second dataset. This represents the first output value of the deep neural network corresponding to the j-th sample in the first dataset. This represents the second output value of the deep neural network corresponding to the j-th sample in the second dataset. and These represent the states of the first and second agents, respectively. Take action below The expected cumulative discount return obtained, j represents the j-th sample in the first sample set and the second sample set, This represents the state of the j-th sample. This represents the action of the j-th sample. This represents the current neural network parameters corresponding to the first agent. This represents the initial network parameters corresponding to the first agent. This represents the current neural network parameters corresponding to the second agent. This represents the initial network parameters corresponding to the second agent. This represents the learning rate corresponding to the first agent. This represents the learning rate corresponding to the second agent.
[0281] It is evident that implementation Figure 4 The described deep reinforcement learning-based radio transmission device can combine the loss function formula to calculate the first loss function corresponding to the first agent based on the first data set, calculate the second loss function corresponding to the second agent based on the second data set, and update the network parameters through the backpropagation algorithm based on the loss function. This improves the accuracy and reliability of the determined loss function and network parameters, thereby ensuring the accuracy of subsequent optimization of the joint optimization model.
[0282] Example 4
[0283] Please see Figure 5 , Figure 5 This is a schematic diagram of another radio transmission device based on deep reinforcement learning disclosed in an embodiment of the present invention. Figure 5 As shown, the deep reinforcement learning-based radio transmission device may include:
[0284] Memory 401 storing executable program code;
[0285] Processor 402 coupled to memory 401;
[0286] The processor 402 calls the executable program code stored in the memory 401 to execute the steps in the radio transmission method based on deep reinforcement learning described in Embodiment 1 or Embodiment 2 of the present invention.
[0287] Example 5
[0288] This invention discloses a computer storage medium storing computer instructions, which, when invoked, are used to execute the steps in the deep reinforcement learning-based radio transmission method described in Embodiment 1 or Embodiment 2 of this invention.
[0289] Example 6
[0290] This invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform the steps in the deep reinforcement learning-based radio transmission method described in Embodiment 1 or Embodiment 2.
[0291] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0292] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0293] Finally, it should be noted that the radio transmission method and apparatus based on deep reinforcement learning disclosed in the embodiments of the present invention are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A radio transmission method based on deep reinforcement learning, characterized in that, The method includes: A joint optimization model for channel selection and power allocation is established, and the number of training rounds, memory pool, deep neural network, and parameter set of the deep neural network are initialized. The parameter set includes the initial network parameters of the deep neural network. For the current round of training of the joint optimization model, initialize the state of the first agent corresponding to the channel selection; The action of the first agent is determined according to a greedy strategy. The state of the second agent corresponding to the power allocation is determined according to the action and state of the first agent. The action of the second agent is then determined according to the greedy strategy. The actions of the first agent and the actions of the second agent are input into the deep neural network for analysis, and the reward content returned by the deep neural network is obtained. Update the state of the agent, and generate a state transition based on the state of the agent, the action of the agent, the report content, and the updated state of the agent, and store the state transition in the memory pool; A preset number of data sets are randomly sampled from the memory pool, and a loss function is calculated based on the data sets. The initial network parameters of the deep neural network are updated based on the loss function and the backpropagation algorithm to obtain the current neural network parameters. The current neural network parameters are determined as the initial network parameters of the deep neural network when training the joint optimization model for the next time, and the training operation on the joint optimization model continues until the number of training rounds of the joint optimization model reaches the number of training rounds, and the joint optimization model obtained from the last training is determined as the target joint optimization model; Channel and power selection are performed using the aforementioned joint optimization model.
2. The radio transmission method based on deep reinforcement learning according to claim 1, characterized in that, The method further includes: The neural network parameters are updated using a soft update method to obtain the optimal network parameters for the current iteration. The optimal network parameters for the current iteration are determined as the initial network parameters of the deep neural network for the next training of the joint optimization model. The training operation of the joint optimization model is continued until the number of training iterations of the joint optimization model reaches the number of training rounds. The joint optimization model obtained from the last training iteration is then determined as the target joint optimization model.
3. The radio transmission method based on deep reinforcement learning according to claim 1 or 2, characterized in that, The initialization formula for the state of the first agent corresponding to the channel selection is: ; in, This represents the initial state of the first agent corresponding to the channel selection in the t-th time slot. This indicates the channel occupancy status of the t-th time slot. This represents the channel gain from the second user in the t-th time slot to the cognitive base station. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. Let t represent the channel gain from the primary user to the cognitive base station in the t-th time slot. When t=0, it represents the first initialization of the state of the first agent corresponding to the channel selection. At the beginning of each time slot, the state of the first agent corresponding to the channel selection is initialized, and the state of the first agent after initialization in each time slot is used for training the joint optimization model in that time.
4. The radio transmission method based on deep reinforcement learning according to claim 1 or 2, characterized in that, The parameter set also includes the threshold for the greedy strategy; Determining the action of the first agent according to a greedy strategy includes: The state of the first agent is input into the deep neural network to obtain the first return value of the deep neural network; A first probability is randomly generated. When the first probability is less than or equal to the threshold of the greedy policy, the action of the first agent is randomly selected. When the first probability is greater than the threshold of the greedy policy, the action of the first agent is selected according to the first action selection formula. The step of determining the state of the second agent corresponding to the power allocation based on the action and state of the first agent, and determining the action of the second agent according to a greedy strategy, includes: The state of the second agent corresponding to the power allocation is determined based on the action and state of the first agent. The state of the second agent is input into the deep neural network to obtain the second return value of the deep neural network; A second probability is randomly generated. When the second probability is less than or equal to the threshold of the greedy policy, the action of the second agent is randomly selected. When the second probability is greater than the threshold of the greedy policy, the action of the second agent is selected according to the second action selection formula.
5. The radio transmission method based on deep reinforcement learning according to claim 4, characterized in that, The formula for selecting the first action is: , In the formula, This represents the action of the first agent in the t-th time slot. This represents the set of actions of the first agent in the t-th time slot. This represents the state of the first agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the first intelligent agent; The formula for selecting the second action is: , In the formula, This represents the action of the second agent in the t-th time slot. This represents the set of actions of the second agent in the t-th time slot. This represents the state of the second agent in the t-th time slot. This represents the initial network parameters of the deep neural network corresponding to the second agent.
6. The radio transmission method based on deep reinforcement learning according to claim 1 or 2, characterized in that, The formula for calculating the reward content is as follows: , in, This indicates the access status of the nth secondary user to channel m in the t-th time slot. This represents the channel dryness ratio of the nth user in the t-th time slot on channel m. This represents the linear fairness index for the t-th time slot. This represents the reachable rate of the nth user in the tth time slot. This represents the transmit power of the nth secondary user in the tth time slot. This represents the channel gain from the secondary user to the primary base station in the t-th time slot. This is the interference threshold.
7. The radio transmission method based on deep reinforcement learning according to claim 1, characterized in that, The process of updating the agent's state, generating a state transition based on the agent's state, the agent's action, the report content, and the updated state of the agent, and storing the state transition in the memory pool includes: Update the state of the first agent and the state of the second agent; A first state transition is generated based on the state of the first agent, the action of the first agent, the report content, and the updated state of the first agent, and the first state transition is stored in the first memory pool corresponding to the channel selection. A second state transition is generated based on the state of the second agent, the action of the second agent, the report content, and the updated state of the second agent. When the time slot is greater than or equal to a preset time slot threshold, the second state transition is stored in the second memory pool corresponding to the power allocation.
8. The radio transmission method based on deep reinforcement learning according to claim 7, characterized in that, The random sampling of a preset number of data sets from the memory pool includes: From the first memory pool, a first data set corresponding to the first agent is randomly sampled; from the second memory pool, a second data set corresponding to the second agent is randomly sampled. The step of calculating the loss function based on the dataset includes: Calculate the first loss function corresponding to the first agent based on the first data set, and calculate the second loss function corresponding to the second agent based on the second data set.
9. The radio transmission method based on deep reinforcement learning according to claim 8, characterized in that, The parameter set also includes the learning rate of the deep neural network; The formula for calculating the first loss function is: , The formula for calculating the second loss function is: , The formula for the backpropagation algorithm is: , , in, This refers to the first data set. This represents the number of samples in the first dataset. This refers to the second data set. This indicates the number of samples in the second dataset. This represents the first output value of the deep neural network corresponding to the j-th sample in the first data set. This represents the second output value output by the deep neural network corresponding to the j-th sample in the second data set. and These represent the states of the first and second intelligent agents, respectively. Take action below The expected cumulative discount benefit obtained, j represents the j-th sample in the first data set and the second data set, This represents the state of the j-th sample. This represents the action of the j-th sample. This represents the current neural network parameters corresponding to the first agent. This represents the initial network parameters corresponding to the first agent. This represents the current neural network parameters corresponding to the second agent. This represents the initial network parameters corresponding to the second agent. This represents the learning rate corresponding to the first intelligent agent. This represents the learning rate corresponding to the second intelligent agent.
10. A radio transmission device based on deep reinforcement learning, characterized in that, The device includes: An initialization module is used to establish a joint optimization model for channel selection and power allocation, and to initialize the number of training rounds, memory pool, deep neural network, and parameter set of the deep neural network for the joint optimization model. The parameter set includes the initial network parameters of the deep neural network. The initialization module is also used to initialize the state of the first agent corresponding to the channel selection for the current round of training the joint optimization model. The first determining module is configured to determine the action of the first agent according to a greedy strategy, determine the state of the second agent corresponding to the power allocation according to the action and state of the first agent, and determine the action of the second agent according to the greedy strategy. An input module is used to input the actions of the first agent and the actions of the second agent into the deep neural network for analysis, and to obtain the reward content returned by the deep neural network; An update module is used to update the state of the agent, generate a state transition based on the state of the agent, the action of the agent, the report content, and the updated state of the agent, and store the state transition in the memory pool; The sampling module is used to randomly sample a preset number of data sets from the memory pool, calculate a loss function based on the data sets, update the initial network parameters of the deep neural network based on the loss function and the backpropagation algorithm, and obtain the current neural network parameters. The first determining module is further configured to determine the current neural network parameters as the initial network parameters of the deep neural network when training the joint optimization model for the next time, and continue to perform training operations on the joint optimization model until the number of training rounds of the joint optimization model reaches the number of training rounds, and determine the joint optimization model obtained from the last training as the target joint optimization model; The selection module is used to select the channel and power through the target joint optimization model.
Citation Information
Patent Citations
Green cognitive radio power distribution method based on deep reinforcement learning
CN114126021A