Method for optimizing indoor aggregation VLC-RF energy efficiency based on post-decision depth Q network
By using a post-decision deep Q-network approach, the access point selection and resource allocation of indoor aggregated VLC-RF networks are optimized, solving the computational complexity and local optima problems of traditional algorithms, and achieving improvements in network energy efficiency and user communication quality.
Patent Information
- Application Number
- CN202511363465.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-01-06
AI Technical Summary
Existing technologies struggle to effectively optimize access point selection and resource allocation in indoor aggregated VLC-RF networks, resulting in low network energy efficiency. Furthermore, traditional optimization algorithms suffer from high computational complexity, are prone to getting trapped in local optima, and have weak environmental adaptability.
A post-decision deep Q-network-based approach is adopted. By training a deep Q-network model, access point selection, sub-channel allocation, and power allocation are optimized. An energy efficiency optimization model is designed, which combines the Bregman ball algorithm and the ε-greedy strategy to generate an agent action scheme that combines imitation learning and local actions. The network parameters are updated through an adaptive moment estimation optimization algorithm to obtain the optimal action strategy.
It improves the energy efficiency of indoor aggregated VLC-RF networks, enhances network learning efficiency and stability, ensures user communication satisfaction, and achieves efficient utilization of network resources.
Smart Images

Figure CN121284599A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of optical wireless communication technology and relates to a method for optimizing indoor aggregated VLC-RF energy efficiency based on post-decision deep Q network. Background Technology
[0002] Visible Light Communication (VLC) technology effectively solves the crisis of scarce spectrum resources in Radio Frequency (RF) communication, and has advantages such as high speed, energy saving, and no electromagnetic interference, making it a core technology for next-generation wireless communication networks. Integrating VLC and RF can increase network capacity, extend communication range, and improve user communication experience. In indoor VLC-RF converged networks, users can simultaneously receive information transmitted by both VLC and RF links, achieving higher data transmission rates and a more stable and reliable communication experience than receiving VLC or RF links separately. Furthermore, selecting appropriate network access points (APs) and reasonable channel resources and transmit power for users can not only improve user communication satisfaction but also reduce network resource waste and improve communication quality.
[0003] With the increasing number and variety of wireless communication devices, energy consumption has gradually become a focus of attention in the global communications industry. Optimizing power allocation to improve network energy efficiency, while simultaneously optimizing user access and channel allocation, has become an inevitable trend in the development of wireless communication networks. Deep reinforcement learning, as a machine learning method that combines reinforcement learning with deep neural networks, possesses powerful learning capabilities and can effectively address the limitations of traditional optimization algorithms, such as high computational complexity, susceptibility to local optima, and weak environmental adaptability. Furthermore, deep Q-networks, as one of the key models in the field of deep reinforcement learning, have the characteristics of experience playback mechanisms and the provision of high-quality experience samples for long-term problems by the target Q-network, which can improve the stability and learning efficiency of the deep Q-network model training process. Therefore, studying the access point selection and resource allocation problems in indoor aggregated VLC-RF networks based on deep Q-network models, and thus improving network energy efficiency, is of significant research importance. Summary of the Invention
[0004] The core of this invention is to provide a method for optimizing the energy efficiency of indoor aggregated VLC-RF based on a post-decision deep Q-network. By training a post-decision deep Q-network model, the optimal action strategies for access point selection, sub-channel allocation, and power allocation are obtained, thereby achieving the best energy efficiency results for the indoor aggregated VLC-RF network.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] 1. A method for optimizing indoor aggregated VLC-RF energy efficiency based on post-decision deep Q-network, characterized in that the method includes the following steps:
[0007] S1: Input user set in the indoor aggregated VLC-RF network VLC Access Point (AP) Collection A set of subchannels for an RFAP and a VLC AP and the sub-channel set of the RF AP In this context, the VLC AP reuses the same set of sub-channels; the network parameters θ of the input decision depth Q-network are... t Based on empirical samples, the network parameters of the target Q-network and empirical samples, and Set the post-decision depth Q network experience replay buffer capacity to X1, and set X3 = 0; calculate the channel gain of each sub-channel of the user accessing the VLC AP according to the Lambert radiation model; calculate the channel gain of the user accessing the RF AP sub-channel according to the Rayleigh fading model.
[0008] According to the Lambert radiation model, the channel gain between user k and subchannel n of VLC AP v is calculated as follows:
[0009]
[0010] In the above formula, This represents the probability that the line-of-sight link between user k and VLC AP v on subchannel n is not blocked. A line-of-sight link refers to a straight-line optical propagation path between the VLC AP and the user without any obstructions. PD This represents the physical area of the photodetector; t = -1 / log2(cos(Φ) 1 / 2 )) represents the order of the Lambert radiation, Φ 1 / 2 d is the half-power angle of the LED; v,k φ represents the distance between VLC AP v and user k; v,k It is the irradiance angle from VLC AP v to user k; ψ v,k It is the angle of incidence from VLC AP v to user k; T o The gain of the optical filter is typically set to T. o =1; G(ψ) v,k )=δ 2 / sin 2 Ψ FoV , where represents the concentrator gain, δ is the refractive index, and Ψ is the refractive index. FoV The field of view of the photodetector;
[0011] According to the Rayleigh fading model, the channel gain between user k and the subchannel m of the RF AP is calculated using the following formula:
[0012]
[0013] In the above formula, h r This indicates small-scale fading due to signal multipath propagation, following a Rayleigh distribution, with an average power of 2.46 dB; d RF,k L(d) represents the distance between user k and the RF AP; RF,k ) represents the large-scale fading loss, calculated using the formula L(d RF,k ) = 20log 10 (d RF,k )+20log 10 (f c -147.5 (dB), f c This indicates the center frequency of the RF carrier, expressed in GHz.
[0014] S2: Calculate the received signal-to-interference-plus-noise ratio (SINR) of the user on the VLC AP sub-channel, and calculate the throughput between the user and the VLC AP sub-channel using the VLC throughput lower bound closure formula; calculate the received SINR of the user on the RF AP sub-channel, and calculate the throughput between the user and the RF AP sub-channel using Shannon's law; in the indoor aggregated VLC-RF network, any user can access the RF AP, and users in the VLC AP coverage area can access one VLC AP and one RF AP simultaneously. Calculate the throughput achieved by each user accessing the indoor aggregated VLC-RF network, the total throughput of the indoor aggregated VLC-RF network, the transmit power allocated to the user by the indoor aggregated VLC-RF network, and the total transmit power consumed by the indoor aggregated VLC-RF network.
[0015] The SINR calculation formula between user k and subchannel n of VLC AP v is as follows:
[0016]
[0017] In the above formula, R is the transmit power that VLC AP v allocates to user k on subchannel n; PD v′ and k′ represent the responsivity of the photodetector; v′ and k′ represent other VLC APs and other users that reuse the same subchannel n. It is the transmit power that VLC APv′ allocates to user k′ when multiplexing the same subchannel n; N represents the interference channel gain for user k when VLC APv′ reuses the same subchannel n; VLC B is the power spectral density of additive white Gaussian noise in a VLC channel; VLC It is the sub-channel bandwidth of VLC AP v;
[0018] The lower bound closure formula for the throughput provided by VLC AP v to user k on subchannel n is:
[0019]
[0020] The SINR calculation formula between user k and subchannel m of the RF AP is as follows:
[0021]
[0022] In the above formula, N is the transmit power that the RF AP allocates to user k on subchannel m. RF B is the power spectral density of additive white Gaussian noise in an RF channel. RF It is the sub-channel bandwidth of the RF AP;
[0023] According to Shannon's law, the formula for calculating the throughput provided by the RF AP to user k on subchannel m is:
[0024]
[0025] The formula for calculating the throughput achieved by user k accessing the indoor aggregated VLC-RF network is as follows:
[0026]
[0027] The formula for calculating the total throughput of the indoor aggregated VLC-RF network is as follows:
[0028]
[0029] The formula for calculating the transmit power allocated to user k in the indoor aggregated VLC-RF network is as follows:
[0030]
[0031] The formula for calculating the total transmit power consumed by the indoor aggregated VLC-RF network is as follows:
[0032]
[0033] In the above formula, P rf and P vlcThese represent the power consumption of the internal circuitry of the RF AP and VLC AP, respectively; V is the total number of VLC APs.
[0034] S3: Establish an energy efficiency optimization problem model for joint access point selection and resource allocation in an indoor aggregated VLC-RF network, expressed as follows: η EE Energy efficiency for indoor aggregated VLC-RF networks; The access point selection vector, whose element x RF,k =1 and x v,k =1 respectively indicates that user k accesses the RF AP and user k accesses the VLC AP v; Represents the sub-channel allocation vector, whose elements and These represent the RF AP allocating subchannel m to user k and the VLCAP v allocating subchannel n to user k, respectively. Represents the power allocation vector, whose elements and These represent the transmit power allocated by the RF AP to user k on subchannel m and the transmit power allocated by VLCAP v to user k on subchannel n, respectively. This represents the communication interruption vector. If the user is in a communication interruption state, then a k =0, otherwise, a k =1;
[0035] The energy efficiency optimization model for joint access point selection and resource allocation in the established indoor aggregated VLC-RF network is as follows:
[0036]
[0037] In the above formula, η EE The energy efficiency of an indoor aggregated VLC-RF network is defined as the ratio of total network throughput to total network transmit power; x = {x RF,k ,x v,k} represents the access point selection vector, whose element x RF,k =1 and x v,k =1 respectively indicates that user k is connected to the RF AP and user k is connected to the VLC AP v; Represents the sub-channel allocation vector, whose elements and These represent the RF AP allocating subchannel m to user k and the VLC AP v allocating subchannel n to user k, respectively. Represents the power allocation vector, whose elements and These represent the transmit power allocated to user k by the RF AP on subchannel m and the transmit power allocated to user k by the VLC AP v on subchannel n, respectively; a = {a k} represents the communication interruption vector. If the user is in a communication interruption state, then a k =0, otherwise, a k =1; Constraint C1 ensures that the AP can only allocate a subchannel to a user when the user accesses the AP; Constraint C2 ensures that the transmit power allocated by the RF AP and each VLC AP to any user on any subchannel should not exceed the AP's maximum transmit power, and transmit power can only be allocated when the AP allocates a subchannel to a user. RF and P v Let C1 and C2 represent the maximum transmit power of the RF AP and VLCAP v, respectively; constraint C3 ensures that the power consumption of each AP does not exceed its maximum transmit power; constraint C4 ensures that any subchannel of each AP is allocated to at most one user; constraint C5 ensures that each user can access the RF AP, but can access at most one VLCAP; constraint C6 requires that the throughput of any user when communication is uninterrupted is greater than or equal to the minimum data rate requirement R. min To ensure the non-negativity of each element in the power allocation vector p, the binary nature of each element in the sub-channel allocation vector s, the binary nature of each element in the access point selection vector x, and the binary nature of each element in the interruption vector a, constraints C7, C8, C9, and C10 are applied respectively.
[0038] S4: This invention designs a method for optimizing the energy efficiency of indoor aggregated VLC-RF based on a post-decision deep Q-network. A post-decision state mechanism is introduced into the deep Q-network, and an energy efficiency optimization model based on the post-decision deep Q-network is designed. The post-decision state is an intermediate state between the agent's current state and the next state, representing the direct result of the action and constituting the deterministic part of the agent's state transition process. In the energy efficiency optimization model based on the post-decision deep Q-network, the agent's learning task can be described by a Markov decision process.
[0039] The Markov decision process in S4 is as follows:
[0040] S401: The central controller of the indoor aggregated VLC-RF network is used as the agent of the post-decision deep Q network. The agent is responsible for learning and making decisions on the access VLC AP or RF AP, sub-channel and power allocation strategies of all users in the VLC-RF network to optimize the energy efficiency of the indoor aggregated VLC-RF network. The network parameters of the post-decision deep Q network are set as the network parameters of the target Q network, and the experience samples of the target Q network are copied to the experience buffer of the post-decision deep Q network.
[0041] S402: Constructing the state space of the intelligent agent The state vector of the agent in time slot t is represented as: in, Let user k be located in time slot t; Let k be the speed at which user k moves in time slot t. This indicates whether the RF AP has allocated subchannel m in time slot t; This indicates whether VLCAP v allocates subchannel n in time slot t; as well as These represent the received signal-to-interference-plus-noise ratio (SINR) of user k from subchannel m of RFAP in time slot t and the received SINR of user k from subchannel n of VLC AP v in time slot t, respectively. They are used to measure the subchannel quality of RF network and VLC network in time slot t. as well as These are used to measure the total transmit power of the RF AP and each VLC AP in time slot t, respectively; R k,t The throughput of user k in time slot t is used to determine whether the minimum data rate requirement R is met. min ;a k,t ={0,1}, indicating whether user k's communication in time slot t is interrupted;
[0042] S403: Constructing the action space of the intelligent agent The agent operates in time slot t based on state s t Execute actions according to the action strategy Includes access point selection x for each user and each access point. t Sub-channels c allocated to each user by each access point t and power distribution p t Three decision vectors;
[0043] S404: Constructing the post-decision state space of the agent The agent takes action Then, the agent moves from state s t Transition to a deterministic post-decision state The constituent elements and meanings of s t same;
[0044] S405: The agent operates based on state s in time slot t. t Take action Then, a known state transition probability is used. From state s t Transition to post-decision state This is the deterministic phase; subsequently, the agent transitions to an unknown state with probabilities. From the post-decision state Transition to the next state s t+1 This is the random phase; among which, The value is estimated offline and Monte Carlo sampled from historical experience samples in the experience replay buffer of the deep Q network. The agent then searches for a value related to its state s from the experience replay buffer using the Bregman sphere algorithm. t The transition probability is determined by referencing the historical experience sample with the highest similarity. The value; and Action values are directly generated by the agent based on an ε-greedy policy. In the ε-greedy policy, the agent randomly selects a non-greedy action with a probability of ε- based on the current state, and selects the optimal action with the largest Q value from the current policy set of employees with a probability of 1-ε.
[0045] Among them, the agent of the ε-greedy strategy is based on state s t Take local action probability for:
[0046]
[0047] In the above formula, For empirical samples Q-value function;
[0048] S5: The agent operates based on state s in time slot t. t The action strategy determines the execution action. The formula for determining vector values is: in, To mimic the learned action values, the agent uses the Bregman ball algorithm to search for values in the deep Q-network's experience replay buffer that correspond to its state s. t The most similar historical experience sample is used to generate imitation learning actions by referencing the action strategies of the historical experience sample. value; These are local action values, generated directly by the agent based on an ε-greedy policy. The value represents the immediate feedback value of the agent's interaction with the environment; These are the weighting coefficients.
[0049] The specific process of S5 is as follows:
[0050] S501: The Bregman sphere algorithm is used to calculate the difference between two data points based on the Bregman divergence, and the agent's state vector s is then used. t Treat it as a data point in a high-dimensional space, and take it as the center point Z of the Bregman sphere.cen Construct a Bregman sphere; represent the set of positions in the Bregman sphere corresponding to the state vector data points of all historical experience samples in the deep Q-network experience replay buffer as follows:
[0051] From the formula Calculate Z cen The Bregman divergence between f and z, where f is a convex function in the vector space. Let f represent the gradient of function f at z, (Z) cen -z) represents the difference between two vectors. Indicates the calculation of gradient With vector (Z) cen The inner product of -z);
[0052] S503: Judgment Whether a condition is true or false is true if and only if the data point z is inside the Bregman sphere, representing the relationship between the historical experience sample corresponding to data point z and the agent's state vector s. t They possess similarity; for data points within the Bregman sphere, the data point with the smallest Bregman divergence is selected, representing the data point with the agent's state vector s. t The agent generates a learned action by referencing the action strategy of the most similar historical experience sample. Among them, the value of the Bregman sphere radius needs to ensure that the Bregman sphere contains a sufficient number of historical experience samples;
[0053] S504: Determine the agent based on state s t Local actions taken The agent is based on state s t Adopting a non-greedy exploration strategy with a probability of ε helps the agent discover actions that may bring higher rewards in an unknown environment, preventing the agent from prematurely converging to a local optimum; this is exploration. The agent will select the optimal action that maximizes the Q-value function under the current strategy with a probability of 1-ε; this is exploitation. A local action is then generated. Furthermore, to reduce the exploration probability of the agent during the learning process, ε generally decreases gradually with the increase of time steps, and the calculation formula is ε = max{ε - ε δ ,ε min}, where ε δ It is the decrease in ε after each time step, ε min It is the minimum value that ε can be reduced to.
[0054] Among them, the agent of the ε-greedy strategy is based on state s tTake local action probability for:
[0055]
[0056] In the above formula, For empirical samples Q-value function;
[0057] in, The calculation formula is:
[0058]
[0059] In the above formula, V * (s t+1 ) indicates that the agent is based on state s t Transition to the next state s t+1 The optimal Q-value function of the system at that point, i.e.
[0060] S505: From the formula Determine the action learned by imitation and local actions Combined agent actions value;
[0061] S6: This agent is based on state s t Take action Receive immediate reward value afterward The immediate reward value is determined by the known reward for the known state transition probability and the unknown reward for the unknown state transition probability; and each As empirical samples, they are stored in the empirical replay buffer of the post-decision deep Q-network;
[0062] Among them, timely reward value The calculation formula is:
[0063]
[0064] In the above formula, Given an agent with a known state transition probability From state s t Transition to post-decision state s t The known reward value was then obtained; The probability that the agent will subsequently transition to an unknown state. From the post-decision state s t Transition to the next state s t+1 An unknown reward value was subsequently obtained;
[0065] in, The formula for calculating the reward value is:
[0066]
[0067] In the above formula, η EE Represents the agent from state s t Transferred to Energy efficiency of indoor aggregated VLC-RF networks is a revenue item; Indicates that the agent changes state s t Transferred to The probability of communication interruption for user k. Indicates that the agent changes state s t Transferred to The probability that user k's minimum data rate requirement is not met is a cost term; ζ1 and ζ2 both represent weighting coefficients used to balance the benefits and costs in the reward function, and their values are both between 0 and 1.
[0068] in, The formula for calculating the reward value is:
[0069]
[0070] In the above formula, η EE Representing the agent from the state Transfer to s t+1 Energy efficiency of indoor aggregated VLC-RF networks is a revenue item; Indicates that the agent changes state Transfer to s t+1 The probability of communication interruption for user k. Indicates that the agent changes state s t Transfer to s t+1 The probability that user k's minimum data rate requirement is not met is a cost item.
[0071] S7: Calculate the average sample loss function value between the Q value of the post-decision deep Q network and the Q value of the target Q network, minimize the average sample loss function value using the adaptive moment estimation optimization algorithm, and output the action policy with the minimum average sample loss function value as the optimal action policy of the post-decision deep Q network; when the number of samples in the experience buffer reaches X1, update the neural network parameters of the post-decision deep Q network, and update the experience samples of the target Q network and the deep Q network.
[0072] The specific process of S7 is as follows:
[0073] S701: State-action pairs of each empirical sample in the empirical buffer of the post-calculated decision-making deep Q-network. and Q-value function;
[0074] Among them, the state-action pairs of the experience samples and The formula for calculating the Q-value function is:
[0075]
[0076] In the above formula, State-action pairs for experience samples Q-value function, State-action pairs for experience samples Q-value function, V * (s t+1 ) indicates that the agent transitions from the state of the current experience sample to the next state s. t+1 The optimal Q-value function of the system at that point, i.e.
[0077] S702: The next state s of each empirical sample in the empirical buffer of the decision deep Q network after updating using the Bellman equation. t+1 Q value at and the next state s t+1 The corresponding Q value
[0078] Among them, the next state s of the empirical sample t+1 place The formula for calculating the value is:
[0079]
[0080] In the above formula, α∈(0,1) represents the learning rate; γ∈[0,1] represents the reward discount factor, which is used to measure the importance of future rewards. The closer γ is to 1, the more the agent values future rewards.
[0081] Among them, the next state s of the empirical sample t+1 The corresponding Q value The calculation formula is:
[0082]
[0083] S703: Calculate the sample average loss function value between the Q value of the post-decision deep Q network and the Q value of the target Q network. In order to train the post-decision deep Q network model, the sample average loss function measures the difference between the estimated Q value of the deep Q network and the Q value of the target Q network. The smaller the value of the sample average loss function, the closer the predicted result of the model output is to the real result, and the closer the action strategy learned by the agent is to the optimal strategy.
[0084] The formula for calculating the average loss function value between the best Q-value of the empirical samples in the empirical replay buffer of the post-decision deep Q-network and the best Q-value of the empirical samples in the target Q-network is as follows:
[0085]
[0086] In the above formula, θ t The parameters of the post-decision depth Q-network in the current time slot t; The network parameters of the target Q-network; The network parameters θ of X1 experience samples in the agent's experience buffer. t Choose the optimal action in the next state. The average maximum expected return after that, i.e. in, The network parameters θ represent the X1 experience samples in the agent's experience buffer. t The average expected return value of choosing the optimal action in the next state; This represents the average Q-value of X2 samples in the target Q-network. Indicates rounding up to the nearest integer;
[0087] S704: The state-action pair corresponding to the post-decision depth Q network with the minimum sample average loss function value is taken as the optimal action policy output of the indoor aggregated VLC-RF network.
[0088] The optimal action strategy for determining the state-action pairs in the indoor aggregated VLC-RF network is as follows:
[0089]
[0090] In the above formula, The state and action corresponding to the post-decision depth Q network that minimizes the average loss function value of the samples are represented by , and the optimal action set of user access state, channel selection state and power allocation state that minimizes energy efficiency in indoor aggregated VLC-RF network is represented by .
[0091] S705: When the number of samples in the experience buffer reaches X1, in order to minimize the sample average loss function between the Q-value of the deep Q-network and the Q-value of the target Q-network, this invention uses the adaptive moment estimation optimization algorithm to update the network parameters of the decision deep Q-network to θ. t+1 That is, let θ t =θ t+1The idea is to minimize the sample average loss function, thereby gradually reducing the difference between the expected return value and the actual result of the post-decision deep Q network, so that the post-decision deep Q network model can gradually approach the optimal Q value function, and finally obtain the optimal action strategy for access point selection, sub-channel allocation and power allocation, so as to achieve the best energy efficiency optimization result for the indoor aggregated VLC-RF network; let X3 = X3 + 1;
[0092] Among them, the network parameter θ is updated using the adaptive moment estimation optimization algorithm. t+1 The calculation formula is:
[0093]
[0094] In the above formula, α t The learning rate in the adaptive moment estimation optimization algorithm; This represents the first-order moment estimate after bias correction, where, This is a first-moment estimate of the gradient of the sample average loss function, i.e., the mean of the gradient. It is the exponential decay rate estimated by the first moment. It is the sample average loss function L(θ) t Regarding the post-decision depth Q-network parameter θ t The gradient value; This represents the second-order moment estimate after bias correction, where, The second moment estimate of the gradient of the sample average loss function is represented by the mean of the squared gradient. It is the exponential decay rate estimated by the second moment. It is the sample average loss function with respect to the network parameters θ t The gradient squared value; κ represents a very small constant to avoid the denominator being zero; this invention uses α t , The initial values of κ are set to 0.001, 0.9, 0.999, and 10, respectively. -8 ;ω t and χ t Initially set to 0, introducing a deviation correction can increase ω in the initial stage. t and χ t The value of is adjusted to mitigate the impact of bias; as the time step increases, and The bias correction factor tends to 0. and As the value approaches 1, the impact of the bias correction will gradually decrease.
[0095] S706: Copy the two largest Q-value empirical samples from the empirical replay buffer of the post-decision deep Q-network to the target Q-network, replacing the empirical samples in the target Q-network, and update the network parameters of the target Q-network as follows: Set X3 = 0; if a new connection request is made to access the indoor aggregated VLC-RF network, proceed to step S1.
[0096] The beneficial effects of this invention are as follows: This invention provides a method for optimizing the energy efficiency of indoor aggregated VLC-RF networks based on post-decision deep Q-networks, to solve the energy efficiency optimization problem of joint access point selection and resource allocation in indoor aggregated VLC-RF networks. First, an energy efficiency optimization model based on post-decision deep Q-networks is designed. Second, based on the Bregman ball algorithm, historical experience samples with the highest similarity to the agent's current state are found. By referring to the action strategies of historical experience samples, an imitation learning action is generated, and a local action of the agent after interaction with the environment is generated based on an ε-greedy strategy, thus designing an agent action scheme that combines the imitation learning action with the local action. Furthermore, a reward function is designed that aims to maximize network energy efficiency while also considering user communication satisfaction. Finally, by training the post-decision deep Q-network model, the optimal action strategies for access point selection, sub-channel allocation, and power allocation are obtained, thereby achieving the best optimization result for the energy efficiency of the indoor aggregated VLC-RF network. Attached Figure Description
[0097] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein:
[0098] Figure 1 This is an indoor aggregated VLC-RF network model;
[0099] Figure 2 This is a schematic diagram of the post-decision state transition process;
[0100] Figure 3 This is an energy efficiency optimization model based on a post-decision deep Q-network;
[0101] Figure 4 The flowchart shows the indoor aggregated VLC-RF energy efficiency optimization method based on post-decision deep Q-network. Detailed Implementation
[0102] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0103] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures, and should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0104] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0105] Appendix Figure 1 This is an indoor aggregated VLC-RF network model. An indoor aggregated VLC-RF network consists of a central controller, one RF AP, multiple VLC APs, and multiple users. All APs are connected to the central controller. The downlink of the VLC APs is used for data transmission. All users indoors can access the RF APs. Users within the VLC AP coverage area can simultaneously access one VLC AP and one RF AP to achieve higher data transmission rates. Users in overlapping VLC AP coverage areas will be affected by inter-cell interference. The VLC APs are aggregated for... Indicated by |·|, where |·| represents the cardinality of the set; RF APs and all VLC APs employ Orthogonal Frequency Division Multiple Access (OFDMA), and the sub-channel set of an RF AP is represented by... This indicates that the sub-channel set of each VLC AP is used... Representation; User set This indicates that a random pathpoint movement model is used to simulate user movement behavior indoors; any user's receiving device includes a photodetector and an RF antenna, used to receive VLC and RF signals respectively. Φ 1 / 2 The half-power angle of the light-emitting diode (LED) in a VLC AP; φ v,k It is the irradiance angle from VLC AP v to user k; ψ v,k It is the angle of incidence from VLC AP v to user k; Ψ FoV This refers to the receiver's field of view (FoV).
[0106] Appendix Figure 2 This is a schematic diagram of the post-decision state transition process. The agent, based on state s, operates in time slot t. t Take action Then, with a known state transition probability. From state s t Transition to post-decision state s t Obtain known rewards This is the deterministic part of the state transition process, representing the direct result after the action is executed; subsequently, the agent transitions to a state with an unknown probability. From the post-decision state s t Transition to the next state s t+1 Receive unknown rewards This is the stochastic part of the state transition process, taking into account the real-time dynamic changes in the environment and the influence of other uncertainties; the final immediate reward obtained... From the formula The calculation involves obtaining known state transition probabilities by offline estimation or Monte Carlo sampling of historical experience samples in the experience replay buffer of the deep Q-network. Introducing a post-decision state mechanism into the deep Q-network can separate the deterministic and stochastic parts of the agent's state transition process, allowing subsequent Q-value function updates to focus on the stochastic part of the state transition process, thereby improving the agent's learning efficiency and accelerating the convergence of the learning model.
[0107] Appendix Figure 3 For the energy efficiency optimization model based on the post-decision deep Q-network, the experience samples obtained by the agent interacting with the VLC-RF environment are stored in the experience replay buffer of the post-decision deep Q-network. A batch of experience samples is randomly selected from this buffer for model training. The experience replay mechanism helps to break the correlation between data, improve the stability of the model training process, and enhance learning efficiency. The post-decision deep Q-network consists of a Q-network and a target Q-network. The Q-network outputs the estimated Q-value after taking an action in the current state, representing the true result of the Q-value of the post-decision deep Q-network. The target Q-network is used to stably estimate the target Q-value for the next state, representing the prediction result. The sample average loss function is used to intuitively reflect the error between the estimated Q-value and the target Q-value. This invention uses an adaptive moment estimation optimization algorithm to continuously update the network parameters θ of the post-decision deep Q-network. tThis minimizes the average loss function of the samples, thereby obtaining the optimal action strategy for access point selection, sub-channel allocation, and power allocation. After a certain number of training rounds or when the number of samples in the experience replay buffer reaches X1, the Q network periodically copies the network parameters to the target Q network, replacing the samples in the target Q network with high-quality experience samples, thus updating the parameters and samples of the target Q network. The agent's action selection consists of two parts: on the one hand, based on the Bregman ball algorithm, it searches for the historical experience sample with the highest similarity to the agent's current state in the experience replay buffer of the deep Q network, and generates an imitation learning action by referring to the action strategy of the historical experience sample; on the other hand, it directly generates a local action based on the ε-greedy strategy, representing the immediate feedback of the agent's interaction with the environment, and combines the imitation learning action and the local action as the agent's action plan. The state transition process of the agent follows the post-decision state mechanism.
[0108] Appendix Figure 4 The flowchart below shows the indoor aggregated VLC-RF energy efficiency optimization method based on post-decision deep Q-networks. A detailed explanation follows:
[0109] Input: Input the indoor aggregated VLC-RF network environment and resource information, including user sets. VLC Access Point (AP) Collection A set of subchannels for an RF AP and a VLC AP and the sub-channel set of the RF AP In this context, the VLC AP reuses the same set of sub-channels; the network parameters θ of the input decision depth Q-network are... t Based on empirical samples, the network parameters of the target Q-network and empirical samples, and Set the post-decision depth Q network experience replay buffer capacity to X1 and the target Q network experience capacity to X2; set the relevant parameters of the ε-greedy policy, learning rate, and discount factor;
[0110] Output: Access point information, sub-channel allocation, and power allocation information for all users in the indoor aggregated VLC-RF network.
[0111] Step 1: Calculate the channel gain of each user accessing each sub-channel of the VLC AP based on the Lambert radiation model; calculate the channel gain of each user accessing each sub-channel of the RF AP based on the Rayleigh fading model.
[0112] Step 2: Calculate the received signal-to-interference-plus-noise ratio (SINR) for each user on the VLC AP sub-channel, and use the VLC throughput lower bound closure formula to calculate the throughput between the user and the VLC AP sub-channel; calculate the received SINR for each user on the RF AP sub-channel, and use Shannon's law to calculate the throughput between the user and the RF AP sub-channel; calculate the throughput achieved by each user accessing the indoor aggregated VLC-RF network. Total throughput R of indoor aggregated VLC-RF network total The transmit power allocated to users in an indoor aggregated VLC-RF network The total transmit power P consumed by the indoor aggregated VLC-RF network total ;
[0113] Step 3: Represent the energy efficiency optimization problem model for joint access point selection and resource allocation in an indoor aggregated VLC-RF network.
[0114] In an indoor aggregated VLC-RF network, the parameters to be optimized are: access point selection vector x, sub-channel allocation vector c, power allocation vector p, and user communication interruption vector a.
[0115] Step 5: The agent operates in time slot t based on state s t , definite execution action The formula for determining vector values is:
[0116] in, The action values for imitation learning are the imitation learning actions generated by the agent based on the Bregman ball algorithm. value; These are local action values, generated directly by the agent based on an ε-greedy policy.
[0117] Step 4: In each time slot, the agent searches for the historical experience sample with the highest similarity to its current state in the deep Q network experience replay buffer based on the Bregman ball algorithm. By referring to the action strategy of the historical experience sample, it generates an imitation learning action.
[0118] Step 5: Design an agent action scheme that combines imitation learning actions with local actions; design a reward function that aims to maximize network energy efficiency while also taking into account user communication satisfaction.
[0119] Step 6: In each time slot, the agent, based on state s t Take action Receive immediate reward value afterward and each tuple As empirical samples, they are stored in the empirical replay buffer of the post-decision deep Q-network;
[0120] Step 7: Calculate the average sample loss function between the Q-value of the post-decision deep Q-network and the Q-value of the target Q-network. Minimize the average sample loss function using the adaptive moment estimation optimization algorithm. Output the action strategy with the minimum average sample loss function as the optimal action strategy of the post-decision deep Q-network. When the number of samples in the experience buffer reaches X1, update the neural network parameters of the post-decision deep Q-network and update the experience samples of the target Q-network and the deep Q-network.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for optimizing indoor aggregated VLC-RF energy efficiency based on a post-decision deep Q network, characterized in that: The method comprises the following steps: S1: input the user set in indoor hybrid VLC-RF network A set of VLC access points APs 1 RF AP, a set of sub-channels of VLC APs and a set of sub-channels of RF APs wherein the VLC APs multiplex the same set of sub-channels; input the network parameters θ of the post-decision deep Q network t and experience samples, the network parameters θ of the target Q network and experience samples, and set the experience replay buffer capacity of the post-decision deep Q network as X1, set X3 = 0; calculate the channel gain of each sub-channel of the VLC AP accessed by the user according to the Lambert radiation model; calculate the channel gain of the sub-channel of the RF AP accessed by the user according to the Rayleigh fading model; S2: calculate the received signal to interference and noise ratio (SINR) value of the user on the VLC AP subchannel, calculate the throughput between the user and the VLC AP subchannel by using the lower bound closed formula of the throughput of VLC; calculate the received SINR of the user on the RF AP subchannel, calculate the throughput between the user and the RF AP subchannel by using the Shannon law; in the indoor aggregated VLC-RF network, any user can access the RF AP, the user in the coverage area of the VLC AP can simultaneously access a VLC AP and an RF AP, and the throughput R achieved by each user accessing the indoor aggregated VLC-RF network is calculated respectively k , The total throughput R of the indoor aggregated VLC-RF network total , the transmit power P allocated to the user by the indoor aggregated VLC-RF network k , The total transmit power P consumed by the indoor aggregated VLC-RF network total ; S3: Establish the energy efficiency optimization problem model of joint access point selection and resource allocation in indoor hybrid VLC-RF network, which is expressed as η EE RF,k , v,k , x RF,k = 1 and x v,k = 1 represent that user k accesses RF AP and user k accesses VLC AP v, respectively; represent that RF AP allocates subchannel m to user k and VLC AP v allocates subchannel n to user k, respectively; represent that RF AP allocates transmit power on subchannel m to user k and VLC AP v allocates transmit power on subchannel n to user k, respectively; a = {a k , a k = 0 if user is in communication interruption state, otherwise a k = 1; S4: introducing a post-decision state mechanism in the deep Q network, representing the energy efficiency optimization problem model of joint access point selection and resource allocation in the indoor aggregated VLC-RF network as an energy efficiency optimization problem model based on the post-decision deep Q network, abstracting the central controller of the indoor aggregated VLC-RF network as an agent of the deep Q network, and regarding the post-decision state as an intermediate state between the current state and the next state of the agent, representing the deterministic part in the state transition process after the agent performs an action; in the energy efficiency optimization model based on the post-decision deep Q network, the learning task of the agent is described by using a Markov decision process to describe the transition between states; The specific steps in S4 are: S401: regarding the central controller of the indoor aggregated VLC-RF network as an agent of the post-decision deep Q network, and regarding the agent as being responsible for learning and deciding the VLC AP or RF AP, subchannel and power allocation strategy of all users in the VLC-RF network, so as to optimize the energy efficiency of the indoor aggregated VLC-RF network; regarding the network parameters of the post-decision deep Q network as the network parameters of a target Q network, and copying the experience samples of the target Q network to the experience buffer of the post-decision deep Q network; S402: Constructing the state space of the agent The state vector of the agent at time slot t is denoted as where, is the location of user k at time slot t; is the moving speed of user k at time slot t; denotes the state of whether the RF AP allocates subchannel m at time slot t; denotes the state of whether the VLC AP v allocates subchannel n at time slot t; and denote the received signal-to-interference-and-noise ratio (SINR) value of user k from subchannel m of the RF AP and the received SINR value of user k from subchannel n of the VLC AP v at time slot t, respectively; and are used to measure the total transmit power of the RF AP and each VLC AP at time slot t, respectively; k,t is the throughput of user k at time slot t, which is used to ensure that the minimum data rate requirement R min is satisfied;a k,t = {0, 1}, denotes whether the communication of user k at time slot t is interrupted or not. S403: Constructing the action space of the agent The agent performs an action at time slot t based on state s t According to the action policy The access point selection x containing the association of each user with each access point t The sub-channel c assigned to each user by each access point t And the power allocation p t Three decision vectors; S404: Construct the post-decision state space of the agent The agent takes an action After that, the agent moves from state s t to a deterministic post-decision state The constituent elements and meaning of s t are the same as s S405: The agent operates based on state s in time slot t. t Take action Then, a known state transition probability is used. From state s t Transition to post-decision state This is the deterministic phase; subsequently, the agent transitions to an unknown state with probabilities. From the post-decision state Transition to the next state s t+1 This is the random phase; among which, The value is estimated offline and Monte Carlo sampled from historical experience samples in the experience replay buffer of the deep Q network. The agent then searches for a value related to its state s from the experience replay buffer using the Bregman sphere algorithm. t The transition probability is determined by referencing the historical experience sample with the highest similarity. The value; and Action values are directly generated by the agent based on an ε-greedy policy. In the ε-greedy policy, the agent randomly selects a non-greedy action with a probability of ε- based on the current state, and selects the optimal action with the largest Q value from the current policy set of employees with a probability of 1-ε. S5: the agent determines to perform an action t based on the state s at time slot t The determination formula of the vector value is: is the action value of the imitation learning, which is found by the agent from the historical experience samples in the deep Q network experience replay buffer based on the Bregman ball algorithm, and the state s t of the agent is most similar to the state s of the historical experience sample, and the imitation learning action value is generated by referring to the action strategy of the historical experience sample; is the local action value, which is directly generated by the agent based on the ε-greedy strategy; is a weight coefficient, The specific process of S5 is: S501: Calculate the difference between two data points based on Bregman divergence using Bregman ball algorithm, take the state vector s t of the agent as a data point in high-dimensional space, and take it as the center point Z cen of the Bregman ball, and construct a Bregman ball; the set of positions corresponding to the state vector data points of all historical experience samples in the deep Q network experience replay buffer in the Bregman ball is represented as S502: Compute Z = argminzBregman divergence between z and Z cen Bregman divergence between z and Z denotes the gradient of the function f at z, (Z cen denotes the difference between two vectors, denotes the computation of the gradient denotes the inner product of the vector (Z cen -z). S503: judging is established, and only when the condition is established, the data point z is inside the Bregman ball, representing that the historical experience sample corresponding to the data point z is similar to the agent state vector s t has the similarity; for the data points inside the Bregman ball, the data point with the minimum Bregman divergence is selected, representing that the historical experience sample with the highest similarity to the agent state vector s t has the highest similarity; the agent generates an imitation learning action by referring to the action policy of the historical experience sample with the highest similarity wherein the value of the Bregman ball radius needs to ensure that the Bregman ball contains a sufficient number of historical experience samples; S504: determining an action of the agent based on the state s t the local action taken the agent determines an action based on the state s t the agent takes a local action by randomly selecting a non-greedy exploration strategy with a probability of ε, and the agent will select an optimal action that can maximize the Q value function under the current strategy with a probability of 1-ε In addition, in order to reduce the exploration probability of the agent in the learning process, ε is generally gradually reduced with the increase of the time step, and the calculation formula is ε = max{ε-ε δ ,ε min}, wherein ε δ is the reduction amount of ε after each time step, and ε min is the minimum exploration value to which ε can be reduced; where the ε-greedy policy agent takes an action based on state s t Take local action The probability is: In the above formula, Q-value function for the empirical sample Q-value function for the empirical sample in, The calculation formula is: In the above formula, V * (s t+1 ) represents the optimal Q-value function of the system at the state transition of the agent to the next state s t+1 S505: determining the value of the action a by the formula from the imitation learning action and the local action combined with the agent action value; S6: the agent bases the action a on the state s t Take action Post-obtain timely reward value And each tuple As experience samples, stored in the experience replay buffer of the post-decision deep Q network; Wherein, the timely reward value The calculation formula is: In the above formula, is a known state transition probability for the agent from state s t to a post-decision state after which a known reward value is obtained; is an unknown state transition probability for the agent from the post-decision state to the next state s t+1 after which an unknown reward value is obtained; wherein, The calculation formula of the reward value of the i-th user is: In the above formula, η EE represents the probability of the agent moving from state s t to s , which is a benefit term. represents the probability of the agent moving from state s t to s , which is a cost term. represents the probability of the agent moving from state s t to s , which is a cost term. Both ζ1 and ζ2 represent weight coefficients for balancing the benefits and costs in the reward function, and take values between 0 and 1. wherein, The calculation formula of the reward value of the i-th user is: In the above equation, η EE represents the probability that the agent moves from state s to s t+1 when the energy efficiency of the indoor aggregated VLC-RF network is considered as the benefit term; represents the probability that the agent moves from state s to s t+1 when the probability of communication interruption for user k is considered as the cost term; represents the probability that the agent moves from state s t to s t+1 when the probability that the minimum data rate requirement for user k is not satisfied is considered as the cost term; S7: calculating the sample average loss function value between the Q value of the post-decision deep Q network and the Q value of the target Q network, minimizing the sample average loss function value by using an adaptive matrix estimation optimization algorithm, and outputting the optimal action strategy of the post-decision deep Q network as the action strategy with the minimum sample average loss function value; when the number of samples in the experience buffer reaches X1, updating the neural network parameters of the post-decision deep Q network, and updating the experience samples of the target Q network and the deep Q network; The specific process of S7 is: S701: Calculate the Q-value function for the state-action pair of each experience sample in the experience buffer of the post-decision depth Q network and the Q-value function where the Q-value function for a state-action pair and is calculated as In the above formula, Q-value function for state-action pairs of the experience samples, Q-value function for state-action pairs of the experience samples, Q-value function for state-action pairs of the experience samples, Q-value function for state-action pairs of the experience samples, V * (s t+1 ) represents the optimal Q-value function of the system in which the agent transitions from the state of this experience sample to the next state s t+1 S702: update the Q value of each experience sample in the experience buffer of the post-decision deep Q network using the Bellman equation t+1 at the next state s and the corresponding Q value t+1 at the next state s wherein the next state s t+1 of the experience sample is calculated by the formula: wherein the value of the next state s In the above formula, α∈(0,1) represents a learning rate; γ∈[0,1] represents a discount factor of a reward, and is used to measure the importance of future rewards; the closer γ is to 1, the more the agent values future rewards; where s is the next state of the experience sample t+1 The corresponding Q value The calculation formula is: S703: calculating the sample average loss function value between the Q value of the post-decision deep Q network and the Q value of the target Q network; in order to train the post-decision deep Q network model, the difference between the Q value of the estimated deep Q network and the Q value of the target Q network is measured by using the sample average loss function; the smaller the value of the sample average loss function, the closer the prediction result output by the model is to the true result, and the closer the action strategy learned by the agent is to the optimal strategy; The calculation formula of the sample average loss function value between the optimal Q value of the experience sample in the experience replay buffer of the post-decision deep Q network and the optimal Q value of the experience sample in the target Q network is as follows: In the above formula, θ t is the network parameter of the post-decision deep Q network at the current time slot t; is the network parameter of the target Q network; is the network parameter of the X1 experience samples in the experience buffer of the agent t The average maximum expected return value after selecting the optimal action in the next state, that is, wherein, represents the network parameter θ t of the X1 experience samples in the experience buffer of the agent represents the average Q value of the X2 samples in the target Q network, represents the upward rounding integer; S704: regarding the state-action pair corresponding to the post-decision deep Q network with the minimum sample average loss function value as the optimal action strategy of the indoor aggregated VLC-RF network and outputting the optimal action strategy; The optimal action strategy of the state-action pair of the indoor aggregated VLC-RF network is determined as follows: In the above formula, , and represent the state and action corresponding to the post-decision deep Q network with the minimum sample average loss function value, and represent the optimal action set of the user access state, channel selection state, and power allocation state that enable the minimum energy efficiency in the indoor aggregated VLC-RF network. S705: When the number of samples in the experience buffer reaches X1, the network parameters of the updated decision deep Q network are updated to θ t+1 , the network parameters of the updated decision deep Q network are updated to θ t = θ t+1 , and X3 = X3 + 1. Wherein, the adaptive moment estimation optimization algorithm is used to update the network parameter θ t+1 The calculation formula is: In the above formula, α t is the learning rate in the adaptive matrix estimation optimization algorithm; is the first moment estimation value after bias correction, wherein, is the first moment estimation of the gradient of the sample average loss function, i.e., the mean of the gradient, is the exponential decay rate of the first moment estimation, is the gradient value of the sample average loss function L(θ t ) with respect to the post-decision depth Q network parameter θ t ; is the second moment estimation value after bias correction, wherein, is the second moment estimation of the gradient of the sample average loss function, i.e., the mean of the square of the gradient, is the exponential decay rate of the second moment estimation, is the square of the gradient of the sample average loss function with respect to the network parameter θ t ; κ represents a very small constant, to avoid the case where the denominator is zero; the initial values of α t , and κ are set to 0.001, 0.9, 0.999 and 10 -8 , respectively; ω t and χ t are set to 0 at the initial moment; the introduction of bias correction can increase the values of ω t and χ t in the initial stage, to reduce the influence of bias; with the increase of the time step, and tend to 0, and the bias correction factors and tend to 1, and the influence of bias correction gradually decreases; S706: When X3=X2, copy the X2 experience samples with the maximum Q value in the experience replay buffer of the post-decision depth Q network to the target Q network, replace the experience samples of the target Q network, and update the network parameters of the target Q network as: Let X3=0; if there is a new connection request to access the indoor aggregation VLC-RF network, go to step S1.