VLC-RF network resource matching method based on yoke theory and reinforcement learning

By applying martingale theory and reinforcement learning in VLC-RF networks, building user satisfaction utility functions and obtaining optimal service rate combinations, the problems of single QoS indicators and insufficient consideration of user satisfaction in the existing technology are solved, and the optimal allocation of resources and the improvement of network service efficiency are achieved.

CN120090705AActive Publication Date: 2025-06-03JILIN INST OF CHEM TECH

Patent Information

Application Number
CN202510233341.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing VLC-RF network resource matching scheme only focuses on system delay performance, with single QoS indicators, and the nonlinear mapping relationship between user satisfaction and service quality parameters is not fully considered, and it is difficult to achieve optimal resource allocation and network service efficiency improvement.

Method used

Using a method based on martingale theory and reinforcement learning, the VLC-RF heterogeneous network switching service mechanism is constructed, the state transfer matrix of the system service process is derived, the QoS index is evaluated using martingale theory, and the user satisfaction utility function is constructed. At the same time, the optimal service rate combination required by the reinforcement learning is used to obtain the system.

Benefits of technology

It realizes optimal allocation of resources, improves network service efficiency, ensures users' diversified QoS needs, and improves user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090705A_ABST
    Figure CN120090705A_ABST
Patent Text Reader

Abstract

The invention relates to a VLC-RF (Visible Light Communication-Radio Frequency) network resource matching method based on a yoke theory and reinforcement learning, and the method comprises the steps: deducing a state transition matrix of a system service process through constructing a VLC-RF heterogeneous network switching service mechanism; network QoS parameters are evaluated based on a yoke theory, and a user satisfaction utility function is constructed in combination with queue reliability QoS, time delay reliability QoS and jitter reliability QoS. And finally, acquiring a service rate combination required to be provided by a system network by adopting reinforcement learning so as to realize optimal allocation of resources and improve the network service efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technology applications, and particularly to a VLC-RF network resource matching method based on martingale theory and reinforcement learning. Background Art

[0002] The integration of VLC-RF networks is an important trend in the development of communication network technologies, enabling more efficient, flexible, and reliable network communication. VLC-RF resource matching schemes often only focus on the system delay performance, and the QoS metrics for measuring network performance are relatively single, and the non-linear mapping relationship between user satisfaction and service quality parameters is not fully considered. The present invention derives the queue reliability QoS, delay reliability QoS, and jitter reliability QoS based on martingale theory; and constructs a user satisfaction utility function by jointly considering multi-dimensional QoS metrics, achieving the expected user satisfaction on the premise of ensuring the diverse QoS requirements of users. In addition, relying on the ability of reinforcement learning to handle continuous and complex decision-making problems, it can solve the optimal service rate combination of dynamic systems, ensure the QoS requirements of users, and efficiently improve network service efficiency. Summary of the Invention

[0003] The object of the present invention is a VLC-RF network resource matching method based on martingale theory and reinforcement learning to solve the problems existing in the above-mentioned prior art. This method can achieve the optimal allocation of resources and improve network service efficiency.

[0004] To achieve the above object, the present invention provides the following solution:

[0005] A VLC-RF network resource matching method based on martingale theory and reinforcement learning includes:

[0006] Step 1, construct a VLC-RF heterogeneous network handover service mechanism and derive the state transition matrix of the system service process; wherein, the VLC-RF heterogeneous network handover service mechanism is described by a Markov chain model with 2N - 1 states;

[0007] Step 2, based on the state transition matrix, use martingale theory to derive QoS metrics and establish a user satisfaction utility function; wherein, the QoS metrics include: martingale domain queue reliability QoS, martingale domain delay reliability QoS, and martingale domain jitter reliability QoS;

[0008] Step 3, based on user satisfaction, use reinforcement learning to obtain the optimal service rate combination required by the system.

[0009] Optionally, deriving the state transition matrix of the system service process includes:

[0010] Based on the VLC-RF heterogeneous network handover service mechanism, establish a communication scenario model for the VLC-RF system;

[0011] Based on the communication scenario model of the VLC-RF system, derive the state transition matrix of the system service process.

[0012] Optionally, the communication scenario model of the VLC-RF system includes: N VLC APs and 1 RF AP;

[0013] Among them, the VLC AP signals overlap with each other and cover a preset range, and the RF AP signal covers the entire indoor environment;

[0014] When the user is in the VLC signal overlapping area, the VLC AP is in an unavailable state, and the system will promptly switch to the RF network to provide services for the user; otherwise, the VLC AP is in an available state, and the system preferentially uses the VLC network to provide services for the user.

[0015] Optionally, the state transition matrix is:

[0016]

[0017] Among them, t i,j represents the transition probability from state k i to state k j state,

[0018] Optionally, obtaining the user satisfaction includes:

[0019] Based on the state transition matrix, use the supermartingale theory to construct the probability bound of queue overflow;

[0020] Based on the probability bound of queue overflow, derive the queue reliability QoS requirements; based on the probability bound of queue overflow, combined with Little's theorem, obtain the probability bound of martingale domain delay violation of the system, and derive the delay reliability QoS requirements; based on the martingale theory, use the probability generating function to derive the jitter reliability QoS requirements;

[0021] Optionally, the probability bound of queue overflow is:

[0022]

[0023] Among them, σ represents the threshold of the system queue length, l(n) represents the system queue length at the nth moment, θ * represents the supermartingale exponent, h a , h s are the characteristic functions of the arrival process and the service process respectively, E[h a (a(0))], E[h s(s(0)) represent the initial value expectations of the arrival and service process characteristic functions respectively, and H represents the threshold value;

[0024] Based on the QoS metrics, construct a user satisfaction utility function;

[0025] Based on the user satisfaction utility function, obtain the user satisfaction.

[0026] Optionally, the QoS metrics include:

[0027] Martingale domain queue reliability QoS:

[0028]

[0029] where δ is the system violation probability threshold;

[0030] Martingale domain delay reliability QoS:

[0031]

[0032] where D represents the target delay, d(n) represents the system delay, C represents the arrival rate; ε represents the delay violation probability threshold;

[0033] Martingale domain jitter reliability QoS:

[0034]

[0035] where J r represents the system steady-state jitter, J ε is the jitter constraint threshold. Ψ(z) is the probability generating function of the network queue size, ψ'(z) represents ψ'(z) and ψ”(z) respectively represent the first and second derivative functions of Ψ(z), z represents the martingale parameter, and C represents the average arrival rate.

[0036] The user satisfaction utility function is expressed as:

[0037]

[0038] where U l (L(n)) represents the current network system queue reliability QoS normalization function, U d (D(n)) represents the current network system delay reliability QoS normalization function, U jr (J(n)) represents the current network system jitter reliability QoS normalization function, and R represents the system average service rate.

[0039] Optionally, based on the user satisfaction, using reinforcement learning to obtain the optimal service rate combination that the system needs to provide includes:

[0040] Construct the environment of the reinforcement learning agent based on user satisfaction;

[0041] Use the Advantage Actor-Critic algorithm to update the parameters of the Actor policy network and the Critic value network; the Actor network is used to learn the resource matching policy and generate the actions of the selection network. The Critic network is used to evaluate the value corresponding to the network environment state or state-action; through multiple rounds of interaction and iteration between the agent and the environment, obtain the optimal service rate combination required by the system.

[0042] Optionally, constructing the environment of the reinforcement learning agent based on user satisfaction includes:

[0043] Define the state vector of the network environment where the agent is located at time slot n;

[0044] Set the action of the agent at time slot n; where the agent refers to the user and the action refers to the network selected by the user;

[0045] Based on user satisfaction, obtain the immediate reward R(n) obtained by the agent at time slot n; where, when the user satisfaction obtained by the agent performing action a is greater than the minimum satisfaction required by the user, it is a positive reward, and when the user satisfaction obtained by the agent performing action a is less than the minimum satisfaction required by the user, it is a negative reward;

[0046] The immediate reward R(n) is:

[0047]

[0048] where, ζ represents the reward adjustment factor, QoS Γ represents the current user satisfaction, QoS min represents the minimum user satisfaction.

[0049] Optionally, using the Advantage Actor-Critic algorithm to update the parameters of the policy network and the value network includes:

[0050] The agent collects state, action, and reward data through interaction with the environment, and uses the policy gradient method in the Advantage Actor-Critic algorithm to update the Actor policy network to optimize the action selection policy, and at the same time update the Critic value network.

[0051] The beneficial effects of the present invention are:

[0052] The present invention discloses a VLC-RF network resource matching method based on martingale theory and reinforcement learning. This method uses martingale theory to evaluate network QoS parameters and constructs a user satisfaction utility function by combining queue reliability QoS, delay reliability QoS, and jitter reliability QoS in the martingale domain. At the same time, a reinforcement learning algorithm is used to match corresponding service rates for users to achieve optimal resource allocation and improve network service efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 It is a schematic flowchart of a VLC-RF network resource matching method based on martingale theory and reinforcement learning according to an embodiment of the present invention;

[0055] Figure 2 It is a schematic diagram of the VLC-RF network switching process according to an embodiment of the present invention;

[0056] Figure 3 It is a schematic diagram of the area corresponding to the network state according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0058] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0059] As Figure 1 shown, this embodiment proposes a VLC-RF network resource matching method based on martingale theory and reinforcement learning, including:

[0060] Step 1: Construct a VLC-RF heterogeneous network handover service mechanism and deduce the state transition matrix of the system service process; among them, the VLC-RF heterogeneous network handover service mechanism of the VLC-RF network resource matching method based on martingale theory and reinforcement learning is described by a Markov chain model with 2N - 1 states;

[0061] Step 2: Based on the state transition matrix of the VLC-RF network resource matching method using martingale theory and reinforcement learning, deduce the QoS metrics using martingale theory, construct a user satisfaction utility function, and obtain the user satisfaction. Among them, the QoS metrics of the VLC-RF network resource matching method based on martingale theory and reinforcement learning include: martingale domain queue reliability QoS, martingale domain delay reliability QoS, and martingale domain jitter reliability QoS.

[0062] Step 3: Based on the user satisfaction, use reinforcement learning to obtain the optimal service rate combination that the system needs to provide.

[0063] In this embodiment, a handover service mechanism for a visible light communication (VLC)-radio frequency (RF) heterogeneous network is constructed to deduce the state transition matrix of the system service process. Based on martingale theory, the network QoS parameters are evaluated, and a user satisfaction utility function combining queue reliability QoS, delay reliability QoS, and jitter reliability QoS is constructed. Finally, reinforcement learning is used to obtain the service rate combination that the system network needs to provide, so as to achieve optimal resource allocation and improve network service efficiency.

[0064] Furthermore, deducing the state transition matrix of the system service process includes:

[0065] Based on the VLC-RF heterogeneous network handover service mechanism of the VLC-RF network resource matching method using martingale theory and reinforcement learning, establish a VLC-RF system communication scenario model;

[0066] Based on the VLC-RF system communication scenario model of the VLC-RF network resource matching method using martingale theory and reinforcement learning, deduce the state transition matrix of the system service process.

[0067] Furthermore, the VLC-RF system communication scenario model of the VLC-RF network resource matching method using martingale theory and reinforcement learning includes: N VLC APs and 1 RF AP;

[0068] Among them, the VLC AP signals overlap with each other and cover a preset range, and the RF AP signal covers the entire indoor environment;

[0069] When the user is in the VLC signal overlapping area, the VLC AP is in an unavailable state, and the system will promptly switch to the RF network to provide services for the user; otherwise, the VLC AP is in an available state, and the system preferentially uses the VLC network to provide services for the user.

[0070] Specifically, in this embodiment, step 1: Establish a communication scenario model of the VLC-RF system by designing a VLC-RF heterogeneous network handover service mechanism, and deduce the state transition matrix of the system service process; specifically including:

[0071] This embodiment mainly considers the communication scenario of a VLC-RF heterogeneous network composed of N VLC APs and 1 RF AP. The signals of VLC APs overlap with each other and have a limited coverage range; the signal of the RF AP covers the entire indoor environment. When the user is in the VLC signal overlapping area, the VLC AP is in an unavailable state, and the statistical QoS of the user is difficult to guarantee. The system will promptly switch to the RF network to provide services for the user. Otherwise, the VLC AP is in an available state, and the system preferentially uses the VLC network to provide services for the user. The present invention models the process of the system providing services for the user into a handover service mechanism of a heterogeneous network, so that the user always maintains an online state during the communication process.

[0072] Each AP in the system can only switch to other adjacent APs. The handover process between networks is only related to the network QoS state at the previous moment and has nothing to do with the historical state. This fact conforms to the Markov property. Therefore, the present invention uses a Markov chain model with 2N-1 states to describe the handover service process of the system, as shown in Figure 2 . The set of all network states in the system is represented by , where the network state (k 1 , k 3 ,..., k (2N-1) ) is the available state of the VLC AP; and the network state (k 2 , k 4 ,..., k (2N-2) ) is the unavailable state of the VLC network. The area covered by the c-th network state is S c (c = 1, 2,..., 2N-1). The areas covered by specific network states are as shown in Figure 3 .

[0073] When there is one user in the system, let i represent the row of the system state transition probability matrix X, and j represent the column of the system state transition probability matrix X. Then the matrix X can be modeled as follows:

[0074]

[0075] where t i,j represents the transition probability from the k i state to the k j state. t i,j needs to satisfy the following several constraint conditions:

[0076]

[0077] Represents the set of networks adjacent to the currently connected network, and There is no state k in the system 0 with k 2N , so the coverage area S corresponding to its network state 0 and S N also does not exist, that is

[0078] Constraint (2) indicates that t i,j is a real number between 0 and 1; Constraint (3) indicates that as long as the handover is successful, the user must be served by a specific AP; Constraint (4) indicates that the handover probability between all non-adjacent networks is 0; Constraint (5) is the expression of the transition probability between adjacent networks.

[0079] Therefore, the exponential column transformation of matrix X is:

[0080]

[0081] where, (R VLC1 , R VLC2 , …, R VLCN ) is the service rate corresponding to (VLC AP 1 , VLC AP 2 ,..., VLC AP N ), and R RF is the service rate of the RF network.

[0082] When there are M users in the system, it is only necessary to sequentially analyze the probability of the number of users existing within the coverage area of the available VLC AP and multiply it by the corresponding elements of Equation (1) to obtain the system state transition probability matrix for the existence of multiple users When there are m (0 ≤ m ≤ M) users in the S c area, it can be obtained from the following formula:

[0083]

[0084] where S is the total area of the indoor scene.

[0085] The S c area is the coverage area corresponding to the (c - 1)th VLC AP, and the service rate it provides is When the S c area is the coverage area corresponding to the RF AP, the service rate it provides is R RF / m. Assume that when there are no users in the S c area, the system does not provide services.

[0086] The corresponding exponential sequence transformation matrix can be obtained by the calculation method of Equation (6), and its dimension will be expanded to (((m + 1)(2N - 1)) × ((m + 1)(2N - 1))).

[0087] Furthermore, obtaining user satisfaction includes:

[0088] Based on the state transition matrix of the VLC-RF network resource matching method based on martingale theory and reinforcement learning, using the supermartingale theory, construct the probability bound of queue overflow;

[0089] Based on the probability bound of the queue overflow, deduce the QoS requirements of queue reliability; based on the probability bound of the queue overflow, combined with Little's theorem, obtain the probability bound of the martingale domain delay violation of the system, and deduce the QoS requirements of delay reliability; based on martingale theory, use the probability generating function to deduce the QoS requirements of jitter reliability;

[0090] Based on the QoS metrics of the VLC-RF network resource matching method based on martingale theory and reinforcement learning, construct a user satisfaction utility function;

[0091] Based on the user satisfaction utility function of the VLC-RF network resource matching method based on martingale theory and reinforcement learning, obtain user satisfaction.

[0092] Specifically, in this embodiment, step 2: Deduce the martingale domain queue reliability QoS, martingale domain delay reliability QoS, and martingale domain jitter reliability QoS based on martingale theory, and establish a user satisfaction utility function. Specifically, it includes:

[0093] Step (2.1): Construction of the probability bound of martingale domain queue overflow

[0094] The construction of the probability bound of queue overflow can be obtained from the following formula:

[0095]

[0096] θ * : = sup{θ > 0: K a (θ) = K s (θ)} (9)

[0098] H: min{h a (a(n))h s (s(n)): a(n) - s(n) > 0} (10)

[0100] where the threshold of the system queue length is σ, and l(n) is the system queue length at the nth moment. θ is a QoS parameter and θ > 0, and θ * is the supermartingale exponent.

[0101] Let the supermartingale correction function \(K\) of the system arrival process a (\(\theta\)) take the value of the arrival rate \(C\), that is, \(K\) a (\(\theta\)) = \(C\). The supermartingale correction function \(K\) of the service process s (\(\theta\)) takes the following value:

[0102]

[0103] where \(sp(X\) θ ) is the spectral radius.

[0104] Step (2.2): Martingale domain queue reliability QoS

[0105] Define the queue evolution process of the network system as the queue reliability QoS. To ensure the queue reliability QoS, the following equation needs to hold:

[0106]

[0107] where \(\delta\) is the system violation probability threshold.

[0108] Step (2.3): Martingale domain delay reliability QoS

[0109] The delay reliability QoS is defined as the probability that the system delay exceeds the target delay. Combining Little's theorem with Equation (12), the requirements for the delay reliability QoS need to satisfy the following equation:

[0110]

[0111] where \(D\) is the target delay, \(d(n)\) is the system delay, and \(\varepsilon\) is the delay violation probability threshold.

[0112] Step (2.4): Martingale domain jitter reliability QoS:

[0113] Define the martingale process related to the queue length, in the following form:

[0114]

[0115] where \(z\) is the martingale parameter, \(X(n)\) represents the system queue length observed at the \(n\)th departure, \(I[\cdot]\) is the indicator function, and \(b(z)\) is the probability generating function of the number of data packets arriving at the system, expressed as:

[0116] \(b(z)=E(z\) B(n) ) (15)

[0117] In the formula, \(B(n)\) represents the number of data packets arriving at the system when the \(n\)th data packet is served.

[0118] The probability generating function \(\psi(z)\) of the network queue size is:

[0119]

[0120] where ρ = C / R, and R is the average service rate of the system.

[0121]

[0122] The jitter reliability QoS constraint is defined as that the steady-state jitter must be lower than the maximum allowable jitter threshold of the system. That is, it is necessary to ensure that the following formula holds:

[0123]

[0124] where J r is the steady-state jitter, and J ε is the jitter constraint threshold.

[0125] Step (2.5): Construct the user satisfaction utility function;

[0126] The present invention synthesizes the queue reliability QoS, delay reliability QoS, and jitter reliability QoS metrics, constructs the user satisfaction utility function, and uses this to measure the user satisfaction.

[0127]

[0128] where U l (L(n)) is the normalization function of the queue reliability QoS of the current network system, and U d (D(n)) is the normalization function of the delay reliability QoS of the current network system, and U jr (J(n)) is the normalization function of the jitter reliability QoS of the current network system.

[0129] Furthermore, based on the user satisfaction, obtaining the optimal service rate combination required by the system by using reinforcement learning includes:

[0130] Based on the user satisfaction, construct the environment of the reinforcement learning agent;

[0131] Use the Advantage Actor-Critic algorithm to update the parameters of the Actor policy network and the Critic value network; in the reinforcement learning algorithm, the agent needs to continuously experiment in the environment and continuously optimize the state-action correspondence relationship through the feedback (reward) given by the environment.

[0132] Through multiple rounds of interaction and iteration between the agent and the environment, obtain the optimal service rate combination required by the system.

[0133] Furthermore, constructing the environment of the reinforcement learning agent based on the user satisfaction includes:

[0134] Define the state vector of the network environment where the agent is located at time slot n;

[0135] Set the action of the agent at time slot n;

[0136] Based on user satisfaction, obtain the immediate reward R(n) obtained by the agent at time slot n.

[0137] Furthermore, use the Advantage Actor-Critic algorithm to update the policy network and value network parameters, including:

[0138] The agent collects state, action, and reward data through interaction with the environment, and uses the policy gradient method in the Advantage Actor-Critic algorithm to update the Actor policy network to optimize the action selection strategy, while updating the Critic value network.

[0139] Specifically, in this embodiment, step 3 uses reinforcement learning to obtain the optimal service rate combination required by the system, improve the service efficiency of the heterogeneous network, and achieve efficient allocation of bandwidth resources. Specifically, it includes:

[0140] Step (3.1) Construct the environment of the reinforcement learning agent;

[0141] Step (3.1.1) State vector s;

[0142] Define the state vector of the network environment where the agent is located at time slot n as:

[0143] s = {s 1 (n), s 2 (n)…, s m (n), …, s M (n)} (20)

[0144] Where represents the current network condition of the m-th user.

[0145] Where W S is the network set selected by the agent, that is is the service rate set corresponding to the network selected by the agent, that is R RF is the service rate of the RF selected by the agent, D, σ, J r are the various QoS metrics of the agent.

[0146] Step (3.1.2) Action a;

[0147] The present invention sets the action of the agent at time slot n as a ∈ {A 1 (n), A2 (n), …, A m (n), …, A M (n)}, A m (n) represents the network parameters selected by the m-th user in the current state s.

[0148] Among them, A m (n) ∈ {W K , R VK , R RK}}, W K is the set of all candidate networks, and the candidate system must be adjacent to the current access network. R VK is the set of candidate service rates required to be provided by all VLC networks. R RK is the set of candidate service rates required to be provided by the RF network.

[0149] Step (3.1.3) Reward R(n);

[0150] The immediate reward R(n) obtained by the agent at time slot n is respectively expressed as:

[0151]

[0152] At this time, QoS Γ represents the current user satisfaction. QoS min represents the minimum user satisfaction. ζ is the reward adjustment factor. When the agent performs action a and the obtained user satisfaction is greater than the minimum satisfaction required by the user, it is a positive reward. When the agent performs action a and the obtained user satisfaction is less than the minimum satisfaction required by the user, it is a negative reward.

[0153] Step (3.2) Update the policy network and value network parameters using the Advantage Actor-Critic (A2C) algorithm;

[0154] The policy gradient method used in the A2C algorithm is:

[0155]

[0156] Among them, J(φ) represents the performance of the target policy, represents the policy gradient, π(a t |s t ) represents the probability of selecting action a t in state s t . B(s t ) is the baseline function. Q π (s t , a t ) is the action-value function.

[0157] y t = rt +γv(s t+1 ; w) (23)

[0158] μ t = v(s t ; w) - y t (24)

[0159] where r t is the observed reward, γ is the discount factor, v(s t+1 ; w) represents the estimation of the state value of the value network at time t+1, and w are the neural network parameters.

[0160] Next, the approximate policy gradient is used to update the Actor policy network Λ and the Critic value network w:

[0161]

[0162]

[0163] Λ are the neural network parameters, and α, κ are the learning rates.

[0164] Through multiple rounds of interactive iteration between the agent and the environment, the service rate parameters of the VLC-RF network will gradually tend to the optimal state, thereby realizing the reinforcement learning task in the high-dimensional state space with efficient continuous actions. According to the VLC-RF network resource matching method based on martingale theory and reinforcement learning of the present invention, the system can timely match the corresponding service rate for users to achieve the expected user satisfaction.

[0165] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A VLC-RF network resource matching method based on martingale theory and reinforcement learning, characterized in that: include: Step 1, constructing a VLC-RF heterogeneous network switching service mechanism and deriving a state transition matrix of a system service process; wherein the VLC-RF heterogeneous network switching service mechanism is described by a Markov chain model with 2N-1 states; Step 2: Based on the state transfer matrix, QoS indicators are derived using martingale theory, and a user satisfaction utility function is constructed to obtain user satisfaction; wherein the QoS indicators include: martingale domain queue reliability QoS, martingale domain delay reliability QoS, and martingale domain jitter reliability QoS; Step 3: Based on user satisfaction, reinforcement learning is used to obtain the optimal service rate combination that the system needs to provide.

2. The VLC-RF network resource matching method based on martingale theory and reinforcement learning according to claim 1, characterized in that: The state transition matrix of the derivation system service process includes: Based on the VLC-RF heterogeneous network switching service mechanism, a VLC-RF system communication scenario model is established; Based on the VLC-RF system communication scenario model, a state transition matrix of a system service process is derived.

3. The VLC-RF network resource matching method based on martingale theory and reinforcement learning according to claim 2, characterized in that: The VLC-RF system communication scenario model includes: N VLC APs and 1 RF AP; Among them, VLC AP signals overlap each other and cover a preset range, and RF AP signals cover the entire indoor environment; When the user is in the VLC signal overlapping area, the VLC AP is unavailable and the system will switch to the RF network in time to provide services for the user; otherwise, the VLC AP is available and the system will give priority to using the VLC network to provide services for the user.

4. The VLC-RF network resource matching method based on martingale theory and reinforcement learning according to claim 1, characterized in that: The state transfer matrix is: Among them, t i,j Indicates that from k i State to k j The transition probability of the state, 5. The VLC-RF network resource matching method based on martingale theory and reinforcement learning according to claim 1, characterized in that: Obtaining user satisfaction includes: Based on the state transfer matrix, the probability bound of queue overflow is derived using the supermartingale theory; Based on the probability bound of queue overflow, the queue reliability QoS requirement is derived; based on the probability bound of queue overflow, combined with Little's theorem, the system martingale delay violation probability bound is obtained, and the delay reliability QoS requirement is derived; based on the martingale theory, using the probability generating function, the jitter reliability QoS requirement is derived; Based on QoS index requirements, construct user satisfaction utility function; Based on the user satisfaction utility function, user satisfaction is obtained.

6. The VLC-RF network resource matching method based on martingale theory and reinforcement learning according to claim 1, characterized in that: The probability bound of the queue overflow is: Among them, σ represents the threshold of the system captain, l(n) represents the system captain at the nth moment, and θ * represents the supermartingale index, h a ,h s are the characteristic functions of the arrival process and the service process, E[h a (a(0))],E[h s (s(0))] represent the initial value expectation of the arrival and service process characteristic functions, respectively, and H represents the threshold value; The QoS indicators include: Martingale Queue Reliability QoS: Among them, δ is the system violation probability threshold; Martingale Domain Delay Reliability QoS: Where D represents the target delay, d(n) represents the system delay, C represents the arrival rate, and ε represents the delay violation probability threshold. Martingale Domain Jitter Reliability QoS: Among them, J r represents the steady-state jitter of the system, J ε is the jitter constraint threshold, Ψ(z) is the probability generating function of the network queue size, ψ'(z) and ψ”(z) represent the first and second order derivative functions of Ψ(z), respectively, and z represents the martingale parameter.

7. The VLC-RF network resource matching method based on martingale theory and reinforcement learning according to claim 5, characterized in that: The user satisfaction utility function is expressed as: Among them, U l (L(n)) represents the QoS normalization function of the current network system queue reliability, U d (D(n)) represents the delay reliability QoS normalization function of the current network system, U jr (J(n)) represents the jitter reliability QoS normalization function of the current network system, and R represents the average service rate of the system.

8. The VLC-RF network resource matching method based on martingale theory and reinforcement learning according to claim 1, characterized in that: Based on user satisfaction, the optimal service rate combination required by the system using reinforcement learning includes: Build an environment for reinforcement learning agents based on user satisfaction; The Advantage Actor-Critic algorithm is used to update the parameters of the Actor strategy network and the Critic value network. The Actor network is used to learn resource matching strategies and generate actions for selecting networks. The Critic network is used to evaluate the value corresponding to the network environment state or state action. Through multiple rounds of interactive iterations between the agent and the environment, the optimal service rate combination required by the system is obtained.

9. The VLC-RF network resource matching method based on martingale theory and reinforcement learning according to claim 8, characterized in that: Based on user satisfaction, the environment for building a reinforcement learning agent includes: Define the state vector of the network environment where the agent is located in time slot n; Set the action of the agent in time slot n; where the agent refers to the user and the action refers to the network selected by the user; Based on user satisfaction, obtain the instant reward R(n) obtained by the agent in time slot n; where the reward is positive when the user satisfaction obtained by the agent for action a is greater than the minimum satisfaction required by the user, and negative when the user satisfaction obtained by the agent for action a is less than the minimum satisfaction required by the user; The instant reward R(n) is: Among them, ζ represents the reward adjustment factor, QoS Γ Indicates the current user satisfaction, QoS min Indicates the minimum user satisfaction.

10. The VLC-RF network resource matching method based on martingale theory and reinforcement learning according to claim 8, characterized in that: Using the Advantage Actor-Critic algorithm to update the policy network and value network parameters includes: The agent collects state, action, and reward data through interaction with the environment, and uses the policy gradient method in the Advantage Actor-Critic algorithm to update the Actor policy network to optimize the action selection strategy and update the Critic value network at the same time.

Citation Information

Patent Citations

  • Differentiated network traffic bandwidth demand estimation method based on martingale theory

    CN112929217A

  • VLC / RF network random compensation method based on yoke theory

    CN115087116A

  • Routing optimization method and device, equipment and medium

    CN115499365A

  • Multi-thread virtual network mapping and performance adjusting method based on entropy weight method

    CN116405385A

  • Visible light heterogeneous network communication resource allocation method and system based on reinforcement learning

    CN117119593A

Cited By

  • 6G HRLLC-oriented multi-dimensional intelligent service quality assurance method and system

    CN120935668A