A reinforcement learning-based approach to optimizing end-system access delay and jitter
Through the access selection algorithm based on reinforcement learning, the Markov model is used to train network parameters, collect and combine network status information, and optimize network access point selection. This solves the delay jitter problem of 5G terminals in complex network scenarios and realizes intelligent access judgment and resource optimization.
Patent Information
- Application Number
- CN202411601608.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-11-11
AI Technical Summary
Existing 5G terminal access algorithms are insufficient in optimizing latency and jitter, cannot effectively cope with complex and changing network scenarios, and lack intelligent access judgment.
An access selection algorithm based on reinforcement learning is adopted. The reinforcement learning model is obtained through the random access algorithm. The network parameters are trained using the Markov model. The network status information is collected and combined to optimize the network access point selection and achieve the optimal delay and jitter performance.
It realizes access judgment based on delay jitter as the main performance, improves the adaptability of 5G mobile terminals in complex network scenarios, and optimizes network resource utilization.
Smart Images

Figure CN119485398B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of terminal access, and in particular relates to a method for optimizing terminal system access delay jitter based on reinforcement learning. Background Art
[0002] In recent years, the demand for high-quality, anytime, anywhere broadband wireless network access has become increasingly urgent, leading to the rapid development of various wireless network technologies adapted to different scenarios. Mobile cellular networks have evolved from the Global System for Mobile Communications (GSM) to the Universal Mobile Telecommunications System (UMTS) and further to Long Term Evolution (LTE), providing wide-area network coverage and seamless mobility. Furthermore, the development of a series of 802.11 wireless local area network (WLAN) standards and the 802.16 wireless metropolitan area network (WiMAX) standards have provided users with high-speed wireless connections. Within the coverage area of a GSM base station, wireless access points such as LTE, WLAN, and WiMAX have been added simultaneously to facilitate data transmission, becoming a common practice in the industry. This has gradually led to the formation of heterogeneous wireless networks (HWNs), where multiple networks coexist and have overlapping coverage areas. Furthermore, these networks vary in signal coverage, uplink and downlink transmission rates, and optimally supported service types. Therefore, no single network technology can effectively support all different user services. With the development of HWNs, these networks are competing, complementing, and promoting each other during their respective evolutionary processes, ultimately making HWN convergence inevitable. During HWN convergence, each wireless network serves as an access network and connects users using different architectures and protocols. After convergence, these networks are interconnected through a common IP core network. Multimode mobile users can select a single access point to access the Internet through the common IP core network. Network access selection is a key technology in HWN convergence, controlling user access requests and selecting a network to provide connectivity. The question of how to leverage the unique characteristics of HWNs, such as diverse access technologies, overlapping network architectures, and multi-service traffic loads, while providing access selection while ensuring service quality (QoS) and optimizing wireless resource utilization, has become a hot topic in HWN research.
[0003] The design of network access selection algorithms is directly related to user experience and network resource utilization. Extensive research has been conducted on HWN access selection algorithms. This article categorizes various algorithms based on criteria such as received signal strength (RSS), load balancing, and service quality of service (QoS). Furthermore, based on the mathematical model employed, these algorithms are further categorized into those based on multi-attribute decision-making, utility functions, fuzzy logic, and game theory.
[0004] As network terminal applications continue to develop and advance, the demand for network performance is also increasing. Currently, with the increasing emergence of high-real-time traffic, 5G terminal devices have increasingly stringent requirements for network latency and jitter. While ensuring low latency, access nodes must also provide very low latency and jitter performance. Currently, 5G terminal access algorithms have made significant progress in achieving high speeds, low latency, and high reliability. 5G networks utilize a variety of advanced technologies, such as massive multiple-input multiple-output (MIMO), beamforming, and millimeter wave frequency bands, to optimize terminal access performance. Access algorithms have been improved in areas such as resource allocation, interference management, and load balancing, enabling 5G networks to support more terminal devices and higher data rates. However, there is no intelligent access technology that uses latency and jitter as the primary performance factor for access decisions. Summary of the Invention
[0005] In response to the above-mentioned deficiencies in the prior art, the present invention provides a method for optimizing terminal system access delay jitter based on reinforcement learning, which solves the problem that the existing access technology does not use delay jitter as the main performance for access judgment and the problem that the complex and changeable network scenarios of 5G mobile terminals are difficult to cope with.
[0006] To achieve the above objectives, the present invention adopts a technical solution: a method for optimizing end system access delay jitter based on reinforcement learning, comprising the following steps:
[0007] S1. Use a random access algorithm to connect a 5G access device to a network access point, and use an access selection algorithm to obtain a reinforcement learning model based on the access selection algorithm.
[0008] S2. Train a reinforcement learning model based on the access selection algorithm, and transmit the trained reinforcement learning model to all 5G terminal devices that use the access node for network transmission;
[0009] S3. Use 5G terminal equipment to collect network information of network access points;
[0010] S4. Quantify all collected network information and combine it into a network state matrix. Input the network state matrix into the trained reinforcement learning model, output a quantized value, and select the network access point with the best delay and jitter performance based on the output quantized value, and perform network switching.
[0011] S5. Connect the 5G terminal device to the network, obtain the network service of the network access point corresponding to the optimal delay and jitter performance, and complete the end system access delay and jitter optimization.
[0012] The beneficial effects of the present invention are: by using access selection algorithms, reinforcement learning models and data collection, the present invention realizes access judgment with delay jitter as the main performance and enables 5G mobile terminals to effectively cope with complex and changeable network scenarios.
[0013] Furthermore, the specific steps of S1 are as follows:
[0014] Using the random access algorithm, a network access point with the operator access node identifier is randomly generated in the 5G access device. According to the operator access node identifier, the 5G terminal device is used to connect to the 5G network of the corresponding operator, and the access selection algorithm is used according to the network access point to obtain a reinforcement learning model based on the access selection algorithm.
[0015] Furthermore, the S2 includes the following steps:
[0016] S201. Obtain a reinforcement learning model and a training data set based on an access selection algorithm for each access node in a 5G access device, and initialize the reinforcement learning model and the training data set based on the access selection algorithm using a Markov model;
[0017] S202. Using access nodes to collect network data uploaded by 5G terminals, and accessing device parameters and device performance, the data training samples are obtained by combining them. Based on the data training samples, the parameters of the reinforcement learning model are trained using Q learning.
[0018] S203: Determine whether the time reaches the training cycle or whether there is a terminal requesting model parameters. If so, execute S204; otherwise, execute S202.
[0019] S204: using the parameters of the trained reinforcement learning model to optimize the objective function of the reinforcement learning model, training the reinforcement learning model based on the access selection algorithm;
[0020] S205: Transmit the trained reinforcement learning model to all 5G terminal devices that use the access node for network transmission;
[0021] The objective function of the reinforcement learning model is expressed as follows:
[0022]
[0023] Among them, L(θ) represents the objective function of the reinforcement learning model, r t (θ) represents the reward ratio, ε represents the estimated factor of the advantage function, represents a positive number in the estimation factor of the advantage function, clip(·) represents contrastive language-image pre-training, θ represents the model parameters of reinforcement learning, θ old Represents the model parameters of the previous round of reinforcement learning, a t Indicates execution of an action, s t Represents the state, π θ Represents the algorithm strategy of the reinforcement learning model parameters, π θold The algorithm policy representing the model parameters of the previous round of reinforcement learning.
[0024] The beneficial effect of the above further solution is that the present invention utilizes the Markov model to learn and strengthen the previous access status and network performance, and trains the network parameters, thereby improving the network delay jitter performance for node selection.
[0025] Furthermore, the S3 includes the following steps:
[0026] S301, dividing network status information into terminal node network information, access node status information and flow status information;
[0027] S302. Collect terminal node network information using 5G terminal equipment;
[0028] S303. Collect access node status information using a 5G terminal device;
[0029] S304. Collect flow status information using 5G terminal equipment.
[0030] The beneficial effects of the above further scheme are: the present invention collects clear network status information, including terminal node network information, access node status information and flow status information, which improves the accuracy of the network status matrix composed of network status information, and enables the reinforcement learning access selection model to have more accurate actual data, thereby improving the ability of the present invention to cope with complex and changing 5G mobile network scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Flow chart of the method of the present invention.
[0032] Figure 2 This is a flow chart of the process of model training and issuing to access nodes in this embodiment.
[0033] Figure 3 This is a schematic diagram of access selection in this embodiment. DETAILED DESCRIPTION
[0034] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0035] Example
[0036] like Figure 1 As shown, the present invention provides a method for optimizing end system access delay jitter based on reinforcement learning, and its implementation method is as follows:
[0037] S1. Use the random access algorithm to connect the 5G access device to the network access point, and use the access selection algorithm to obtain a reinforcement learning model based on the access selection algorithm. The specific steps are as follows:
[0038] Using the random access algorithm, a network access point with the operator access node identifier is randomly generated in the 5G access device. According to the operator access node identifier, the 5G terminal device is used to connect to the 5G network of the corresponding operator, and the access selection algorithm is used according to the network access point to obtain a reinforcement learning model based on the access selection algorithm.
[0039] In this embodiment, a random access algorithm is used to randomly generate an operator access node identifier. The 5G terminal device connects to the corresponding operator's 5G network based on the identifier and obtains a reinforcement learning model based on the access selection algorithm from the access point. At this time, the 5G access device has not obtained the parameters of the trained reinforcement learning model, so it cannot directly use the reinforcement learning model to select the access point with the best delay and jitter performance and obtain network services. Instead, it first uses the random access algorithm to access an access point that can provide the network.
[0040] S2. Train the reinforcement learning model based on the access selection algorithm and transmit the trained reinforcement learning model to all 5G terminal devices that use the access node for network transmission. The specific steps are as follows:
[0041] S201. Obtain a reinforcement learning model and a training data set based on an access selection algorithm for each access node in a 5G access device, and initialize the reinforcement learning model and the training data set based on the access selection algorithm using a Markov model;
[0042] S202. Using access nodes to collect network data uploaded by 5G terminals, and accessing device parameters and device performance, the data training samples are obtained by combining them. Based on the data training samples, the parameters of the reinforcement learning model are trained using Q learning.
[0043] S203: Determine whether the time reaches the training cycle or whether there is a terminal requesting model parameters. If so, execute S204; otherwise, execute S202.
[0044] S204: using the parameters of the trained reinforcement learning model to optimize the objective function of the reinforcement learning model, training the reinforcement learning model based on the access selection algorithm;
[0045] S205: Transmit the trained reinforcement learning model downward to all 5G terminal devices that use the access node for network transmission.
[0046] In this embodiment, the reinforcement learning model parameters are requested from the access node. From the terminal's perspective, only the trained model parameters are obtained, but the steps of model training and data set collection are specifically included.
[0047] like Figure 2 As shown, the reinforcement learning model and training data set based on the access selection algorithm of each access node in the 5G access device are obtained and initialized using the Markov model. The parameters include<S,A,P,R,γ> ; Where S represents the state space, the set of all possible states, representing the different conditions or positions that the agent can be in the environment, and the agent will migrate between these states; A represents the action space, the set of all possible behaviors or actions that the agent can choose to perform; for each state s∈S, the agent will select an action a from the action space A to perform; P represents the state transition probability, the probability distribution of transitioning from a state s to the next state s′ after performing action a, and P(s′|s,a) is used to represent the probability of the agent reaching the new state s′ given the current state s and the action a taken; R represents the reward function, which defines a function that maps from state-action pairs to immediate rewards. After taking action a, the immediate reward obtained when the agent transitions from state s to the next state is recorded as R(s,a,s′) or abbreviated r; γ represents the discount factor, which is a value between 0 and 1, which determines the importance attached to future rewards. The closer the value of the discount factor γ is to 1, the longer the agent looks at future rewards and the more it pays attention to long-term interests; the closer the value is to 0, the more inclined the agent is to short-term gains.
[0048] In this embodiment, access nodes are used to collect network data uploaded by 5G terminals, and device parameters and performance are accessed to obtain new data training samples. Based on Q learning, the collected data is used to continuously train parameters. The optimal decision sequence of the Markov decision process is solved by the Bellman equation. The state value function V(s) is used to evaluate the quality of the current state. Each state value is determined not only by the current state but also by the subsequent states. Therefore, the state value function V(s) of the current state can be obtained by calculating the expected cumulative reward of the state.
[0049] The expression of the state value function V(s) is as follows:
[0050] V π (s)=E π [R t+1 +γV(s′)|s t =s]
[0051] Among them, π represents the algorithm strategy, E represents the expectation of the reward function, and R t+1 Represents the t+1th reward function; sort out E π Indicates the expected size of the reward that reinforcement learning can obtain when executing the access strategy of the algorithm;
[0052] According to the state value function V(s), the state action value function Q(S,A) is derived;
[0053] The expression of the state-action-value function Q(S,A) is as follows:
[0054] Q π (s,a)=E π [R t+1 +γQ(S t+1 ,A t+1 )|A t =a,s t =s]
[0055] Among them, A t+1 represents the t+1th action; the Q value can be calculated according to the state-action-value function Q(S,A), and learning is performed using the Q-table update process, where α represents the learning rate, β represents the reward decay coefficient, and the time difference method is used for updating; the update expression is as follows:
[0056] Q(s,a)←Q(s,a)+α[β+γmax a′ Q(s′,a′)-Q(s,a)]
[0057] Use L1 regularization parameter reverse derivation to optimize the reinforcement learning model parameters. The expression is as follows:
[0058]
[0059] Among them, L data Represents the data loss of the reinforcement learning model, which is usually the error between the predicted value of the reinforcement learning model and the true label, including mean square error (MSE) or cross-entropy loss. λ represents the regularization parameter, which is used to control the strength of the regularization term. |ω i | represents the absolute value of the weight of the reinforcement learning model.
[0060] In this embodiment, the reinforcement learning model based on the access selection algorithm is trained by continuously optimizing the objective function specially designed for reinforcement learning, that is, repeatedly optimizing the reinforcement learning model parameters. The objective function of the reinforcement learning model is expressed as follows:
[0061]
[0062] Among them, L(θ) represents the objective function of the reinforcement learning model, r t (θ) represents the reward ratio, ε represents the estimated factor of the advantage function, represents a positive number in the estimation factor of the advantage function, clip(·) represents contrastive language-image pre-training, θ represents the model parameters of reinforcement learning, θ old Represents the model parameters of the previous round of reinforcement learning, a t Indicates execution of an action, s t Represents the state, π θ Represents the algorithm strategy of the reinforcement learning model parameters, π θold The algorithmic strategy representing the model parameters of the previous round of reinforcement learning;
[0063] The trained reinforcement learning model is transmitted downward to all 5G terminal devices that use the access node for network transmission.
[0064] S3. Use 5G terminal equipment to collect network information of network access points. The specific steps are as follows:
[0065] In this embodiment, the 5G terminal access device collects network information such as the location, number, available bandwidth, background flow information, number of hops of the upper layer connection, basic delay and bandwidth requirements, and its own location and direction of the network access point; the network status information that needs to be collected is mainly divided into three parts: terminal node network information, access node status information, and flow status information: the network status is defined as a three-part combination expression as follows:
[0066] S network =[S end ,S access ,S flow ]
[0067] Among them, S network Indicates network status information, S end Represents the terminal node network information, S access Indicates access node status information, S flow Indicates flow status information;
[0068] Collect terminal node network information, such as the location of the 5G terminal device, mobile speed, device type, and the types and number of operators supported by the device; the expression of the terminal node network information is as follows:
[0069] S end =[POS end ,RATE,TYPE,OPERAT]
[0070] Among them, S end Indicates the terminal node network information, RATE indicates the mobile rate, TYPE indicates the device type, and OPERAT indicates the types and number of operators supported by the device;
[0071] Collect access node status information, such as the access node number, available bandwidth of the access node, number of background flows of the access node, size of background flows of the access node, and operator and location of the access node; the expression of the access node status information is as follows:
[0072] S access =[ID,BandWidth,NUM access ,SIZE access ,POS access ]
[0073] Among them, S flow Indicates access node status information, ID indicates the access node number, BandWidth indicates the available bandwidth of the access node, NUM access Indicates the number of background flows accessing the node, SIZE access Indicates the background flow size of the access node, POS access Indicates the operator and location of the access node;
[0074] Collect flow status information, such as the receiver of the flow, the number of routing hops, the delay required by the flow, the bandwidth and delay jitter requirements required by the flow, the bit size of the flow, and the duration of the flow; the expression of the flow status information is as follows:
[0075] S flow =[RECV,HOP,DELAY,JUTTER,BYTES,DURATION]
[0076] Among them, S flowIndicates flow status information, RECV indicates the receiver of the flow, HOP indicates the number of routing hops, DELAY indicates the delay required by the flow, JUTTER indicates the bandwidth and delay jitter requirements required by the flow, BYTES indicates the bit size of the flow, and indicates the duration of the flow.
[0077] S4. Quantify all collected network information and combine it into a network state matrix. Input the network state matrix into the trained reinforcement learning model and output the quantized value. Based on the output quantized value, select the network access point corresponding to the optimal delay and jitter performance and perform network switching.
[0078] In this embodiment, all collected network information is quantized and combined into a network state matrix, which is used as input to the reinforcement learning access selection model. The reinforcement learning model outputs a quantized value, and the corresponding access point is found to perform network switching.
[0079] S5. Connect the 5G terminal device to the network, obtain the network service of the network access point corresponding to the optimal delay and jitter performance, and complete the end system access delay and jitter optimization.
[0080] In this embodiment, the access network provides services for upper-layer applications and completes the end system access delay jitter optimization based on reinforcement learning.
[0081] In this embodiment, Figure 3 As shown, the access selection includes: a core network device 103, a first access network device 101 and a second access network device 102 connected to the core network device 103; the first access network device 101 is connected to a first terminal device 104 and a second terminal device 105 respectively; the second access network device 101 is connected to a third terminal device 106 and a fourth terminal device 107 respectively;
[0082] The core network device is connected to a plurality of access network devices, and each access network device is connected to a plurality of terminal devices; the access selection of the end system is realized; and the access delay jitter of the end system is optimized by using the present invention.
[0083] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0084] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
Claims
1. A method for optimizing end system access delay jitter based on reinforcement learning, characterized in that: The following steps are involved: S1. Use a random access algorithm to connect a 5G access device to a network access point, and use an access selection algorithm to obtain a reinforcement learning model based on the access selection algorithm. S2. Train a reinforcement learning model based on an access selection algorithm in the access node, and transmit the trained reinforcement learning model to all 5G terminal devices that use the access node for network transmission; S3. Use 5G terminal equipment to collect network information of network access points; S4. Quantify all collected network information and combine it into a network state matrix. Input the network state matrix into the trained reinforcement learning model, output a quantized value, and select the network access point with the best delay and jitter performance based on the output quantized value, and perform network switching. S5. Connect the 5G terminal device to the network, obtain the network service of the network access point corresponding to the optimal delay and jitter performance, and complete the end system access delay and jitter optimization.
2. The method for optimizing end system access delay jitter based on reinforcement learning according to claim 1, characterized in that: The specific steps of S1 are as follows: Using the random access algorithm, a network access point with the operator access node identifier is randomly generated in the 5G access device. According to the operator access node identifier, the 5G terminal device is used to connect to the 5G network of the corresponding operator, and the access selection algorithm is used according to the network access point to obtain a reinforcement learning model based on the access selection algorithm.
3. The method for optimizing end system access delay jitter based on reinforcement learning according to claim 1, characterized in that: The S2 comprises the following steps: S201. Obtain a reinforcement learning model and a training data set based on an access selection algorithm for each access node in a 5G access device, and initialize the reinforcement learning model and the training data set based on the access selection algorithm using a Markov model; S202. Using access nodes to collect network data uploaded by 5G terminals, and accessing device parameters and device performance, the data training samples are obtained by combining them. Based on the data training samples, the parameters of the reinforcement learning model are trained using Q learning. S203: Determine whether the time reaches the training cycle or whether there is a terminal requesting model parameters. If so, execute S204; otherwise, execute S202. S204: using the parameters of the trained reinforcement learning model to optimize the objective function of the reinforcement learning model, training the reinforcement learning model based on the access selection algorithm; S205: Transmit the trained reinforcement learning model downward to all 5G terminal devices that use the access node for network transmission.
4. The method for optimizing end system access delay jitter based on reinforcement learning according to claim 3, characterized in that: The objective function of the reinforcement learning model is expressed as follows: Among them, L(θ) represents the objective function of the reinforcement learning model, r t (θ) represents the reward ratio, ε represents the estimated factor of the advantage function, represents a positive number in the estimation factor of the advantage function, clip(·) represents contrastive language-image pre-training, θ represents the model parameters of reinforcement learning, θ old Represents the model parameters of the previous round of reinforcement learning, a t Indicates execution of an action, s t Represents the state, π θ Represents the algorithm strategy of the reinforcement learning model parameters, π θold The algorithm policy representing the model parameters of the previous round of reinforcement learning.
5. The method for optimizing end system access delay jitter based on reinforcement learning according to claim 1, characterized in that: The S3 includes the following steps: S301, dividing network status information into terminal node network information, access node status information and flow status information; S302. Collect terminal node network information using 5G terminal equipment; S303. Collect access node status information using a 5G terminal device; S304. Collect flow status information using 5G terminal equipment.
Citation Information
Patent Citations
A network load balancing system and a balancing method based on deep reinforcement learning
CN109039942A
Dynamic multi-channel access method in high-speed moving scene
CN110035478A