Terminal Aggregation Monitoring-Assisted Access Control Method in Ultra-Dense Heterogeneous Wireless Networks

Through the Gauss Markov model, the terminal trajectory is predicted and grid-based processing is performed, and combined with the Dueling-DQN network selection method, the network congestion problem caused by the sudden aggregation of terminals in super-dense heterogeneous wireless networks is solved, timely monitoring and effective response to aggregation situation is achieved, and terminal satisfaction and network load balancing are improved.

CN114286381BActive Publication Date: 2025-05-30CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111500041.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-05-30
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

In ultra-dense heterogeneous wireless networks, sudden congestion of terminals leads to network congestion, and the prior art is difficult to effectively avoid the occurrence of congestion, especially in hotspot events, the problems of low prediction accuracy and high time overhead are difficult to solve.

Method used

The Gauss Markov model is used to predict the trajectory of the mobile terminal, and the throughput change trend in each grid is calculated through grid processing, and the grid with the fastest growth in throughput is positioned to judge the aggregation situation. Then, based on the Dueling-DQN network selection method, terminal migration and network selection are performed to dynamically adjust access policies and alleviate network congestion.

Benefits of technology

Timely monitoring and effective response to terminal aggregation situation is realized, access strategies are dynamically adjusted to balance network load, terminal satisfaction is improved, and time overhead is reduced while ensuring prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114286381B_ABST
    Figure CN114286381B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for terminal aggregation monitoring-assisted access control in an ultra-dense heterogeneous wireless network, belonging to the field of mobile communications. Aiming at the problem of network congestion caused by the sudden aggregation of mobile terminals, a method for aggregation monitoring with an adaptively adjusted update period is proposed. Through grid division, the terminal movement trajectory is associated with the geographical area; then, by calculating whether the access point can accommodate the terminals within the grid, the monitoring of the aggregation location is realized; an adaptive adjustment mechanism for the monitoring period of the method is designed to achieve dynamic monitoring of the network congestion situation. Aiming at the problems of high network dynamics and complex access decisions, an access control method based on deep reinforcement learning is proposed. According to the results of aggregation monitoring, the access strategy is adjusted, which not only balances the load between networks, alleviates or avoids network congestion, but also improves the service experience of the terminals. This method can balance the network load while improving the total network throughput and providing high-quality data transmission for the terminals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of mobile communications, and particularly relates to a method for assisting access control by monitoring terminal aggregation in an ultra-dense heterogeneous wireless network. Background Art

[0002] With the large-scale popularization of intelligent devices, the demand for network traffic has increased sharply, posing a great challenge to network performance. In order to meet the transmission efficiency and quality of service requirements of terminals, an ultra-dense heterogeneous wireless network (UHWN) composed of macro base stations (MBS), small base stations (SBS), and other wireless technologies has emerged. With the continuous development of 5G network technology, it is bound to form a more complex and large-scale heterogeneous 5G network environment with the current wireless network environment, jointly providing wireless network services with different characteristics for mobile terminals. In these complex network environments, the process of a mobile terminal switching from one access point to an access point of a different technology is called vertical handover. How to switch a mobile terminal to a network that meets the requirements in such a large heterogeneous environment has become a hot research issue in this field. UHWN improves network capacity by deploying a large number of miniaturized devices in macro cells to meet the diverse quality of service requirements of terminals. At the same time, this ultra-dense deployment further increases the randomness of terminal access and resource requirements, and the network performance is easily affected by many human factors. If mobile terminals have aggregation activities within the network coverage area, due to the increase in the number of terminals and the increase in data demand, local overload will occur. The traffic volume brought by emergencies such as accidents, news events, and entertainment activities may even cause a rapid increase in local network load, resulting in network congestion, having a negative impact on the performance of the entire network, and unable to guarantee the terminal experience.

[0003] Regarding the problem of network congestion caused by terminal aggregation, the existing solutions can be mainly divided into two categories:

[0004] (1) Post-event mechanism: When network congestion occurs, the system balances the load between networks by controlling the handover of terminals; or uses a migration mechanism to move some terminals out. The literature [Wang S, Deng H, Xiong R, et al. A multi-objective model-based vertical handoff algorithm for heterogeneous wireless networks[J]. EURASIP Journal on Wireless Communications and Networking, 2021, 2021(1): 1-18.] constructs a multi-objective optimization model based on network and terminal states, and uses a multi-objective genetic method to solve it to obtain the terminal access strategy, which optimizes the terminal experience while reducing network congestion. The literature [Gao K, Xu C, Zhang P, et al. GCH-MV: Game-enhanced Compensation Handover Scheme for Multipath TCP in 6G Software Defined Vehicular Networks[J]. IEEE Transactions on Vehicular Technology, 2020.] models the multi-terminal network access control problem as a multi-person game and iteratively obtains the access strategy that can reach the Nash equilibrium to reduce network congestion and ensure terminal benefits. The literature [Feng B, Zhang C, Liu J, et al. D2D communications-assisted traffic offloading in integrated cellular-WiFi networks[J]. IEEE Internet of Things Journal, 2019, 6(5): 8670-8680.5] introduces a Device to Device (D2D) communication-assisted traffic offloading scheme to relieve the burden on the WiFi access point caused by offloading terminals. The literature [Feng B, Zhang C, Liu J, et al. D2D communications-assisted traffic offloading in integrated cellular-WiFi networks[J]. IEEE Internet of Things Journal, 2019, 6(5): 8670-8680.] proposes an automatic traffic offloading scheme based on big data and machine learning to balance the load between networks.The literature [Han S. Congestion-aware WiFi offload algorithm for 5G heterogeneous wireless networks [J]. Computer Communications, 2020, 164: 69-76.] divides priorities for terminals through terminal behavior analysis, and then formulates different offloading strategies based on multi-objective decision-making to relieve network congestion on the premise of meeting the satisfaction of migrating terminals.

[0005] (2) Pre-event mechanism: According to the prediction results of environmental changes at the next moment, adjust and plan the access strategy in advance. The literature [Yap K L, Chong Y W, Liu W. Enhanced handover mechanism using mobility prediction in wireless networks [J]. PloS one, 2020, 15(1): e0227982.] proposed a handover method assisted by mobility prediction. By combining mobility prediction with a multi-path protocol, some terminals are diverted to the WiFi network without degrading performance, improving the service experience of mobile devices. The literature [Y. Qi and H. Wang. Interference-aware user association under cell sleeping for heterogeneous cloud cellular networks [J]. IEEE Wireless Communications Letters, 2017, 6(2): 242-245.] predicts the traffic demand of terminals through the support vector regression method. Terminals can select networks according to the traffic demand to maximize system throughput and load balancing. The literature [Farooq H, Asghar A, Imran A. Mobility prediction based proactive dynamic network orchestration for load balancing with QoS constraint (OPERA) [J]. IEEE Transactions on Vehicular Technology, 2020, 69(3): 3370-3383.] predicts the future load of access points based on a semi-Markov model. Based on the prediction results, a joint genetic method and a pattern search method are used to solve the handover strategy of terminals, improving network capacity and terminal satisfaction.

[0006] If the above methods are used to solve the congestion problem caused by the sudden aggregation of mobile terminals in the present invention, there are some defects. For the ex-post mechanism method based on the handover strategy, the degree of network capacity expansion is limited and it is unable to effectively accommodate mobile terminals that are about to enter the network coverage area. For the ex-post mechanism method based on migration, there are risks and handover overheads during the process of terminal migration. If the handover fails, the services of the migrated terminals will be interrupted. The above two types of solutions are both strategies that are passively executed when the network has already experienced congestion and belong to ex-post mechanisms. Facing the sudden aggregation of mobile terminals in the present invention, the passive ex-post mechanism can only relieve congestion to a limited extent and cannot avoid the occurrence of congestion drops. Although the ex-ante mechanism can actively take control measures according to the prediction results, its method performance depends on the accuracy of the prediction. For the problem of sudden aggregation of mobile terminals in the present invention, the prediction methods adopted by such traditional ex-ante mechanisms have problems of large time overhead and low prediction accuracy. Therefore, such methods cannot detect terminal aggregation in time in this problem, resulting in the failure of subsequent active measures. The reason for congestion in the environment of the present invention is that when a hot event occurs, mobile terminals will go to the place where the event occurs and stay near the hot area, thus resulting in aggregation. Due to the two characteristics of the suddenness of hot events and the rapidity of mobile terminal aggregation, it is difficult for existing prediction methods to simultaneously meet the time overhead and prediction accuracy. At the same time, this problem involves the access control tasks of multiple base stations and multiple users within a period of time, and the network has high dynamicity. The characteristics of deep reinforcement learning to make decisions in chronological order and adapt to the dynamic changes of the environment can effectively solve such tasks. Summary of the Invention

[0007] The present invention aims to solve the above problems of the prior art. A method for terminal aggregation monitoring-assisted access control in an ultra-dense heterogeneous wireless network is proposed.

[0008] The technical solution adopted by the present invention is as follows. The method for terminal aggregation monitoring-assisted access control in an ultra-dense heterogeneous wireless network includes the following steps:

[0009] 101. Predict the trajectory of mobile terminals through a Gaussian Markov model, grid the trajectory of mobile terminals, calculate the change trend of the throughput of mobile terminals in each grid at the next moment, and the degree of load impact on the access points covering the grid; locate the grid with the fastest throughput growth. If its load impact on the access point exceeds the aggregation threshold, it is considered that aggregation is about to occur; finally, dynamically update the time interval according to the aggregation degree of the terminals.

[0010] 102. According to the result obtained in step 101, if the current network monitors an aggregation, it triggers the migration of edge fixed terminals, migrating them in descending order of the rate requirements of the fixed terminals until the number is less than the threshold or there are no fixed terminals available for migration; if there is no migration, it directly enters network selection, taking the transmission rate requirement, network load rate, and RSS as inputs to perform network selection based on Dueling-DQN.

[0011] The advantages and beneficial effects of the present invention are as follows:

[0012] 1. The present invention is directed to an ultra-dense heterogeneous wireless network environment composed of a wireless local area network and a cellular network. According to the predicted and gridified terminal prediction trajectories in step 101, it then calculates the grid with the largest throughput increase. Based on the calculation result, it is converted into the resource occupancy rate of the access point. If it exceeds the threshold, it is determined that an aggregation has occurred. After detecting the aggregation, an access control method is then proposed according to step 102, which balances the load while improving the terminal satisfaction; the invention designs an adaptive adjustment mechanism for the monitoring period of the algorithm, realizes the dynamic monitoring of the network congestion situation, achieves the purpose of monitoring the rapid aggregation of terminals, can be used to accurately grasp the movement trend of the terminals to assist subsequent access control steps; meanwhile, it takes into account both the time overhead and the prediction accuracy.

[0013] 2. Through the access control of the terminals in step 102, first, according to the result of the aggregation monitoring method, if there is a terminal aggregation that the access point cannot bear, it triggers the migration of edge fixed terminals, migrating them in descending order of the rate requirements of the fixed terminals until the number is less than the threshold or there are no fixed terminals available for migration; if there is no migration, it directly enters network selection. Subsequently, taking the transmission rate requirement, network load rate, and RSS as inputs, it performs network selection based on Dueling-DQN; step 102 can flexibly adjust the access strategy according to the result monitored in step 101, and can be used to handle the access control tasks of multiple base stations and multiple users in a high network dynamic environment; in order to achieve load balancing among the networks, relieve or avoid network congestion, and at the same time improve the service experience of the terminals. Description of the Drawings

[0014] Figure 1 is the flowchart of the monitoring and access control method of the present invention;

[0015] Figure 2 is the comparison of the number of correctly predicted terminals;

[0016] Figure 3 is the comparison of the time consumption of the monitoring method;

[0017] Figure 4 is the comparison of the delay satisfaction of different methods;

[0018] Figure 5Bandwidth satisfaction comparison for different methods;

[0019] Figure 6 Network load comparison for different methods;

[0020] Figure 7 Network delay increase comparison for different methods;

[0021] Figure 8 Packet loss rate increase comparison for different methods;

[0022] Figure 9 System throughput comparison for different methods. Specific implementation manner

[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0024] The present invention proposes an access control method for terminal aggregation monitoring based on deep reinforcement learning, which optimizes the congestion problem caused by sudden terminal aggregation in a ultra-dense heterogeneous wireless network. In the terminal aggregation monitoring stage, a mobility monitoring method with an adaptive update time interval is proposed, which can timely and quickly detect the aggregation behavior of terminals. At the same time, according to the severity of terminal aggregation, a trigger-based terminal migration strategy is used to increase the accommodation capacity of the network in the aggregation area. And according to the rate requirements of terminals and network parameters, an optimization model for maximizing the operator's balanced accommodation rate is established, and Dueling DQN is introduced for approximate solution to obtain the target network that meets the terminal requirements.

[0025] The access control method proposed by the present invention includes the following steps:

[0026] Step 1: Predict the trajectory of the mobile terminal through a Gaussian Markov model, and then represent the trajectory in a grid manner, converting the trajectory data points into the ID numbers of the grid. After that, calculate the change trend of the throughput of the mobile terminal in the next moment in each grid, as well as the degree of load impact on the access point covering the grid. Pay attention to the grid with the fastest throughput growth. If its load impact on the access point exceeds the threshold, it is considered that terminal aggregation that the access point cannot bear is about to occur. Finally, dynamically update the time interval of the algorithm according to the aggregation degree of the terminals.

[0027] Step 2: First, according to the result of the aggregation monitoring algorithm, if terminal aggregation that the access point cannot bear occurs, trigger the migration of edge fixed terminals, and migrate them in descending order of the rate requirements of the fixed terminals until it is less than the threshold or there are no fixed terminals to migrate; if not, directly enter network selection. Subsequently, use the transmission rate requirement, network load rate, and received signal strength (RSS) as inputs for network selection based on Dueling-DQN.

[0028] According to the above analysis, the present invention designs Figure 2 the method flowchart shown in the figure. The above steps are analyzed in more detail.

[0029] Furthermore, if the terminal enters a new network range, according to step two, the attribute parameters of the current network and the target network are obtained, specifically including the steps:

[0030] The received signal strength value of mobile terminal i received by base station j can be expressed as:

[0031]

[0032] where ρ is the transmission power of the base station's radio signal; η is the path loss factor; D is the distance from the terminal to the access point; μ is white noise with a variance of σ and a mean of 0.

[0033] The minimum bandwidth requirement of mobile terminal i can be expressed as:

[0034] Considering that the transmission rate is one of the key indicators determining the Quality of Service (QoS) and directly affects the service delay of the terminal, the minimum transmission rate requirement of the terminal is considered as the QoS requirement. Let the minimum transmission rate of terminal u be Then the minimum bandwidth requirement B of terminal u u is:

[0035]

[0036] where is the signal-to-noise ratio of terminal u at time t, expressed as N(k) = -117 + μ(k) dBm, and μ(k) follows a Gaussian distribution with parameters (0, 6).

[0037] The entire simulation space is divided into multiple hexagonal honeycomb grids with an inscribed circle diameter of 60 m; in the case of overlapping coverage of multiple SBSs or wireless access points, some grids will be covered by multiple access points, so the Thiessen polygon division is used to distinguish the boundaries of the overlapping coverage areas of multiple access points; finally, the basic processing unit of the prediction method is obtained: the grid; after dividing the grid, according to the terminal coordinate information, the original trajectory points of the terminal can be mapped to the corresponding grid, so the position of terminal u at time slot t can be expressed as: the corresponding number g of the grid where the terminal is located at time slot t.

[0038] Furthermore, three indicators are defined: grid throughput, grid throughput growth rate, and grid load contribution degree. The definition process is as follows:

[0039] The grid throughput is equal to the sum of the throughputs of the terminals located within grid g at time t, and the grid throughput C g,t is expressed as:

[0040]

[0041] where, is the minimum transmission rate requirement of terminal j; C g,t is used to measure the degree of demand of the terminals within the grid for the access point bandwidth. The larger C g,t is, the more bandwidth is required.

[0042] The grid throughput growth rate is equal to the throughput growth rate of grid g within the most recent update period ΔT. The throughput growth rate C' of grid g within the most recent update period ΔT g is calculated as follows:

[0043]

[0044] C g,t+ΔT represents the throughput of the grid at the next moment; C' g can represent the change speed of the throughput of the grid within ΔT. The larger the throughput growth rate is, the greater the possibility that this grid may potentially become the location of a hotspot event.

[0045] The load contribution degree is equal to the bandwidth demand B of grid g g,t and the available bandwidth of the covered access point of this grid The load contribution degree of grid g located in the overlapping area of n access points is calculated as follows:

[0046]

[0047] where, is the area ratio of grid g within the coverage of access point to the unit grid, is the coverage area of grid g in access point , S g is the area of grid g; is the remaining available bandwidth of access point ; The bandwidth demand of grid g

[0048] During the process of terminal mobility monitoring, the throughput growth rate C' g of the grid with the fastest growth will be obtained in each time slot, and this grid is regarded as the grid where terminal aggregation may occur. According to its load contribution degree the update time ΔT of the next cycle is adjusted, and the policy function is given as:

[0049]

[0050] Among them, T max and T min are the maximum time interval and the minimum time interval for executing method updates respectively; S(·) is the Sigmoid function, is the load contribution degree of grid g; C′ g is the throughput growth rate of grid g.

[0051] The located terminal aggregation grid information is input into the subsequent link to determine whether it is necessary to perform terminal migration on the network covering the grid;

[0052] Suppose there are n access points covering the terminal aggregation grid g, for any access point covering the grid g Therefore, the condition for terminal migration can be expressed as:

[0053]

[0054] Among them, represents the total bandwidth of the k-th access point; α represents the aggregation threshold, taking 0.7. According to the above formula, a trigger-based edge terminal migration strategy is proposed, that is: when the degree of terminal aggregation exceeds the capacity of the network, edge terminals in the network should be migrated; in addition, to avoid unnecessary handovers of edge terminals, this strategy only considers migrating edge terminals from overloaded networks to idle networks adjacent to them and not covering the aggregation grid, so as to minimize the risk during the process of edge terminal migration and more effectively optimize the network.

[0055] First, calculate the load change after migration. Suppose terminal u moves out of the original serving network After that, the bandwidth occupancy of becomes: Among them, is the resource occupied by terminal u in network : c j is the minimum transmission rate requirement of terminal j, represents the channel quality from network to terminal u;

[0056] Since the terminal will have different signal qualities in the target network m, and services are provided according to the minimum rate requirement after migration, the bandwidth required by the terminal will be different; the predicted bandwidth required for terminal j to move into the target network m becomes:

[0057]

[0058] Resource Occupation of Target Network m after Terminal j is Migrated In: When performing migration, the background obtains the information of fixed terminals located at the edge of the high-load network through regularly collected terminal information; first, calculate the load after migrating the edge terminals with high rate requirements. If the load rate of the target network still exceeds α, then gradually migrate the edge terminals with lower rate requirements until sufficient resources can be provided or there are no terminals to migrate.

[0059] The present invention also designs an objective function for maximizing throughput, using to define the association relationship between terminal u and the network such that indicates that terminal u accesses the network at time slot t otherwise it indicates non-access to the network The data volume sent by the terminal is represented as b u , then maximizing throughput can be expressed as:

[0060]

[0061] The network occupied by the already-accessed services The bandwidth, taking the ratio of the occupied bandwidth of the network to the bandwidth as the load rate

[0062]

[0063] where is the total resource number of the network Taking the minimum variance of the load rates of each network as the optimization objective can ensure load balancing among networks. The objective function is as follows:

[0064]

[0065] When a service request accesses, a terminal can only access one network at the same time, expressed as the following constraint condition:

[0066]

[0067] where Φ 1 , Φ 2 and Φ 3 represent the SBS set, the wireless local area network access point set, and the MBS set respectively; the total bandwidth occupied by terminal u accessing the network shall not exceed its allocable upper limit expressed as the following constraint condition:

[0068]

[0069] And when the terminal transmits data, it needs to meet the minimum transmission rate requirement, which is expressed as the following constraint condition:

[0070]

[0071] The operator's balanced accommodation rate represents the ability of the network operator to use the existing access devices to evenly distribute terminals among networks as much as possible to provide communication services with quality assurance. Therefore, the following objective function for maximizing the operator's balanced accommodation rate is constructed. Its meaning is that the network operator hopes to control the terminal handover behavior to maximize the throughput under the premise of load balancing, which is expressed as follows:

[0072]

[0073] Transform the maximization problem into a Markov decision (MDP) process. Utilize the characteristic of reinforcement learning to make decisions in chronological order, and use the Dueling-DQN method to solve the approximate optimal solution of this problem. Dueling-DQN divides the Q function into two parts, namely the state value function and the advantage function. The state-action value function Q π (s t ,a t ) represents the expected return value when making a network selection behavior at according to the policy π in the network state s t . The state value function V(s t ) describes the value of the current state st, which is the expected value of all action values generated by the policy π in this state. Then the difference between the two represents the value of choosing the action a t in the state s t , that is, the advantage function A π (s t ,a t ) is defined as:

[0074] A π (s t ,a t )=Q π (s t ,a t )-V π (s t ) (15)

[0075] There are two data streams in the competing network. One stream outputs the state value V π (s t ), and the other stream outputs the action advantage A π (s t ,a t ). The output of the deep Q network adopting the competing network structure is:

[0076] Q(s t ,at ; θ, ζ, ξ) = V(s t ; θ, ξ) + A π (s t , a t ; θ, ζ) (16)

[0077] Among them, θ represents the network neuron parameters for feature processing of the input layer; ζ and ξ are the parameters of the state value function and the advantage function respectively. Since the network directly outputs the parameterized estimated value of the true Q function, there will be a problem of identifiability in Q(s t , a t ; θ, ζ, ξ). To identify the respective roles of V(s t ; θ, ξ) and A π (s t , a t ; θ, ζ) in the final output, the average value of the advantage function estimator at the selected action is forced to be 0. The modified Q value is expressed as:

[0078]

[0079] Among them, Q(s t , a t ; θ, ζ, ξ) is the output value of the competing Q network; V(s t ; θ, ξ) is the state value function of the current state; A π (s t , a t ; θ, ζ) represents the action advantage; A represents the advantage function estimator;

[0080] The Q network updates the parameters by minimizing the loss function L:

[0081] L = ∑[(y t - Q(s t , a t ; θ, ζ, ξ)) 2 (18)

[0082] Among them, Q(s t , a t ; θ, ζ, ξ) is the output value of the competing Q network;

[0083] represents the estimated Q value at time t;

[0084] r(s t , a t ) represents the reward brought by the action a t in the state s t , and φ represents the discount factor.

[0085] The competitive deep Q-network learns through offline training: The central controller first collects the required information and then conducts offline training. When used online after training, the network resource information and terminal movement information processed by the previous method will be passed into this method. Next, the central controller converts all this information into a system state, which is fed into the competitive deep Q-network method to feedback the optimal action policy π = argmaxQ(x,a) at the current decision time slot t. After obtaining the action, the central agent will send the policy to the network to notify them to adjust the mobile terminal handover policy. The present invention makes the following definitions for the three elements based on Deep Reinforcement Learning (DRL):

[0086] State space That is, the agent needs to observe the terminal states and network states of all terminals within time slot t. Among them, is the minimum rate requirement of all terminals; b tu is the data volume to be transmitted by each terminal at time t; is the RSS of each network at grid g at time t; is the number of available resource blocks of each network at time t;

[0087] The action space of the agent

[0088] The reward function is expressed as:

[0089]

[0090] Among them, W u is the penalty for violating the terminal QoS requirement.

[0091] To verify the present invention, we conduct a simulation experiment on the MATLAB platform and set the following simulation scenario: A 1.2km * 1.2km rectangular network simulation environment composed of three wireless network technologies: wireless local area network, 5G microcell, and 5G macrocell. The simulation scenario is as Figure 1 shown.

[0092] During the simulation process, there are two types of terminals in the scenario: pedestrian terminals and vehicle mobile terminals. The initial positions of both types of terminals are randomly distributed within the simulation area. The number of pedestrian terminals is 64, and the positions are fixed. The number of vehicle mobile terminals will be described in the specific simulation. The initial moving speed is 3m / s, and they move with variable speed and direction within the simulation scenario. The minimum number requirement of each terminal To further highlight the superiority of the present invention, the present invention compares the performance of the access control algorithm based on deep reinforcement learning (Proposed), the access control algorithm based on neural network (Neural Network-based, NN-based), and the access control algorithm based on multi-objective decision-making (Multi-Attribute Decision Making-based, MADM-based) in terms of load balancing degree, blocking rate, throughput, and the number of access terminals.

[0093] Figure 3 The number of correctly predicted terminals varies with the number of mobile terminals in the simulation scenario, where the number of terminals correctly predicted by the method of the present invention is higher than the other two methods, the method based on Kalman filter ranks second, and the number of the fixed time interval method is the lowest.

[0094] As Figure 4 shown, it shows the variation of time consumption with the number of mobile terminals, where the time overhead of the mobility prediction method based on Kalman filter is the largest, the time overhead of the method proposed by the present invention is similar to that of the fixed time interval method, and the time consumption of the fixed time interval method is slightly lower than that of this method. This is because this method can dynamically adjust the update period according to the terminal aggregation situation in the simulation scenario. Therefore, this method can achieve the purpose of ensuring the prediction accuracy while the time consumption does not increase significantly compared with the traditional fixed time interval method.

[0095] Figure 5 and Figure 6 are the network load rates of the three methods when the number of mobile terminals is 20 and 40 respectively. It can be seen that the load of the MADM-based method is mainly concentrated in the 5G microcell and wireless local area network, under the NN-based method, the terminals are mainly concentrated in the 5G microcell and 5G macrocell, while the network load of the method proposed by the present invention is relatively balanced. Figure 7 When the number of terminals is 40, the load of the NN-based method and the MADM-based method does not increase significantly. This is because after an emergency occurs, the terminals gather in the hot spot area, and the above methods fail to effectively balance the load between networks, resulting in network congestion in the hot spot area covered, and a large number of terminals gathered here cannot access, while the surrounding networks are relatively idle, so that the overall network load rate does not increase significantly compared with when the number of mobile terminals is 20.

[0096] Figure 7The blocking rates of the three methods vary with the number of mobile terminals in the simulation scenario. It can be seen that when the number of mobile terminals is 20, the NN-based method starts to be blocked. The MADM-based method starts to be blocked slightly later than the NN-based method. The method proposed in the present invention starts to be blocked when the number of mobile terminals is 35. This is because this method successfully utilizes the results of the terminal aggregation detection method, takes corresponding measures after warning of terminal aggregation, migrates the edge terminals in the network covering the aggregation area in advance, increases the network capacity, and fully considers the load during handover network selection, greatly improving the network resource utilization rate.

[0097] Figure 8 The throughputs of the three methods vary with the number of mobile terminals in the simulation scenario. As the number of mobile terminals increases, the throughputs of these three methods all increase. Starting from when the number of mobile terminals is 20, the rising trend of the throughput of the comparative method gradually flattens out. Generally speaking, the method of the present invention can achieve the highest throughput, followed by the MADM-based method, and the NN-based method is the lowest. This is because when the method of the present invention accesses, it fully considers the load, so that while the handover reduces the blocking rate, it can ensure that more terminals are accessed and provide them with data transmission with quality assurance.

[0098] Figure 9 The number of mobile terminals accessed by the three methods varies with the number of mobile terminals in the simulation scenario. When the number of terminals in the scenario is small, the network resources are sufficient, so each terminal can be accessed. However, as the number of mobile terminals in the scenario increases, the number of accessed terminals all increases to varying degrees. Among them, the increase in the number of accessed terminals by the NN-based method is the lowest, followed by the MADM-based method, and the method proposed in the present invention has the largest number of accessed terminals. This is because the method of the present invention can sense terminal aggregation in advance, make access control in advance for terminal aggregation, expand the network capacity of the aggregation area, and meet the access needs of more terminals.

[0099] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0100] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined in the present invention, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0101] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0102] The above embodiments should be understood as being only for illustrative purposes of the present invention and not for limiting the scope of protection of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A method for terminal aggregation monitoring-assisted access control in a ultra-dense heterogeneous wireless network, characterized in that, it includes the following steps:

101. Grid the trajectories of mobile terminals, calculate the change trend of the throughput of mobile terminals in each grid at the next moment, and the grid load contribution degree; locate the grid with the fastest throughput growth. If the grid load contribution degree exceeds the aggregation threshold, it is considered that aggregation is about to occur; Finally, dynamically update the time interval according to the aggregation degree of the terminals; The trajectory of the grid-based mobile terminal specifically includes: dividing the entire simulation space into multiple hexagonal honeycomb grids. In the case of overlapping coverage of multiple SBSs or APs, some grids will be covered by multiple access points. The Thiessen polygon division is used to distinguish the boundaries of the overlapping coverage areas of multiple access points, and the set is used to represent the set of all grids in the simulation area, and finally the divided grids are obtained; The definitions of the grid throughput, grid throughput growth rate, and grid load contribution degree are as follows: The grid throughput is equal to the sum of the throughputs of the terminals located within grid g at time t, and the grid throughput C g,t is expressed as: Among them, is the minimum transmission rate requirement of terminal j; n represents n access points; The grid throughput growth rate is equal to the throughput growth rate of grid g in the most recent update period ΔT. Then, the throughput growth rate C of grid g in the most recent update period ΔT is g calculated as follows: Among them, C g,t+ΔT represents the throughput of the grid at the next moment, and C g ' represents the change rate of the throughput of the grid within ΔT; The grid load contribution is equal to the bandwidth demand B of grid g g,t and the available bandwidth of the coverage access points of this grid The ratio, the load contribution l of grid g located in the overlapping area of n access points is calculated as follows: g,t is calculated as follows: Among them, is the area ratio of the grid g within the coverage range of the access point to the unit grid; is the remaining available bandwidth of the access point ; B g,t is the bandwidth demand of the grid g; During the process of terminal mobility monitoring, the throughput growth rate C is obtained for each time slot g The grid with the fastest growth g is regarded as the grid where terminal aggregation may occur. According to its load contribution degree l g,t , the update time ΔT of the next cycle is adjusted, and the policy function is given as follows: where, T max and T min are the maximum time interval and the minimum time interval for executing the method update respectively; S(·) is the Sigmoid function, l g,t is the load contribution degree of the grid g; C g ′ is the throughput growth rate of the grid g; 102. According to the result obtained in step 101, if the current network monitors that aggregation occurs, trigger the migration of edge fixed terminals, and migrate them in descending order of the fixed terminal rate requirements until it is less than the threshold or there are no fixed terminals to migrate; if not, directly enter network selection, and use the transmission rate requirement, network load rate, and RSS as inputs to perform network selection based on Dueling-DQN; The trigger for edge fixed terminal migration, the strategy is: The condition for terminal migration is expressed as Among them, represents the total bandwidth of the k-th access point; α represents the aggregation threshold; When the aggregation degree of the terminals exceeds the capacity of the network, the edge terminals in the network should be migrated, and the edge terminals should be migrated from the overloaded network to the idle network adjacent to it and not covering the aggregated grid to avoid the risk of the edge terminal migration process as much as possible; The steps for triggering the migration of edge fixed terminals include: Calculate the load change after migration: Assume that the terminal u moves out of the original service network After that, the bandwidth occupancy becomes: Among them, is the resource occupied by the terminal u in the network : c j is the minimum transmission rate requirement of the terminal j, represents the channel quality from the network to the terminal u; The predicted bandwidth required for terminal j to migrate to target network m becomes: Resource Occupation of Target Network m after Terminal j is Migrated In: When performing migration, the background obtains the information of fixed terminals located at the edge of the high-load network through regularly collected terminal information; first calculates the load after migrating the edge terminals with high rate requirements. If the load rate of the target network still exceeds α, then gradually migrate the edge terminals with lower rate requirements until sufficient resources can be provided or there are no terminals to migrate.

2. The method for terminal aggregation monitoring-assisted access control in a ultra-dense heterogeneous wireless network according to claim 1, characterized in that: When a terminal enters a new network range, according to the attribute parameters of the current network and the target network obtained in step 102, specifically including: The received signal strength value of mobile terminal i received by base station j is expressed as: where ρ is the transmission power of the base station radio signal; η is the path loss factor; D is the distance from the terminal to the access point; μ is white noise with a variance of σ and a mean of 0; Let the minimum transmission rate of terminal u be Then the minimum bandwidth requirement B of terminal u u is as follows: Among them, is the signal-to-noise ratio of terminal u at time t; The position of terminal u at time slot t is: the corresponding number g of the grid where the terminal is located at time slot t.

3. The method for terminal aggregation monitoring-assisted access control in a ultra-dense heterogeneous wireless network according to claim 1, characterized in that: Design the objective function to maximize throughput, using Define the association relationship between terminal u and the network as indicating that terminal u accesses the network in time slot t indicating non-access to the network The amount of data sent by the terminal is denoted as b u , then maximizing throughput is expressed as: Network occupied by the services already connected in of the bandwidth. The ratio of the occupied network bandwidth to the total bandwidth is used as the load rate Among them, is the total number of resources of the network ; taking the minimum variance of each network load rate as the optimization goal to ensure load balancing among networks, the objective function is shown as follows: When a service requests access, a terminal can only access one network at the same time, which is expressed as the following constraint condition: m represents the target network; where, Φ 1 , Φ 2 and Φ 3 represent the SBS set, the wireless local area network access point set, and the MBS set respectively; the total bandwidth occupied by the terminal u of the access network shall not exceed its allocable upper limit which is expressed as the following constraint: B u represents the minimum bandwidth requirement of terminal u; And when the terminal transmits data, it needs to meet the minimum transmission rate requirement, which is expressed as the following constraint condition: Indicates the minimum transmission rate of terminal u; Construct an objective function to maximize the operator's balanced accommodation rate:

4. The method for terminal aggregation monitoring-assisted access control in a ultra-dense heterogeneous wireless network according to claim 3, characterized in that: The Dueling-DQN is adopted to solve the problems of maximizing throughput and maximizing the equilibrium accommodation rate of the operator. The Dueling-DQN divides the Q function into two parts, namely the state value function and the advantage function; the state-action value function Q π (s t ,a t ) represents the expected return value when the network selection behavior a t is made by the policy π in the network state s t . The state value function V(s t ) describes the value of the current state s t , which is the expected value of all action values generated by the policy π in this state. Then the difference between the two represents the value of selecting the action a t in the state s t , that is, the advantage function A π (s t ,a t ) is defined as: A π (s t ,a t )=Q π (s t ,a t )-V π (s t )(15) There are two data streams in the competitive network. One stream outputs the state value V π (s t ), and the other stream outputs the action advantage A π (s t , a t ). The modified Q-value is expressed as: where Q(s t , a t ; θ, ζ, ξ) is the output value of the competitive Q-network; V(s t ; θ, ξ) is the state value function of the current state; A π (s t , a t ; θ, ζ) represents the action advantage; A represents the advantage function estimator; The Q network updates the parameters by minimizing the loss function L: L = ∑[(y t - Q(s t , a t ; θ, ζ, ξ)) 2 (18) Among them, Q(s t , a t ; θ, ζ, ξ) is the output value of the competitive Q-network; Denote the estimated Q value at time t, r(s t ,a t ) represents the reward brought by action a t in state s t , and φ represents the discount factor.

5. The method for terminal aggregation monitoring-assisted access control in a ultra-dense heterogeneous wireless network according to claim 4, characterized in that: The Q network learns through offline training and is used online after training. The central controller converts the information into the system state and sends it to the Q network, and feedbacks the optimal action strategy π = argmaxQ(x,a) at the current decision time slot t; after obtaining the action, the central agent will send the strategy to the network and notify them to adjust the mobile terminal handover strategy.

Citation Information

Patent Citations

  • Method for optimizing VM migration between MEC nodes in ultra-dense network

    CN107919986A

  • Network selection method based on improved deep Q learning

    CN112367683A