Ultra-Dense Heterogeneous Wireless Network Access Control Method, System, Device and Medium
By constructing a user satisfaction model based on psychological curve function and deep reinforcement learning, the problems of user personalized needs and network status changes in super-intensive heterogeneous wireless networks are solved, and the best choice of network access is achieved, improving user experience and network efficiency.
Patent Information
- Application Number
- CN202510458267.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing ultra-intensive heterogeneous wireless network access selection algorithm fails to fully consider user personalized needs and network status changes, resulting in the inability to guarantee the ping-pong effect and service quality.
Build a user satisfaction model based on psychological curve function and deep reinforcement learning, combine user expected returns and current network benefits, and train deep reinforcement learning neural networks through evolutionary strategy algorithms to achieve the best choice of network access.
Dynamically adapt to network status and user needs, improve user satisfaction, reduce the ping-pong effect, and improve the intelligence and efficiency of network selection.
Smart Images

Figure CN119997151B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless network access, and in particular to an ultra-dense heterogeneous wireless network access control method, system, device and medium. Background Art
[0002] A wireless communication system with overlapping signal coverage composed of wireless networks with different architectures such as mobile cellular networks, wireless local area networks, wireless metropolitan area networks, and satellite communication networks is called an ultra-dense heterogeneous wireless network. The coexistence and integration of wireless networks with various different access technologies have become the development trend of the next-generation mobile Internet. In the ultra-dense heterogeneous wireless network environment, how to connect users to the most suitable network has always been one of the key research points and difficulties.
[0003] Traditional ultra-dense heterogeneous wireless network access selection algorithms mainly use the received signal strength (RSS) as the judgment parameter for network selection, and mobile users choose to access the network with the highest RSS. Although the access selection algorithm based on RSS has low complexity and is easy to implement, it often causes a relatively serious ping-pong effect. Secondly, some access selection algorithms use the network load as the judgment basis for network selection, and connect users to the network with the lowest load to achieve the purpose of load balancing. Although such algorithms improve the resource utilization rate of heterogeneous wireless networks, these algorithms do not consider the needs of user services, and users may be connected to a network with poor quality. Therefore, the quality of service of services and the quality of user experience cannot be effectively guaranteed. In addition, there are also designs of access selection algorithms using the multiple attribute decision making (MADM) theory.
[0004] However, by designing the access selection algorithm of the heterogeneous wireless network through the above methods, only a network with the best comprehensive performance is selected for the user among all candidate networks, without fully considering various factors such as network performance, user service requirements, and user preferences. Therefore, users cannot be connected to the most suitable network. Summary of the Invention
[0005] The purpose of the present invention is to provide an ultra-dense heterogeneous wireless network access control method, system, device and medium, which can better meet the personalized needs of users, construct a more objective user satisfaction model, quantify user satisfaction, enable the network selection strategy to dynamically adapt to changes in network status and user needs, and achieve the best selection and access of the network.
[0006] To achieve the above purpose, an embodiment of the present invention provides an ultra-dense heterogeneous wireless network access control method, including:
[0007] Construct a network state space based on the attribute parameters of the candidate networks;
[0008] Construct an action space based on the candidate networks that can be selected for access;
[0009] Based on the psychological curve function, combine the expected benefits of the user and the current benefits of the candidate networks to construct a user satisfaction model, and use the user satisfaction model to construct a reward function;
[0010] According to the network state space, the action space, and the reward function, combine with a backpropagation neural network to construct a deep reinforcement learning neural network model;
[0011] Train the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model;
[0012] Adopt the ultra-dense heterogeneous wireless network access control model to execute the handover strategy of the ultra-dense heterogeneous wireless network.
[0013] As an improvement to the above solution, the attribute parameters of the candidate networks include user personalized demand parameters and network state information parameters;
[0014] The user personalized demand parameters include the price, bandwidth, and received signal strength of the candidate networks;
[0015] The network state information parameters include the load rate, bit error rate, and blocking rate of the candidate networks.
[0016] As an improvement to the above solution, the construction of the user satisfaction model based on the psychological curve function according to the expected benefits of the user and the current benefits of the candidate networks includes:
[0017] Normalize the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter, and calculate the current benefits of the candidate networks according to the normalized value of each attribute parameter;
[0018] By analyzing the historical network access data of the user, calculate the satisfaction of the user with each attribute parameter to obtain the expected benefits of the user;
[0019] Based on the psychological curve function, construct a user satisfaction model according to the expected benefits of the user and the current benefits of the candidate networks.
[0020] As an improvement to the above solution, the normalization of the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter, and the calculation of the current benefits of the candidate networks according to the normalized value of each attribute parameter include:
[0021] Normalize the attribute parameters of the candidate networks respectively to obtain the normalized values of each attribute parameter;
[0022] Classify the attribute parameters of the candidate networks into three categories: cost, system performance, and stability. Among them, the cost - type attribute parameters include the price of the candidate network, the system - performance - type attribute parameters include the bandwidth and received signal strength of the candidate network, and the stability - type attribute parameters include the load rate, bit - error rate, and blocking rate of the candidate network;
[0023] Use the fuzzy analytic hierarchy process to calculate the weight values of each parameter in the system - performance - type attribute parameters and stability - type attribute parameters respectively;
[0024] Define the normalized value of the price as the current cost benefit of the candidate network, calculate the current system - performance benefit of the candidate network according to the normalized values and weight values of the bandwidth and received signal strength respectively, and calculate the current stability benefit of the candidate network according to the normalized values and weight values of the load rate, bit - error rate, and blocking rate respectively.
[0025] As an improvement of the above - mentioned scheme, the psychological curve function is:
[0026] ;
[0027] where, is the user satisfaction; is the expected benefit of the user; is the current benefit of the candidate network;
[0028] The user satisfaction model is:
[0029] ;
[0030] ;
[0031] ;
[0032] where, represents selecting the th network, represents the state of the network, represents the action taken; is the cost satisfaction model, represents the expected cost benefit after taking action in state , represents the current cost benefit of the th network after taking action in state ; is the system - performance satisfaction model, represents in state Take actions below The expected benefit of system performance after Indicates the th network in state Take actions below The current benefit of system performance after Is the stability satisfaction model, Indicates in state Take actions below The expected benefit of stability after Indicates the th network in state Take actions below The current benefit of stability after
[0033] As an improvement to the above solution, the reward function is:
[0034] ;
[0035] Where , And Are the weight values of the cost - type attribute parameter, system - performance attribute parameter, and stability - attribute parameter respectively.
[0036] As an improvement to the above solution, training the deep reinforcement learning neural network model through the evolutionary strategy algorithm to obtain a hyper - dense heterogeneous wireless network access control model, including:
[0037] Initialize the parameters of the deep reinforcement learning neural network model through the evolutionary strategy algorithm;
[0038] Construct a training data set using the historical handover decision data of the hyper - dense heterogeneous wireless network;
[0039] Set a main network and a target network for the deep reinforcement learning neural network model;
[0040] Use the training data set to train the main network and the target network respectively, update the parameters of the main network, and stop training when the preset training conditions are met to obtain a hyper - dense heterogeneous wireless network access control model.
[0041] To achieve the above objectives, an embodiment of the present invention also provides a hyper - dense heterogeneous wireless network access control system, including:
[0042] A first model - element construction module, configured to construct a network state space according to the attribute parameters of candidate networks;
[0043] A second model - element construction module, configured to construct an action space according to candidate networks that can be selected for access;
[0044] A third model element construction module, configured to construct a user satisfaction model based on a psychological curve function, in combination with the expected benefit of a user and the current benefit of a candidate network, and construct a reward function by using the user satisfaction model;
[0045] A deep reinforcement learning neural network model construction module, configured to construct a deep reinforcement learning neural network model by combining a backpropagation neural network according to the network state space, the action space, and the reward function;
[0046] A deep reinforcement learning neural network model training module, configured to train the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain a hyper-dense heterogeneous wireless network access control model;
[0047] A network access control module, configured to execute a handover strategy of a hyper-dense heterogeneous wireless network by using the hyper-dense heterogeneous wireless network access control model.
[0048] To achieve the above object, an embodiment of the present invention further provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the hyper-dense heterogeneous wireless network access control method described in any one of the above is implemented.
[0049] To achieve the above object, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. When the computer program runs, it controls a device where the computer-readable storage medium is located to execute the hyper-dense heterogeneous wireless network access control method described in any one of the above.
[0050] Compared with the prior art, an access control method, system, device and medium for a ultra-dense heterogeneous wireless network provided by an embodiment of the present invention first constructs a network state space, an action space and a reward function. Among them, based on a psychological curve function, combining the expected benefit of a user and the current benefit of a candidate network, a user satisfaction model is constructed, and the reward function is constructed by using the user satisfaction model. Then, according to the network state space, the action space and the reward function, a deep reinforcement learning neural network model is constructed by combining with a backpropagation neural network. Then, the deep reinforcement learning neural network model is trained by an evolutionary strategy algorithm to obtain an access control model for a ultra-dense heterogeneous wireless network. Finally, the access control model for a ultra-dense heterogeneous wireless network is used to execute a handover strategy for a ultra-dense heterogeneous wireless network. The present invention fully considers the personalized needs of users, comprehensively considers the attribute parameters on the user side and the network side, constructs a more objective user satisfaction model for quantifying user satisfaction, enables the network selection strategy to dynamically adapt to changes in network states and user needs, and realizes the best selection and access of the network. At the same time, the present invention introduces deep reinforcement learning into the access selection problem of a heterogeneous wireless network, models the network selection and access problem as a reinforcement learning model, can learn the optimal strategy through interaction with the environment, and intelligently makes access selection decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the present invention, the drawings to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 is a flowchart of an access control method for a ultra-dense heterogeneous wireless network provided by an embodiment of the present invention;
[0053] Figure 2 is a hierarchical analysis structure diagram provided by an embodiment of the present invention;
[0054] Figure 3 is a schematic diagram of a main network structure provided by an embodiment of the present invention;
[0055] Figure 4 is a structural block diagram of a deep reinforcement learning neural network model provided by an embodiment of the present invention;
[0056] Figure 5 is another flowchart of an access control method for a ultra-dense heterogeneous wireless network provided by an embodiment of the present invention;
[0057] Figure 6It is a block diagram of an ultra-dense heterogeneous wireless network access control system provided by an embodiment of the present invention;
[0058] Figure 7 It is a block diagram of a terminal device provided by an embodiment of the present invention. Specific embodiments
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] See Figure 1 , Figure 1 It is a flowchart of an ultra-dense heterogeneous wireless network access control method provided by an embodiment of the present invention. The ultra-dense heterogeneous wireless network access control method includes steps S1 to S6:
[0061] S1. Construct a network state space according to the attribute parameters of the candidate network;
[0062] It can be understood that in the embodiments of the present invention, the elements of the deep reinforcement learning model in the ultra-dense heterogeneous wireless network access environment are first constructed. The construction of the elements of the deep reinforcement learning model mainly includes three parts: the definition of the state space, the definition of the action space, and the design of the reward function.
[0063] It should be noted that in deep reinforcement learning, the state space is the set of all possible states that an agent can be in the environment. In the embodiments of the present invention, the state is a complete description of the ultra-dense heterogeneous wireless network access environment, including all relevant information required for the agent to make a decision.
[0064] In an alternative embodiment, the attribute parameters of the candidate network include user personalized demand parameters and network state information parameters;
[0065] The user personalized demand parameters include the price, bandwidth, and received signal strength of the candidate network;
[0066] The network state information parameters include the load rate, bit error rate, and blocking rate of the candidate network.
[0067] It should be noted that in specific implementation, the state space includes not only the attribute parameters of the candidate network, but also the service type currently initiated by the user and the currently selected network.
[0068] It should be noted that among the attribute parameters of the candidate networks, price belongs to the cost attribute, bandwidth and received signal strength belong to the system performance attributes, and bit error rate, blocking rate, and load rate belong to the stability attributes. These parameters cover the key aspects of network performance and can meet the different focuses of different users' network requirements.
[0069] S2. Construct an action space based on the candidate networks that can be selected for access;
[0070] It should be noted that in deep reinforcement learning, the action space is the set of all possible actions that an agent can take in a given state. An action is the way an agent changes the state of the environment, and the action space defines the scope of these behaviors. In the embodiments of the present invention, assuming that the user is under multiple network coverage areas and can only select one candidate network for access at the decision-making moment, the action space is defined as the set of candidate networks that the user can select at the handover decision-making moment.
[0071] S3. Based on the psychological curve function, combine the expected benefit of the user and the current benefit of the candidate network to construct a user satisfaction model, and use the user satisfaction model to construct a reward function;
[0072] It should be noted that the reward function is a key component in the deep reinforcement learning model. It is a feedback mechanism of the environment to the agent's behavior. The reward function defines the immediate reward value obtained by the agent after taking a certain action in a certain state. This reward value is used to measure whether the agent's behavior is "good" or "bad" and guides the agent to develop towards the expected target behavior.
[0073] In the embodiments of the present invention, at each decision-making moment, the mobile terminal will make a handover decision, select a target network based on the current network state, and obtain a feedback reward value. However, when using the existing access selection algorithms for heterogeneous wireless networks, only a network with the best comprehensive performance is selected for the user among all candidate networks, without fully considering various factors such as network performance, user service requirements, and user preferences, and it is impossible to connect the user to the most suitable network. Therefore, in the embodiments of the present invention, in order to better characterize the personalized needs and satisfaction of the user, a psychological curve function is introduced, and the expected benefit of the user and the current benefit of the candidate network are considered to construct a user satisfaction model, and then the user satisfaction model is used to construct a reward function.
[0074] In an alternative embodiment, the constructing a user satisfaction model based on the psychological curve function, combining the expected benefit of the user and the current benefit of the candidate network, includes:
[0075] Normalize the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter, and calculate the current revenue of the candidate network according to the normalized value of each attribute parameter.
[0076] By analyzing the user's historical network access data, calculate the satisfaction degree of the user for each attribute parameter to obtain the expected revenue of the user.
[0077] Based on the psychological curve function, construct a user satisfaction model according to the expected revenue of the user and the current revenue of the candidate network.
[0078] It should be noted that in the embodiments of the present invention, six parameters including price, bandwidth, received signal strength, load rate, bit error rate, and blocking rate are selected as the construction parameters of the user satisfaction model. To solve the differences in the value ranges and units of each parameter in different wireless networks, the above six parameters are normalized respectively. Among them, the attribute parameters of the network can be divided into benefit type and cost type. The smaller the cost type parameters, the better, such as price, bit error rate, blocking rate, and load rate; while the larger the benefit type parameters, the better, such as received signal strength and bandwidth.
[0079] Specifically, the step of respectively normalizing the attribute parameters of the candidate network to obtain the normalized value of each attribute parameter and calculating the current revenue of the candidate network according to the normalized value of each attribute parameter includes:
[0080] Normalize the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter;
[0081] Divide the attribute parameters of the candidate network into three categories: cost, system performance, and stability. Among them, the cost type attribute parameters include the price of the candidate network, the system performance type attribute parameters include the bandwidth and received signal strength of the candidate network, and the stability type attribute parameters include the load rate, bit error rate, and blocking rate of the candidate network;
[0082] Use the fuzzy analytic hierarchy process to calculate the weight values of each parameter in the system performance type attribute parameters and the stability type attribute parameters respectively;
[0083] Define the normalized value of the price as the current cost revenue of the candidate network, calculate the current system performance revenue of the candidate network according to the normalized values and weight values of the bandwidth and received signal strength respectively, and calculate the current stability revenue of the candidate network according to the normalized values and weight values of the load rate, bit error rate, and blocking rate respectively.
[0084] Exemplarily, define the normalized value of the price as the comprehensive value of the cost attribute to obtain the current cost revenue of the candidate network. The comprehensive values of the system performance attribute and the stability attribute are the sum of the products of the weights and the corresponding parameter normalized values, and the current system performance revenue of the candidate network and the current stability revenue of the candidate network are obtained respectively.
[0085] It should be noted that, in order to obtain the weights of different attributes, the embodiments of the present invention use the Fuzzy Analytic Hierarchy Process (FAHP) to calculate the weights of broadband and received signal strength within the system performance and the weights of bit error rate, blocking rate, and load rate within the stability respectively.
[0086] The Fuzzy Analytic Hierarchy Process is a multi-criteria decision-making method that combines the traditional Analytic Hierarchy Process (AHP) and fuzzy set theory. In practical problems, the evaluation and comparison of many attribute parameters often have fuzziness. For example, a user may describe that "this wireless network signal seems to be okay", and "seems to be okay" is a fuzzy expression. It does not specifically indicate how many dBm the signal strength is, and "seems to be okay" may represent different actual signal strength ranges in the cognition of different users. Perhaps for a user with low requirements, around -70 dBm is "seems to be okay", but for a user who needs to conduct high-definition video conferencing and has high requirements for network quality, below -60 dBm will be considered "seems to be okay". Using the Fuzzy Analytic Hierarchy Process can handle this kind of fuzzy information well. By constructing a fuzzy judgment matrix, people's fuzzy evaluations are transformed into a form that can be quantified and calculated, so as to more accurately reflect the actual situation.
[0087] In addition, the Fuzzy Analytic Hierarchy Process retains the hierarchical structure characteristics of the traditional Analytic Hierarchy Process. Taking the selection of heterogeneous wireless networks in the embodiments of the present invention as an example, the factors affecting the selection can be divided into three levels, and each level corresponds to the target level, the criterion level, and the scheme level respectively. The target level is the user benefit of system performance and stability; the criterion level is each decision factor affecting system performance and stability; the scheme level is the normalized values of each parameter in each candidate network. This hierarchical structure enables complex multi-attribute decision-making problems to be systematically decomposed, facilitating decision-makers to analyze problems from different levels and perspectives, clarifying the mutual relationships between various attribute parameters, and avoiding the situation of only considering a single factor and ignoring other important factors.
[0088] Specifically, the psychological curve function is:
[0089] ;
[0090] where is the user satisfaction; is the expected benefit of the user; is the current benefit of the candidate network;
[0091] The user satisfaction model is:
[0092] ;
[0093] ;
[0094] ;
[0095] Among them, represents selecting the th network, represents the state of the network, represents the action taken; is the cost satisfaction model, represents the expected cost benefit after taking action in state ; represents the th network's current cost benefit after taking action in state ; is the system performance satisfaction model, represents the expected system performance benefit after taking action in state ; represents the th network's current system performance benefit after taking action in state ; is the stability satisfaction model, represents the expected stability benefit after taking action in state ; represents the th network's current stability benefit after taking action in state ;
[0096] Furthermore, the reward function is:
[0097] ;
[0098] Among them, , and are the weight values of the cost - type attribute parameter, system performance attribute parameter, and stability attribute parameter respectively.
[0099] It should be noted that for the weight calculation of the three major types of attributes of cost, system performance, and stability, the fuzzy analytic hierarchy process can also be used to construct a fuzzy complementary judgment matrix for calculation.
[0100] S4. According to the network state space, the action space, and the reward function, construct a deep reinforcement learning neural network model in combination with a back - propagation neural network;
[0101] It should be noted that the present invention uses Q-learning to construct a deep reinforcement learning model, and then uses a BP (BackPropagation) neural network, that is, a backpropagation neural network, to approximate the state-action value function in Q-learning. The BP neural network selects a switching target network according to the Q value using the ε-greedy strategy.
[0102] Exemplarily, the number of nodes in the input layer of the neural network is determined according to the state space dimension and is used to receive state information; the number of hidden layers and nodes is determined according to the problem complexity and resources, and a multi-layer perceptron using a non-linear activation function is used to increase the expressiveness; the number of nodes in the output layer depends on the action space. For a discrete action space, Softmax is used to output the action probability distribution, and for a continuous action space, it is designed according to the specific action representation; the state information is input into the input layer through forward propagation, calculated and transmitted through the hidden layer to the output layer to obtain the action probability distribution or specific action value. According to the target design, corresponding rewards are given to different access results, a loss function is constructed based on the reward function, the gradient is calculated starting from the output layer, and the weights are updated by backpropagation. An optimization algorithm is used to control the learning rate.
[0103] S5. Train the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain a hyper-dense heterogeneous wireless network access control model;
[0104] In an optional embodiment, the training of the deep reinforcement learning neural network model through the evolutionary strategy algorithm to obtain a hyper-dense heterogeneous wireless network access control model includes:
[0105] Initialize the parameters of the deep reinforcement learning neural network model through the evolutionary strategy algorithm;
[0106] Construct a training data set using the historical handover decision data of the hyper-dense heterogeneous wireless network;
[0107] Set a main network and a target network for the deep reinforcement learning neural network model;
[0108] Use the training data set to train the main network and the target network respectively, update the parameters of the main network, and stop training when the preset training conditions are met to obtain a hyper-dense heterogeneous wireless network access control model.
[0109] It should be noted that the deep reinforcement learning neural network model includes a main network and a target network. Among them, the main network is constructed using a BP neural network, and the target network has the same structure as the main network and the same initial weights. The difference is that the main network is updated in each iteration, while the target network is updated only every once in a while. Moreover, the main network selects an access network through the ε-greedy strategy, while the target network selects an access network through the greedy strategy.
[0110] S6. Execute the handover strategy of the ultra-dense heterogeneous wireless network using the described ultra-dense heterogeneous wireless network access control model.
[0111] Exemplarily, in a modern intelligent office park, multiple types of wireless network access points are deployed within the park, including traditional Wi-Fi 6 access points, 5G micro base stations, and future 6G experimental base stations, forming an ultra-dense heterogeneous wireless network environment. The park is densely populated with a large number of mobile office devices, such as smartphones, tablets, laptops, etc. These devices need to seamlessly switch between different networks to ensure the continuous and efficient operation of services.
[0112] First, continuously monitor the status information of mobile devices through the ultra-dense heterogeneous wireless network access control model. For example, an employee is walking in the park with a smartphone. It is detected that the signal strength of the Wi-Fi 6 access point currently connected to the phone is gradually weakening, the packet loss rate is increasing, and the phone is moving towards the coverage area of the 5G micro base station. At the same time, a video conferencing application is running on the phone. Then, based on the collected status information, make a decision according to the predefined handover strategy. In the above example, considering the current download task (high bandwidth requirement) being executed on the laptop, decide to start the network handover program according to the handover strategy, and the target handover network is the 5G micro base station with the strongest signal nearby.
[0113] Through such an ultra-dense heterogeneous wireless network access control model and its handover strategy, it is possible to effectively cope with complex and changing network environments, meet the network connection requirements of different users in different scenarios, and improve the overall utilization efficiency of the ultra-dense heterogeneous wireless network.
[0114] In summary, an ultra-dense heterogeneous wireless network access control method provided by an embodiment of the present invention fully considers the personalized needs of users, comprehensively considers the attribute parameters on the user side and the network side, constructs a more objective user satisfaction model to quantify user satisfaction, enables the network selection strategy to dynamically adapt to changes in network status and user needs, and realizes the optimal selection and access of the network; at the same time, the present invention introduces deep reinforcement learning into the access selection problem of heterogeneous wireless networks, models the network selection and access problem as a reinforcement learning model, and can learn the optimal strategy through interaction with the environment and make intelligent access selection decisions.
[0115] To make those skilled in the art more clear about the implementation process of an ultra-dense heterogeneous wireless network access control method provided by an embodiment of the present invention, the following provides a more specific description.
[0116] First, the construction of the deep reinforcement learning model elements mainly includes three parts: state space definition, action space definition, and reward function design. Define the above three elements as follows:
[0117] 1. State Space (State)
[0118] Exemplarily, the considered state space includes network parameter M, the type of service G currently initiated by the user, and the currently selected network .
[0119] Then at decision-making moment t, the state space can be represented as a set in Equation (1);
[0120] #(1).
[0121] 2. Action Space (Action)
[0122] Assume that the user is under multiple network coverage areas and can only select one network for access at decision-making moment t. Then the action space A can be represented as a set in Equation (2) , where is the action taken by the user at handover decision-making moment t , represents the i-th alternative network.
[0123] #(2).
[0124] 3. Reward Function (Reward)
[0125] At each decision-making moment t, the mobile terminal makes a handover decision. Based on the current network state, it selects a target network and obtains a feedback reward value. To better characterize the personalized needs and satisfaction of the user, the embodiment of the present invention introduces a psychological curve function, considers the expected benefit of the user and the current benefit of the candidate network, constructs a user satisfaction model, and constructs a reward function using the user satisfaction model.
[0126] Among them, the psychological curve function is expressed as:
[0127] #(3);
[0128] In Equation (3) is the user satisfaction; is the expected benefit of the user; is the current benefit of the candidate network.
[0129] Furthermore, considering the expected benefit of the user and the current benefit of the candidate network, a user satisfaction model is constructed as follows:
[0130] (1) Normalization of data
[0131] Select six parameters of the wireless network, namely price P, bandwidth B, received signal strength RSS, load factor F, bit error rate D, and blocking rate Z, as the construction parameters of the user satisfaction model. To address the differences in the value ranges and units of the parameters in different wireless networks, the above six parameters are normalized respectively. Among them, for cost-type parameters, the smaller the better, such as price, bit error rate, blocking rate, and load factor; for benefit-type parameters, the larger the better, such as received signal strength and bandwidth.
[0132] The normalized values of cost-type and benefit-type parameters are as shown in equations (4) and (5):
[0133] #(4);
[0134] #(5);
[0135] Among them, is the normalized value of the th network parameter , which is used to represent the six parameters of price P, bandwidth B, received signal strength RSS, bit error rate D, blocking rate Z, and load factor F respectively, is the set of numerical values of the parameters of all candidate networks , and is the numerical value of the parameter of network
[0136] (2) Calculate the weight values using the fuzzy analytic hierarchy process
[0137] The system performance attributes consist of bandwidth B and received signal strength RSS, and the stability attributes consist of bit error rate D, blocking rate Z, and load factor F. To obtain the weights of different attributes, the embodiments of the present invention use the fuzzy analytic hierarchy process to calculate the weights of bandwidth B and received signal strength RSS within system performance, and the weights of bit error rate D, blocking rate Z, and load factor F within stability respectively.
[0138] Specifically, first, establish an analytic hierarchy structure diagram, as shown in Figure 2 , Figure 2 which is an analytic hierarchy structure diagram provided by the embodiments of the present invention. As shown in Figure 2 , it is divided into three levels, corresponding to the goal level, criterion level, and scheme level respectively. The goal level is the user benefit of system performance and stability; the criterion level is each decision factor affecting system performance and stability; the scheme level is the normalized values of the parameters in each candidate network.
[0139] Secondly, for the system performance attributes and stability attributes respectively, construct a system performance fuzzy complementary judgment matrix and a stability fuzzy complementary judgment matrix 。
[0140] Among them, , obtained by the 0.1 - 0.9 scale method, and the 0.1 - 0.9 scale method is shown in Table 1, = 1, + = 1.
[0141] Table 1: 0.1 - 0.9 Scale Method
[0142]
[0143] The weight calculation formula of the fuzzy complementary judgment matrix is as shown in Equation (6):
[0144] #(6);
[0145] Using the above - mentioned weight calculation of the fuzzy complementary judgment matrix to calculate the weights of each parameter, we get:
[0146] , ;
[0147] Among them, , are the weight vectors of system performance and stability respectively, 、 、 、 、 are the weights of the received signal strength RSS, bandwidth B, bit error rate D, blocking rate Z, and load rate F respectively.
[0148] Finally, perform a consistency test for the comparison and judgment. First, construct the characteristic matrices of the fuzzy matrices 、 respectively:
[0149] ;
[0150] Among them, 。 represents the weight of the a - th row and b - th column.
[0151] Use Equation (7) to solve the compatibility value:
[0152] #(7);
[0153] If , it is considered that the consistency meets the standard and the weight value result is reasonable; otherwise, the fuzzy complementary judgment matrix needs to be adjusted and the weights need to be solved again.
[0154] (3) Calculate the current income
[0155] Define the normalized value of the price as the comprehensive value of the cost attribute, and the comprehensive value of the system performance attribute and the stability attribute is the cumulative sum of the products of the weights and the normalized values of the corresponding parameters.
[0156] #(8);
[0157] #(9);
[0158] #(10);
[0159] In equations (8)-(10), 、 、 are respectively the comprehensive values of the cost pc, system performance sp, and stability st of the th network.
[0160] Define the comprehensive value as the current revenue, that is, based on the real-time data of the current candidate network, including the current revenues of the cost, system performance, and stability of the th network 、 、 . Among them, represents the current revenue of the cost after the th network takes the action in the state , represents the current revenue of the system performance after the th network takes the action in the state , represents the current revenue of the stability after the th network takes the action in the state .
[0161] (4) Calculate the expected revenue
[0162] By analyzing the user's historical access data, calculate the user's satisfaction with each attribute in the historical situation, and obtain the expected revenue of the user for each attribute.
[0163] The calculation formula for the expected revenue is as shown in equations (11)-(13):
[0164] #(11);
[0165] #(12);
[0166] #(13);
[0167] Among them, , , are the comprehensive values of cost, system performance, and stability under the th network respectively. h is the number of times of the overlapping network situation, is the average value of the comprehensive value of the user's network attribute requirements under the historical situation of h - times overlap, defined as the expected revenue; represents the expected revenue of cost after taking action in state , represents the expected revenue of system performance after taking action in state , represents the expected revenue of stability after taking action in state .
[0168] (5) Specific satisfaction calculation
[0169] Based on the psychological curve function, combined with the expected revenue and current revenue of the user in the three major categories of attributes, a satisfaction model for the three major categories of attributes of the candidate network is proposed.
[0170] #(14);
[0171] #(15);
[0172] #(16);
[0173] Equations (14)-(16) are the satisfaction models for cost, system performance, and stability respectively.
[0174] Furthermore, the reward function is defined as in Equation (17):
[0175] #(17);
[0176] where , , are the weight values of the three major categories of attributes respectively. A fuzzy complementary judgment matrix is constructed for the three major categories of attributes of cost, system performance, and stability, and the calculation process is the same as that of calculating the weight value by the above - mentioned fuzzy analytic hierarchy process.
[0177] Secondly, fuse the BP (Back Propagation) neural network and the deep Q - network DQN, and use the BP neural network to approximate the state - action value function. The process includes the following:
[0178] Specifically, the initial weight parameters of the BP neural network are obtained by an Evolution Strategies (ES) with global optimization ability, and then the parameters of the neural network are trained and adjusted by the gradient descent method.
[0179] 1. Main network modeling:
[0180] See Figure 3 , Figure 3 which is a schematic diagram of a main network structure provided by an embodiment of the present invention. As Figure 3 shown, in the embodiment of the present invention, a BP neural network is used to construct the main network, and the main network is composed of an input layer, two hidden layers and an output layer.
[0181] Each state-action pair corresponds to a value function , and using the main network to approximate the state-action value function can be expressed as:
[0182] #(18);
[0183] where represents the expectation of the long-term cumulative return obtained by taking the switching action when the network state is , is the parameter of the main network, is the switching strategy adopted for switching, which can be defined as:
[0184] #(19);
[0185] #(20);
[0186] where represents the expectation under the policy , is the cumulative discounted reward, is the discount factor (0 ≤ ≤ 1), indicating the degree of emphasis on the rewards obtained in the future, represents the reward at time
[0187] From the attribute parameters of each network and the total number of all candidate networks, it can be known that the number of neurons in the input layer is 6n + 2. After the forward propagation of the neural network, the values of all state-action pairs in the current state are output: , then the number of neurons in the output layer is n. Based on the ε-greedy strategy, the switching target network is selected according to the output value. At this time, based on the current network state and the switching decision made The end - user will get an immediate feedback reward value r, and the network state will also change to reach the next state Store the switching decision data generated by this interaction process into the training database for updating the main network, and then enter the next step, continuously looping until the BP network converges.
[0188] 2. Target network modeling:
[0189] The target network has the same structure as the main network and the same initial weights. The difference is that the main network is updated every iteration, while the target network is updated only after a certain interval. The parameters of the target network are represented by . The input of the target network is , and the output is . The target network uses the greedy strategy to select the switching action , is the action that makes value the largest.
[0190] See Figure 4 , Figure 4 which is the structural block diagram of a deep reinforcement learning neural network model provided by an embodiment of the present invention.
[0191] As Figure 4 shown, in a heterogeneous wireless network, the algorithm selects an action (i.e., selects to switch to a specific network) according to the current network state . Using the BP neural network as the main network, it receives the current state of the heterogeneous wireless network and the action as inputs, and outputs , that is, the expected return of executing this action in the current state. Every N steps, the parameters of the main network are copied to the target network. The target network, which has the same structure as the main network but a lower parameter update frequency, is used to calculate the target value, that is . By comparing and to calculate the loss, the gradient descent method is used to guide the update of the main network parameters to improve the value prediction accuracy. The training database stores the data generated during the interaction between the agent and the environment , where r is the immediate reward obtained after executing the action , and is the next state.
[0192] 3. Establish a training database:
[0193] Specifically, an experience replay mechanism can be introduced. At each decision-making moment, the historical decision-making data obtained from the interaction between the terminal and the environment is stored in the database for repeated learning and training of the neural network. During the training process, a small batch of training samples is randomly drawn from the database, which reduces the correlation between samples and improves the algorithm stability.
[0194] Among them, the main network selects the access network through the ε-greedy strategy, while the target network selects the access network through the greedy strategy. In DQN, the ε-greedy strategy is used to balance exploration and exploitation. Among them, the process of selecting the current optimal action is exploitation, and the process of selecting other non-optimal actions is exploration, that is, randomly selecting one from all candidate networks with a probability of ε, and selecting the one that makes the value the largest with a probability of 1 - ε. When the neural network converges and can accurately evaluate the values of all state actions, it should be completely inclined to exploitation in exploration and exploitation. Therefore, as the number of training times increases, the value of 1 - ε can be gradually increased. The greedy strategy is a greedy algorithm that selects the access network that makes the value the largest each time, that is .
[0195] 4. Training the main network
[0196] 4.1 Initializing the neural network parameters of the main network
[0197] It should be noted that the Evolution Strategy (ES) algorithm is an optimization algorithm based on the idea of natural evolution. It optimizes mathematical functions through iterative search and is mainly used to solve parameter optimization problems. This algorithm is a black-box optimization algorithm and a heuristic search process inspired by natural evolution: in each iteration, a random perturbation is applied to the parameter vector. Subsequently, the algorithm evaluates its objective function value, i.e., fitness, and then recombines the parameter vector with the highest objective function value to form the next generation of vector groups (populations), and iterates this process until the objective is fully optimized.
[0198] It is worth noting that the neural network is pre-trained using the Evolution Strategy (ES) to obtain a relatively optimal initial neural network parameter. The initial neural network parameter of the main network is the optimization objective. The distribution of each generation of the population follows an isotropic multivariate Gaussian distribution with a mean of and a variance of , where represents the parameters of the main network in the k-th generation. To simplify the optimization process, additive Gaussian noise is applied to the current parameter vector. A mirrored approach is used to generate Gaussian noise vectors that are always evaluated in pairs: and 。
[0199] The offspring parameters can be expressed as:
[0200] #(21);
[0201] where , represents the parameter of the -th offspring in the k-th generation.
[0202] In the Evolutionary Strategy (ES), the fitness of each offspring individual is evaluated by placing it in a heterogeneous wireless network environment, and the fitness function represents . To improve the performance of the neural network parameters , the embodiments of the present invention adopt the Stochastic Gradient Ascent method to maximize the average reward of the population, and its mathematical expression is as follows:
[0203] #(22);
[0204] #(23);
[0205] In equation (23), represents the learning rate, and n refers to half of the number of offspring populations. Through the process of repeated iteration, when the cumulative reward value gradually reaches a stable state, the obtained weight parameters will be used to initialize the parameter settings of the main network.
[0206] 4.2 Update of Main Network Parameters
[0207] To optimize the handover decision, it is necessary for the user terminal to continuously interact with the heterogeneous network environment and adjust the handover strategy according to the feedback reward provided by the environment, that is, to optimize the weight parameters of the neural network through training, so as to continuously strengthen the handover decision-making ability of the terminal. The training of the neural network is essentially a process of minimizing the loss function, and the loss function reflects the deviation between the model prediction and the actual label.
[0208] The widely used Mean Square Error model is adopted to construct the loss function, and the error backpropagation and gradient descent methods are used to iteratively solve it to achieve the minimization of the loss function. The loss function of the ES-DQN (Deep Reinforcement Learning Neural Network Model) in the embodiments of the present invention is defined as:
[0209] #(24);
[0210] where is the estimated value, is the output value of the main network, is the discount factor.
[0211] 4.3 Loop training
[0212] See Figure 5 , Figure 5 which is another flowchart of an access control method for an ultra-dense heterogeneous wireless network provided by an embodiment of the present invention. As Figure 5 shown, after constructing the reinforcement learning model, use the BP neural network to approximate the state-action value function in Q learning, pre-train the initial parameters of the neural network through the evolutionary strategy (ES), and construct a training data set as the training samples of the neural network. Then set the main network to approximate the current state-action value function and continuously update the parameters, and set the target network to calculate Q reality, and the parameters are periodically copied from the main network; randomly select a batch of samples from the training database to start loop training, update the parameters of the main network, check whether the preset number of training rounds or convergence conditions are reached. If so, the training ends and the handover decision is executed. Otherwise, continue the training loop.
[0213] In summary, in the embodiments of the present invention, first, by constructing a network state space based on the attribute parameters of candidate networks, the state of the ultra-dense heterogeneous wireless network can be comprehensively and meticulously described, and the real-time state of each candidate network can be accurately captured, providing detailed network condition information for subsequent access control and handover strategies, which helps to better adapt to the characteristics of complex and changeable network states in the ultra-dense heterogeneous network environment; moreover, by constructing a user satisfaction model based on the psychological curve function and constructing a reward function therefrom, this is an innovative way that takes into account the subjective feelings of users. The user satisfaction model combines the expected benefits of users and the current benefits of candidate networks, and can more accurately reflect the true evaluation of users on network services. This reward mechanism enables the access control strategy to better meet the needs of users, optimize network access and handover decisions, and improve the user experience; then, by using the network state space, action space, and reward function, combined with a backpropagation neural network to construct a deep reinforcement learning neural network model, the optimal strategy for network access and handover can be effectively learned. This automatic learning ability can adapt to the complexity and dynamics of the ultra-dense heterogeneous network, reduce manual intervention, and improve the intelligent level of access control; at the same time, by training the deep reinforcement learning neural network model with an evolutionary strategy algorithm, the performance of the model can be further optimized. The evolutionary strategy algorithm can search for the optimal solution in a complex parameter space, and it can evaluate and optimize the parameters of the model according to the fitness function. Compared with traditional training methods, this algorithm can better handle the high-dimensional parameters and complex optimization objectives in the ultra-dense heterogeneous network access control model, improve the convergence speed and accuracy of the model, and thus obtain a more effective ultra-dense heterogeneous wireless network access control model; finally, by using the obtained ultra-dense heterogeneous wireless network access control model to execute the network handover strategy, more intelligent and efficient network handover can be achieved, and the user experience in a complex network environment can be improved.
[0214] Based on the above method items, the present invention correspondingly provides an embodiment of the system item.
[0215] See Figure 6 , Figure 6 FIG. is a structural block diagram of an ultra-dense heterogeneous wireless network access control system provided by an embodiment of the present invention. The ultra-dense heterogeneous wireless network access control system includes:
[0216] A first model element construction module 21, configured to construct a network state space according to the attribute parameters of candidate networks;
[0217] A second model element construction module 22, configured to construct an action space according to candidate networks that can be selected for access;
[0218] The third model element construction module 23 is used to construct a user satisfaction model based on a psychological curve function, in combination with the expected benefit of the user and the current benefit of the candidate network, and construct a reward function using the user satisfaction model;
[0219] The deep reinforcement learning neural network model construction module 24 is used to construct a deep reinforcement learning neural network model according to the network state space, the action space, and the reward function, in combination with a backpropagation neural network;
[0220] The deep reinforcement learning neural network model training module 25 is used to train the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model;
[0221] The network access control module 26 is used to execute the handover strategy of the ultra-dense heterogeneous wireless network by using the ultra-dense heterogeneous wireless network access control model.
[0222] In an alternative embodiment, the attribute parameters of the candidate network include user personalized demand parameters and network state information parameters;
[0223] The user personalized demand parameters include the price, bandwidth, and received signal strength of the candidate network;
[0224] The network state information parameters include the load rate, bit error rate, and blocking rate of the candidate network.
[0225] In an alternative embodiment, the third model element construction module 23 includes a user satisfaction model construction unit and a reward function construction unit;
[0226] Among them, the user satisfaction model construction unit is specifically used for:
[0227] Normalize the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter, and calculate the current benefit of the candidate network according to the normalized value of each attribute parameter;
[0228] By analyzing the user's historical network access data, calculate the user's satisfaction with each attribute parameter to obtain the user's expected benefit;
[0229] Based on the psychological curve function, construct a user satisfaction model according to the user's expected benefit and the current benefit of the candidate network.
[0230] Specifically, the normalizing the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter, and calculating the current benefit of the candidate network according to the normalized value of each attribute parameter includes:
[0231] Normalize the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter;
[0232] Classify the attribute parameters of the candidate network into three categories: cost, system performance, and stability. Among them, the cost - type attribute parameters include the price of the candidate network, the system - performance - type attribute parameters include the bandwidth and received signal strength of the candidate network, and the stability - type attribute parameters include the load rate, bit - error rate, and blocking rate of the candidate network;
[0233] Use the fuzzy analytic hierarchy process to calculate the weight values of each parameter in the system - performance - type attribute parameters and stability - type attribute parameters respectively;
[0234] Define the normalized value of the price as the current cost benefit of the candidate network, calculate the current system - performance benefit of the candidate network according to the normalized values and weight values of the bandwidth and received signal strength respectively, and calculate the current stability benefit of the candidate network according to the normalized values and weight values of the load rate, bit - error rate, and blocking rate respectively.
[0235] In an alternative embodiment, the psychological curve function is:
[0236] ;
[0237] Among them, is the user satisfaction; is the expected benefit of the user; is the current benefit of the candidate network;
[0238] The user - satisfaction model is:
[0239] ;
[0240] ;
[0241] ;
[0242] Among them, represents selecting the th network, represents the state of the network, represents the action taken; is the cost - satisfaction model, represents the expected cost benefit after taking action in state ; represents the th network's current cost benefit after taking action in state ; is the system - performance - satisfaction model, represents in state Take an action below The expected benefit of system performance after Indicates the th network in state Take an action below The current benefit of system performance after Is the stability satisfaction model Indicates in state Take an action below The expected benefit of stability after Indicates the th network in state Take an action below The current benefit of stability after
[0243] Furthermore, the reward function is:
[0244] ;
[0245] Wherein, , And Are the weight values of the cost - type attribute parameter, the system performance attribute parameter, and the stability attribute parameter respectively
[0246] It should be noted that a super - dense heterogeneous wireless network access control system provided by an embodiment of the present invention is used to execute all the process steps of a super - dense heterogeneous wireless network access control method in the above - mentioned embodiment. The working principles and beneficial effects of the two correspond one by one, so they will not be elaborated here
[0247] An embodiment of the present invention also provides a terminal device. As Figure 7 Shown, it is a structural block diagram of a terminal device provided by an embodiment of the present invention. The terminal device includes a processor 31, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 31. When the processor 31 executes the computer program, it implements the super - dense heterogeneous wireless network access control method described in any of the above embodiments
[0248] In addition, an embodiment of the present invention also provides a computer - readable storage medium. The computer - readable storage medium includes a stored computer program. Wherein, when the computer program runs, it controls the device where the computer - readable storage medium is located to execute the super - dense heterogeneous wireless network access control method described in any of the above embodiments
[0249] When the processor 31 executes the computer program, it implements the steps in the above - mentioned super - dense heterogeneous wireless network access control method embodiment. For example Figure 1All steps of the ultra-dense heterogeneous wireless network access control method shown. Alternatively, when the processor 31 executes the computer program, it implements the functions of each module in the above ultra-dense heterogeneous wireless network access control system embodiment, such as Figure 6 The functions of each module of the ultra-dense heterogeneous wireless network access control system shown.
[0250] Preferably, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 32 and executed by the processor 31 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the terminal device.
[0251] The processor 31 can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor 31 can also be any conventional processor. The processor 31 is the control center of the terminal device and connects various parts of the terminal device through various interfaces and lines.
[0252] The memory 32 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc., and the data storage area can store relevant data, etc. In addition, the memory 32 can be a high-speed random access memory, or can also be a non-volatile memory, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., or the memory 32 can also be other volatile solid-state storage devices.
[0253] It should be noted that the above terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 7 The structural block diagram shown is only an example of the structure of the above terminal device and does not constitute a limitation on the structure of the above terminal device. The above terminal device may include more or fewer components than shown, or combine certain components, or different components.
[0254] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A method for access control in an ultra-dense heterogeneous wireless network, characterized in that Including: Construct a network state space according to the attribute parameters of the candidate network; Construct an action space according to the candidate networks that can be selected for access; Based on the psychological curve function, combine the expected revenue of the user and the current revenue of the candidate network to construct a user satisfaction model, and use the user satisfaction model to construct a reward function; According to the network state space, the action space and the reward function, combine a backpropagation neural network to construct a deep reinforcement learning neural network model; Train the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain a hyper-dense heterogeneous wireless network access control model; Adopt the hyper-dense heterogeneous wireless network access control model to execute the handover strategy of the hyper-dense heterogeneous wireless network; Wherein, the psychological curve function is: ; Among them, is the user satisfaction; is the expected benefit of the user; is the current benefit of the candidate network; The user satisfaction model is: ; ; ; Among them, represents selecting the th network, represents the status of the network, represents the action taken; is the cost satisfaction model, represents the expected cost benefit after taking action in state ; represents the th network's current cost benefit after taking action in state ; is the system performance satisfaction model, represents the expected system performance benefit after taking action in state ; represents the th network's current system performance benefit after taking action in state ; is the stability satisfaction model, represents the expected stability benefit after taking action in state ; represents the th network's current stability benefit after taking action in state ; The reward function is: ; Among them, , and are the weight values of the cost - type attribute parameter, the system performance attribute parameter, and the stability attribute parameter, respectively.
2. The ultra-dense heterogeneous wireless network access control method according to claim 1, wherein The attribute parameters of the candidate network include user personalized demand parameters and network state information parameters; The user personalized demand parameters include the price, bandwidth and received signal strength of the candidate network; The network state information parameters include the load rate, bit error rate and blocking rate of the candidate network.
3. The ultra-dense heterogeneous wireless network access control method according to claim 2, wherein The constructing a user satisfaction model based on the psychological curve function according to the expected revenue of the user and the current revenue of the candidate network includes: Normalize the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter, and calculate the current revenue of the candidate network according to the normalized value of each attribute parameter; By analyzing the historical network access data of the user, calculate the satisfaction of the user with each attribute parameter to obtain the expected revenue of the user; Based on the psychological curve function, construct a user satisfaction model according to the expected revenue of the user and the current revenue of the candidate network.
4. The ultra-dense heterogeneous wireless network access control method according to claim 3, wherein The respectively normalizing the attribute parameters of the candidate network to obtain the normalized value of each attribute parameter and calculating the current revenue of the candidate network according to the normalized value of each attribute parameter includes: Normalize the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter; Classify the attribute parameters of the candidate network into three categories: cost, system performance and stability; among them, the cost category attribute parameters include the price of the candidate network, the system performance category attribute parameters include the bandwidth and received signal strength of the candidate network, and the stability category attribute parameters include the load rate, bit error rate and blocking rate of the candidate network; Use the fuzzy analytic hierarchy process to calculate the weight values of each parameter in the system performance category attribute parameters and the stability category attribute parameters respectively; Define the normalized value of the price as the current cost revenue of the candidate network, calculate the current system performance revenue of the candidate network according to the normalized values and weight values of the bandwidth and received signal strength respectively, and calculate the current stability revenue of the candidate network according to the normalized values and weight values of the load rate, bit error rate and blocking rate respectively.
5. The ultra-dense heterogeneous wireless network access control method according to claim 1, characterized in that The training the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain a hyper-dense heterogeneous wireless network access control model includes: Initialize the parameters of the deep reinforcement learning neural network model through an evolutionary strategy algorithm; Construct a training data set by using the historical handover decision data of the hyper-dense heterogeneous wireless network; Set up a main network and a target network for the deep reinforcement learning neural network model; Use the training data set to train the main network and the target network respectively, update the parameters of the main network, and stop training when the preset training conditions are met to obtain a hyper-dense heterogeneous wireless network access control model.
6. A super-dense heterogeneous wireless network access control system, characterized in that, It includes: The first model element construction module is used to construct a network state space according to the attribute parameters of the candidate network; The second model element construction module is used to construct an action space according to the candidate networks that can be selected for access; The third model element construction module is used to construct a user satisfaction model based on the psychological curve function, combining the expected benefits of the user and the current benefits of the candidate network, and construct a reward function using the user satisfaction model; The deep reinforcement learning neural network model construction module is used to construct a deep reinforcement learning neural network model by combining the network state space, the action space and the reward function with a backpropagation neural network; The deep reinforcement learning neural network model training module is used to train the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain a hyper-dense heterogeneous wireless network access control model; The network access control module is used to execute the handover strategy of the hyper-dense heterogeneous wireless network using the hyper-dense heterogeneous wireless network access control model; Among them, the psychological curve function is: ; Among them, is the user satisfaction; is the expected benefit of the user; is the current benefit of the candidate network; The user satisfaction model is: ; ; ; Among them, represents selecting the th network, represents the status of the network, represents the action taken; is the cost satisfaction model, represents the expected cost benefit after taking action in the state ; represents the th network's current cost benefit after taking action in the state ; is the system performance satisfaction model, represents the expected system performance benefit after taking action in the state ; represents the th network's current system performance benefit after taking action in the state ; is the stability satisfaction model, represents the expected stability benefit after taking action in the state ; represents the th network's current stability benefit after taking action in the state ; The reward function is: ; Among them, , and are the weight values of the cost - type attribute parameter, the system performance attribute parameter, and the stability attribute parameter, respectively.
7. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the hyper-dense heterogeneous wireless network access control method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the hyper-dense heterogeneous wireless network access control method according to any one of claims 1 to 5.
Citation Information
Patent Citations
5G ultra-dense network multi-user access selection method based on deep reinforcement learning
CN114449536A
Edge computing unloading method based on dynamic user satisfaction in ultra-dense network
CN114641076A