Ultra-dense heterogeneous wireless network access control method, system, device and medium
Through deep reinforcement learning methods, the user satisfaction model and reward function are constructed, combined with backpropagation neural networks, the problem that users cannot access the most suitable network in super-dense heterogeneous wireless networks is solved, and the optimal access strategy to dynamically adapt to network status and user needs is realized, improving user experience and network efficiency.
Patent Information
- Application Number
- CN202510458267.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
In the ultra-intensive heterogeneous wireless network environment, the existing access selection algorithm cannot fully consider network performance, user service needs and user preferences, resulting in users that may be connected to poor quality networks and cannot effectively guarantee the service quality of services and the user experience quality.
The deep reinforcement learning method is adopted to build a user satisfaction model and reward function, and combine a backpropagation neural network to build a deep reinforcement learning neural network model, dynamically adapt to the network state and user needs, and achieve the best choice of network access.
It has achieved better meeting users' personalized needs, and a more objective user satisfaction model has been built, so that network selection strategies can dynamically adapt to changes in network status and user needs, improving user experience and network usage efficiency.
Smart Images

Figure CN119997151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless network access technology, and in particular to an ultra-dense heterogeneous wireless network access control method, system, device and medium. Background Art
[0002] A wireless communication system with overlapping signal ranges composed of wireless networks of different architectures such as mobile cellular networks, wireless local area networks, wireless metropolitan area networks, and satellite communication networks is called an ultra-dense heterogeneous wireless network. The coexistence and integration of wireless networks with different access technologies has become the development trend of the next generation of mobile Internet. In an ultra-dense heterogeneous wireless network environment, how to connect users to the most suitable network has always been one of the key and difficult points of research.
[0003] Traditional ultra-dense heterogeneous wireless network access selection algorithms mainly use received signal strength (RSS) as a parameter for network selection, and mobile users choose to access the network with the highest RSS. Although the RSS-based access selection algorithm has low complexity and is easy to implement, it often causes a serious ping-pong effect. Secondly, some access selection algorithms use network load as the basis for network selection, connecting users to the network with the lowest load to achieve load balancing. Although such algorithms improve the resource utilization rate of heterogeneous wireless networks, they do not consider the needs of user services. Users may be connected to networks with poor quality. Therefore, the service quality of the service and the user experience quality cannot be effectively guaranteed. In addition, there are also access selection algorithms designed using the Multiple Attribute Decision Making (MADM) theory.
[0004] However, the above method is used to design an access selection algorithm for heterogeneous wireless networks, which only selects a network with the best comprehensive performance for users among all candidate networks, without fully considering various factors such as network performance, user business needs, and user preferences. Therefore, it is impossible to connect users to the most suitable network. Summary of the invention
[0005] The purpose of the present invention is to provide an ultra-dense heterogeneous wireless network access control method, system, device and medium to better meet the personalized needs of users, build a more objective user satisfaction model, quantify user satisfaction, enable the network selection strategy to dynamically adapt to changes in network status and user needs, and achieve the best selection of network access.
[0006] In order to achieve the above objectives, an embodiment of the present invention provides an ultra-dense heterogeneous wireless network access control method, comprising: Constructing a network state space based on the attribute parameters of the candidate network; Construct an action space based on the candidate networks that can be selected for access; Based on the psychological curve function, a user satisfaction model is constructed by combining the user's expected benefit and the current benefit of the candidate network, and a reward function is constructed using the user satisfaction model; According to the network state space, the action space and the reward function, a deep reinforcement learning neural network model is constructed in combination with a back propagation neural network; The deep reinforcement learning neural network model is trained by an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model; The ultra-dense heterogeneous wireless network access control model is adopted to execute the switching strategy of the ultra-dense heterogeneous wireless network.
[0007] As an improvement of the above solution, the attribute parameters of the candidate network include user personalized demand parameters and network status information parameters; The user personalized demand parameters include the price, bandwidth and received signal strength of the candidate network; The network status information parameters include the load rate, bit error rate and blocking rate of the candidate network.
[0008] As an improvement of the above solution, the user satisfaction model is constructed based on the psychological curve function according to the user's expected benefits and the current benefits of the candidate network, including: Normalizing the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter, and calculating the current benefit of the candidate network according to the normalized value of each attribute parameter; By analyzing the user's historical network access data, the user's satisfaction with each attribute parameter is calculated to obtain the user's expected benefits; Based on the psychological curve function, a user satisfaction model is constructed according to the expected benefits of the user and the current benefits of the candidate networks.
[0009] As an improvement of the above solution, normalizing the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter, and calculating the current benefit of the candidate network according to the normalized value of each attribute parameter includes: Normalizing the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter; The attribute parameters of the candidate network are divided into three categories: cost, system performance and stability. Among them, the cost attribute parameters include the price of the candidate network, the system performance attribute parameters include the bandwidth and received signal strength of the candidate network, and the stability attribute parameters include the load rate, bit error rate and blocking rate of the candidate network. The fuzzy hierarchical analysis method is used to calculate the weight value of each parameter in the system performance attribute parameter and stability attribute parameter respectively; The normalized value of the price is defined as the current benefit of the candidate network's cost. The current benefit of the candidate network's system performance is calculated based on the normalized values and weight values of the bandwidth and received signal strength. The current benefit of the candidate network's stability is calculated based on the normalized values and weight values of the load rate, bit error rate, and blocking rate.
[0010] As an improvement of the above solution, the psychological curve function is: ; in, For user satisfaction; Expected benefits for users; is the current revenue of the candidate network; The user satisfaction model is: ; ; ; in, Indicates selection A network, Indicates the status of the network. Indicates the action taken; is the cost satisfaction model, Indicates in status Take action The expected cost benefit after Indicates The network is in state Take action Costs after current benefits; is the system performance satisfaction model, Indicates in status Take action The expected cost benefit after Indicates The network is in state Take action The current benefits of system performance after is the stability satisfaction model, Indicates in status Take action The expected return after stability is Indicates The network is in state Take action The stability of current earnings after.
[0011] As an improvement of the above solution, the reward function is: ; in, , and They are the weight values of cost attribute parameters, system performance attribute parameters and stability attribute parameters respectively.
[0012] As an improvement of the above solution, the deep reinforcement learning neural network model is trained by an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model, including: Initializing the parameters of the deep reinforcement learning neural network model through an evolutionary strategy algorithm; The training dataset is constructed using historical handover decision data of ultra-dense heterogeneous wireless networks; Setting a main network and a target network for the deep reinforcement learning neural network model; The training data set is used to train the main network and the target network respectively, and the parameters of the main network are updated. The training is stopped when a preset training condition is reached to obtain an ultra-dense heterogeneous wireless network access control model.
[0013] In order to achieve the above objectives, an embodiment of the present invention further provides an ultra-dense heterogeneous wireless network access control system, comprising: A first model element construction module is used to construct a network state space according to attribute parameters of the candidate network; A second model element construction module is used to construct an action space based on the candidate networks that can be selected for access; A third model element construction module is used to construct a user satisfaction model based on the psychological curve function, combined with the user's expected benefits and the current benefits of the candidate network, and use the user satisfaction model to construct a reward function; A deep reinforcement learning neural network model construction module, used to construct a deep reinforcement learning neural network model in combination with a back propagation neural network according to the network state space, the action space and the reward function; A deep reinforcement learning neural network model training module is used to train the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model; The network access control module is used to implement a switching strategy of the ultra-dense heterogeneous wireless network by adopting the ultra-dense heterogeneous wireless network access control model.
[0014] In order to achieve the above objectives, an embodiment of the present invention further provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the ultra-dense heterogeneous wireless network access control method as described in any one of the above.
[0015] In order to achieve the above objectives, an embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the ultra-dense heterogeneous wireless network access control method as described in any one of the above.
[0016] Compared with the prior art, the embodiment of the present invention provides an ultra-dense heterogeneous wireless network access control method, system, device and medium. The method first constructs a network state space, an action space and a reward function, wherein a user satisfaction model is constructed based on a psychological curve function, combined with the user's expected benefit and the current benefit of a candidate network, and a reward function is constructed using the user satisfaction model. Then, according to the network state space, the action space and the reward function, a deep reinforcement learning neural network model is constructed in combination with a back propagation neural network, and then the deep reinforcement learning neural network model is trained by an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model. Finally, the ultra-dense heterogeneous wireless network access control model is used to execute the switching strategy of the ultra-dense heterogeneous wireless network. The present invention fully considers the personalized needs of users and comprehensively considers the attribute parameters of the user side and the network side, and constructs a more objective user satisfaction model for quantifying user satisfaction, so that the network selection strategy can dynamically adapt to changes in network status and user needs, and realize the best selection access of the network; at the same time, the present invention introduces deep reinforcement learning in the access selection problem of heterogeneous wireless networks, and models the network selection access problem as a reinforcement learning model, which can learn the optimal strategy through interaction with the environment and make access selection decisions intelligently. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solution of the present invention, the drawings used in the implementation mode will be briefly introduced below. Obviously, the drawings described below are only some implementation modes of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flow chart of an ultra-dense heterogeneous wireless network access control method provided by an embodiment of the present invention; Figure 2 It is a hierarchical analysis structure diagram provided by an embodiment of the present invention; Figure 3 is a schematic diagram of a main network structure provided by an embodiment of the present invention; Figure 4 It is a structural block diagram of a deep reinforcement learning neural network model provided by an embodiment of the present invention; Figure 5 is another flow chart of an ultra-dense heterogeneous wireless network access control method provided by an embodiment of the present invention; Figure 6 It is a structural block diagram of an ultra-dense heterogeneous wireless network access control system provided by an embodiment of the present invention; Figure 7 It is a structural block diagram of a terminal device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] See also Figure 1 , Figure 1 1 is a flow chart of an ultra-dense heterogeneous wireless network access control method provided by an embodiment of the present invention, wherein the ultra-dense heterogeneous wireless network access control method comprises steps S1 to S6: S1. Construct the network state space according to the attribute parameters of the candidate network; It can be understood that the embodiment of the present invention first constructs the deep reinforcement learning model elements in an ultra-dense heterogeneous wireless network access environment. The construction of the deep reinforcement learning model elements mainly includes three parts: state space definition, action space definition and reward function design.
[0021] It should be noted that in deep reinforcement learning, the state space (State) refers to the set of all states that the agent (agent) may be in the environment. In the embodiment of the present invention, the state is a complete description of the ultra-dense heterogeneous wireless network access environment, which contains all the relevant information needed for the agent to make decisions.
[0022] In an optional embodiment, the attribute parameters of the candidate network include user personalized demand parameters and network status information parameters; The user personalized demand parameters include the price, bandwidth and received signal strength of the candidate network; The network status information parameters include the load rate, bit error rate and blocking rate of the candidate network.
[0023] It should be noted that, in a specific implementation, the state space includes not only the attribute parameters of the candidate networks, but also the service type currently initiated by the user and the currently selected network.
[0024] It is worth noting that among the attribute parameters of the candidate networks, price belongs to cost attributes, bandwidth and received signal strength belong to system performance attributes, bit error rate, blocking rate and load rate belong to stability attributes. These parameters cover the key aspects of network performance and can meet the different emphases of different users on network needs.
[0025] S2, constructing an action space based on the candidate networks that can be selected for access; It should be noted that in deep reinforcement learning, the action space refers to the set of all possible actions that an agent can take in a given state. The action is the way the agent changes the state of the environment, and the action space defines the scope of these behaviors. In the embodiment of the present invention, assuming that the user is in the coverage of multiple networks and can only choose one candidate network to access at the decision time, the action space is defined as the set of candidate networks that the user can choose at the switching decision time.
[0026] S3. Based on the psychological curve function, a user satisfaction model is constructed in combination with the user's expected benefit and the current benefit of the candidate network, and a reward function is constructed using the user satisfaction model; It should be noted that the reward function is a key component in the deep reinforcement learning model. It is a feedback mechanism of the environment on the behavior of the agent. The reward function defines the immediate reward value obtained by the agent after taking a certain action in a certain state. This reward value is used to measure whether the behavior of the agent is "good" or "bad" and guide the agent to develop towards the desired target behavior.
[0027] In the embodiment of the present invention, at each decision moment, the mobile terminal will make a switching decision, select a target network based on the current network status, and obtain a feedback reward value. However, the existing heterogeneous wireless network access selection algorithm only selects a network with the best comprehensive performance for the user among all candidate networks, without fully considering various factors such as network performance, user business requirements, and user preferences, and it is impossible to connect the user to the most suitable network. Therefore, in order to better characterize the personalized needs and satisfaction of users, the embodiment of the present invention introduces a psychological curve function, considers the user's expected benefits and the current benefits of the candidate networks, constructs a user satisfaction model, and then uses the user satisfaction model to construct a reward function.
[0028] In an optional embodiment, the user satisfaction model is constructed based on the psychological curve function and in combination with the user's expected benefit and the current benefit of the candidate network, including: Normalizing the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter, and calculating the current benefit of the candidate network according to the normalized value of each attribute parameter; By analyzing the user's historical network access data, the user's satisfaction with each attribute parameter is calculated to obtain the user's expected benefits; Based on the psychological curve function, a user satisfaction model is constructed according to the expected benefits of the user and the current benefits of the candidate networks.
[0029] It is worth noting that, in the embodiment of the present invention, six parameters, namely price, bandwidth, received signal strength, load rate, bit error rate, and blocking rate, are selected as the construction parameters of the user satisfaction model. In order to solve the differences in the value ranges and units of each parameter in different wireless networks, the above six parameters are normalized respectively. Among them, the attribute parameters of the network can be divided into benefit type and cost type. The smaller the cost type parameters, the better, such as price, bit error rate, blocking rate, and load rate; while the larger the benefit type parameters, the better, such as received signal strength and bandwidth.
[0030] Specifically, normalizing the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter, and calculating the current benefit of the candidate network according to the normalized value of each attribute parameter includes: Normalizing the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter; The attribute parameters of the candidate network are divided into three categories: cost, system performance and stability. Among them, the cost attribute parameters include the price of the candidate network, the system performance attribute parameters include the bandwidth and received signal strength of the candidate network, and the stability attribute parameters include the load rate, bit error rate and blocking rate of the candidate network. The fuzzy hierarchical analysis method is used to calculate the weight value of each parameter in the system performance attribute parameter and stability attribute parameter respectively; The normalized value of the price is defined as the current benefit of the candidate network's cost. The current benefit of the candidate network's system performance is calculated based on the normalized values and weight values of the bandwidth and received signal strength. The current benefit of the candidate network's stability is calculated based on the normalized values and weight values of the load rate, bit error rate, and blocking rate.
[0031] Exemplarily, the normalized value of the price is defined as the comprehensive value of the cost attribute to obtain the cost current benefit of the candidate network. The comprehensive value of the system performance attribute and the stability attribute is the product and accumulation of the weight and the normalized value of the corresponding parameter to obtain the system performance current benefit of the candidate network and the stability current benefit of the candidate network respectively.
[0032] It should be noted that, in order to obtain the weights of different attributes, the embodiment of the present invention adopts the fuzzy analytic hierarchy process (FAHP) to respectively calculate the weights of bandwidth and received signal strength in system performance and the weights of bit error rate, blocking rate and load rate in stability.
[0033] Fuzzy analytic hierarchy process is a multi-criteria decision-making method that combines the traditional analytic hierarchy process (AHP) and fuzzy set theory. In practical problems, the evaluation and comparison of many attribute parameters are often ambiguous. For example, a user may describe "this wireless network signal looks good". "Good" is a fuzzy expression that does not specifically indicate the signal strength in dBm. "Good" may represent different actual signal strength ranges in the cognition of different users. Perhaps for a user with low requirements, -70dBm is "good", but for users who need to conduct high-definition video conferencing and have high requirements for network quality, -60dBm or less is considered "good". The fuzzy analytic hierarchy process can be used to handle this kind of fuzzy information well. By constructing a fuzzy judgment matrix, people's fuzzy evaluation can be converted into a form that can be quantified and calculated, thereby more accurately reflecting the actual situation.
[0034] In addition, the fuzzy hierarchical analysis method retains the hierarchical structure characteristics of the traditional hierarchical analysis method. Taking the selection of heterogeneous wireless networks in the embodiment of the present invention as an example, the factors affecting the selection can be divided into three levels, each of which corresponds to the target layer, the criterion layer, and the solution layer. The target layer is the user benefits of system performance and stability; the criterion layer is the various decision factors that affect system performance and stability; the solution layer is the normalized value of each parameter in each candidate network. This hierarchical structure enables complex multi-attribute decision-making problems to be systematically decomposed, making it convenient for decision makers to analyze problems from different levels and angles, clarify the relationship between various attribute parameters, and avoid considering only a single factor while ignoring other important factors.
[0035] Specifically, the psychological curve function is: ; in, For user satisfaction; Expected benefits for users; is the current revenue of the candidate network; The user satisfaction model is: ; ; ; in, Indicates selection A network, Indicates the status of the network. Indicates the action taken; is the cost satisfaction model, Indicates in status Take action The expected cost benefit after Indicates The network is in state Take action Costs after current benefits; is the system performance satisfaction model, Indicates in status Take action The expected cost benefit after Indicates The network is in state Take action The current benefits of system performance after is the stability satisfaction model, Indicates in status Take action The expected return after stability is Indicates The network is in state Take action The stability of current earnings after.
[0036] Furthermore, the reward function is: ; in, , and They are the weight values of cost attribute parameters, system performance attribute parameters and stability attribute parameters respectively.
[0037] It should be noted that the weight calculation of the three major attributes of cost, system performance and stability can also be calculated by constructing a fuzzy complementary judgment matrix using the fuzzy hierarchical analysis method.
[0038] S4, constructing a deep reinforcement learning neural network model in combination with a back propagation neural network according to the network state space, the action space and the reward function; It is worth noting that the present invention adopts Q learning to construct a deep reinforcement learning model, and then adopts a BP (Back Propagation) neural network, that is, a back propagation neural network, to approximate the state-action value function in Q learning. The BP neural network uses an ε-greedy strategy to select the target network to switch based on the Q value.
[0039] Exemplarily, the input layer of the neural network determines the number of nodes according to the dimension of the state space, which is used to receive state information; the number of hidden layers and the number of nodes are determined according to the complexity of the problem and resources, and a multilayer perceptron with a nonlinear activation function is used to increase the expressiveness; the number of nodes in the output layer depends on the action space, and the discrete action space uses Softmax to output the action probability distribution, while the continuous action space is designed according to the specific action representation; the state information is input into the input layer through forward propagation, and is transferred to the output layer through hidden layer calculation to obtain the action probability distribution or the specific action value. According to the target design, different access results are given corresponding rewards, and the loss function is constructed based on the reward function. The gradient is calculated starting from the output layer, and the weights are updated through reverse transmission, and the learning rate is controlled by the optimization algorithm.
[0040] S5. Training the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model; In an optional embodiment, the deep reinforcement learning neural network model is trained by an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model, including: Initializing the parameters of the deep reinforcement learning neural network model through an evolutionary strategy algorithm; The training dataset is constructed using historical handover decision data of ultra-dense heterogeneous wireless networks; Setting a main network and a target network for the deep reinforcement learning neural network model; The training data set is used to train the main network and the target network respectively, and the parameters of the main network are updated. The training is stopped when a preset training condition is reached to obtain an ultra-dense heterogeneous wireless network access control model.
[0041] It should be noted that the deep reinforcement learning neural network model includes a main network and a target network. The main network is constructed using a BP neural network, and the target network has the same structure as the main network and the same initial weights. The difference is that the main network is updated every iteration while the target network is updated every once in a while. The main network selects access to the network through the ε-greedy strategy, while the target network selects access to the network through the greedy strategy.
[0042] S6. Use the ultra-dense heterogeneous wireless network access control model to execute a switching strategy for the ultra-dense heterogeneous wireless network.
[0043] For example, in a modern smart office park, various types of wireless network access points are deployed in the park, including traditional Wi-Fi 6 access points, 5G micro base stations and future 6G experimental base stations, forming an ultra-dense heterogeneous wireless network environment. The park is densely populated and has many mobile office devices, such as smart phones, tablets, laptops, etc. These devices need to switch seamlessly between different networks to ensure business continuity and efficient operation.
[0044] First, the ultra-dense heterogeneous wireless network access control model continuously monitors the status information of mobile devices. For example, when an employee walks in the campus with a smartphone, it is detected that the signal strength of the Wi-Fi 6 access point to which the phone is currently connected gradually weakens, the packet loss rate increases, and the phone moves to the coverage area of the 5G micro base station. At the same time, a video conferencing application is running on the phone. Then, based on the collected status information, decisions are made according to the predefined switching strategy. In the above example, combined with the download task (high bandwidth requirement) currently being executed by the laptop, the network switching program is started according to the switching strategy, and the target switching network is the 5G micro base station with the strongest signal nearby.
[0045] Such an ultra-dense heterogeneous wireless network access control model and its switching strategy can effectively cope with complex and changing network environments, meet the network connection needs of different users in different scenarios, and improve the overall utilization efficiency of ultra-dense heterogeneous wireless networks.
[0046] In summary, an ultra-dense heterogeneous wireless network access control method provided by an embodiment of the present invention fully considers the personalized needs of users and comprehensively considers the attribute parameters of the user side and the network side, and constructs a more objective user satisfaction model for quantifying user satisfaction, so that the network selection strategy can dynamically adapt to changes in network status and user needs, and realize the best selection access of the network; at the same time, the present invention introduces deep reinforcement learning into the access selection problem of heterogeneous wireless networks, models the network selection access problem as a reinforcement learning model, and can learn the optimal strategy through interaction with the environment, and intelligently make access selection decisions.
[0047] In order to make those skilled in the art more aware of the implementation process of the ultra-dense heterogeneous wireless network access control method provided by the embodiment of the present invention, a more specific description is given below.
[0048] First of all, the construction of deep reinforcement learning model elements mainly includes three parts: state space definition, action space definition and reward function design. The above three elements are defined as follows: 1. State Space For example, the state space considered includes network parameters M, the service type G currently initiated by the user, the currently selected network .
[0049] Then at the decision time t, the state space It can be expressed as a set of formula (1); #(1)
[0050] 2. Action Space Assuming that the user is in the coverage of multiple networks and can only choose one network to access at the decision time t, the action space A can be expressed as the set of formula (2): , is the action taken by the user at the switching decision time t , represents the i-th candidate network.
[0051] #(2)
[0052] 3. Reward function At each decision time t, the mobile terminal will make a switching decision. Based on the current network status, a target network is selected and a feedback reward value is obtained. In order to better characterize the personalized needs and satisfaction of users, the embodiment of the present invention introduces a psychological curve function, considers the expected benefits of users and the current benefits of candidate networks, constructs a user satisfaction model, and uses the user satisfaction model to construct a reward function.
[0053] Among them, the psychological curve function is expressed as: #(3); In formula (3) For user satisfaction; Expected benefits for users; is the current revenue of the candidate network.
[0054] Furthermore, considering the user's expected benefits and the current benefits of the candidate network, a user satisfaction model is constructed. The process is as follows: (1) Data normalization Six wireless network parameters, including price P, bandwidth B, received signal strength RSS, load rate F, bit error rate D, and blocking rate Z, are selected as the construction parameters of the user satisfaction model. In order to solve the differences in the value range and unit of each parameter in different wireless networks, the above six parameters are normalized respectively. Among them, the smaller the cost parameters, the better, such as price, bit error rate, blocking rate, and load rate; the larger the benefit parameters, the better, such as received signal strength and bandwidth.
[0055] The normalized values of cost-effectiveness and benefit-effectiveness parameters are as shown in formula (4) and (5): #(4); #(5); in, For the Network parameters The normalized value of It is used to represent the six parameters of price P, bandwidth B, received signal strength RSS, bit error rate D, blocking rate Z and load rate F respectively. are the parameters of all candidate networks A set of values of For the network Parameters The numerical value of .
[0056] (2) Calculate the weight value using fuzzy analytic hierarchy process The system performance attribute is composed of bandwidth B and received signal strength RSS, and the stability attribute is composed of bit error rate D, blocking rate Z, and load rate F. In order to obtain the weights of different attributes, the embodiment of the present invention uses fuzzy hierarchical analysis method to calculate the weights of bandwidth B and received signal strength RSS in system performance, and the weights of bit error rate D, blocking rate Z, and load rate F in stability.
[0057] Specifically, first, establish a hierarchical analysis structure diagram, see Figure 2 , Figure 2 This is a hierarchical analysis structure diagram provided by an embodiment of the present invention. Figure 2 As shown in the figure, there are three levels, corresponding to the target layer, criterion layer and solution layer. The target layer is the user benefits of system performance and stability; the criterion layer is the various decision factors that affect system performance and stability; the solution layer is the normalized value of each parameter in each candidate network.
[0058] Secondly, for the system performance attributes and stability attributes, the system performance fuzzy complementary judgment matrix is constructed and stability fuzzy complementary judgment matrix .
[0059] in, , It is obtained by the 0.1-0.9 scaling method, as shown in Table 1. =1, + =1.
[0060] Table 1: 0.1-0.9 Scaling Method The weight calculation formula of the fuzzy complementary judgment matrix is as follows: #(6); The weights of each parameter are calculated using the above fuzzy complementary judgment matrix weights: , ; in, , are the weight vectors of system performance and stability, respectively. , , , , They are the weights of received signal strength RSS, bandwidth B, bit error rate D, blocking rate Z, and load rate F respectively.
[0061] Finally, the consistency test of the comparison judgment is performed. First, the fuzzy matrix is constructed respectively. , The characteristic matrix of ; in, . Represents the weight of the a-th row and b-th column.
[0062] Use formula (7) to solve the compatibility value: #(7); like , then it is determined that the consistency meets the standard and the weight value result is reasonable, otherwise it is necessary to adjust the fuzzy complementary judgment matrix and re-solve the weight.
[0063] (3) Calculate current income The normalized value of the price is defined as the comprehensive value of the cost attribute, and the comprehensive value of the system performance attribute and the stability attribute is the product and accumulation of the weight and the normalized value of the corresponding parameter.
[0064] #(8); #(9); #(10); In formulas (8)-(10), , , Respectively The comprehensive value of the network cost pc, system performance sp, and stability st.
[0065] The comprehensive value is defined as the current benefit, which is based on the real-time data of the current candidate network, including the Cost, system performance, stability, and current benefits of each network , , .in, Indicates The network is in state Take action After the cost of current benefits, Indicates The network is in state Take action The current system performance gain after Indicates The network is in state Take action The stability of current earnings after.
[0066] (4) Calculate expected returns By analyzing the user's historical access data, the user's satisfaction with each attribute in historical situations is calculated, and the user's expected benefit for each attribute is obtained.
[0067] The expected return calculation formula is as follows: #(11); #(12); #(13); in, , , In the The comprehensive value of cost, system performance, and stability under a network, h is the number of overlapping network situations, is the average value of the comprehensive value of users' demand for network attributes under the historical situation of h overlaps, which is defined as the expected benefit; Indicates in status Take action The expected benefit after the cost is Indicates in status Take action The expected benefit of the system performance after Indicates in status Take action The expected return after stability.
[0068] (5) Specific satisfaction calculation Based on the psychological curve function and combining the expected benefits and current benefits of users in the three major attributes, a satisfaction model for the three major attributes of the candidate network is proposed.
[0069] #(14); #(15); #(16); Equations (14)-(16) are the cost, system performance, and stability satisfaction models, respectively.
[0070] Furthermore, the reward function is defined as follows: #(17); in , , are the weight values of the three major attributes respectively. A fuzzy complementary judgment matrix is constructed for the three major attributes of cost, system performance and stability. The calculation process is the same as the above-mentioned fuzzy hierarchical analysis method to calculate the weight value.
[0071] Secondly, the BP (Back Propagation) neural network and the deep Q network DQN are integrated, and the BP neural network is used to approximate the state-action value function. The following process is included: Specifically, the initial weight parameters of the BP neural network are obtained through evolutionary strategies (ES) with global optimization capabilities, and then the parameters of the neural network are trained and adjusted through the gradient descent method.
[0072] 1. Main network modeling: See also Figure 3 , Figure 3 FIG. 1 is a schematic diagram of a main network structure provided by an embodiment of the present invention. Figure 3 As shown, in the embodiment of the present invention, a BP neural network is used to construct a main network, and the main network consists of an input layer, two hidden layers and an output layer.
[0073] Each state-action pair corresponds to a value function , using the main network to approximate the state action value function, it can be expressed as: #(18); in, Indicates the network status hour, Is the parameter of the main network, taking switching action The expected long-term cumulative return, is the switching strategy adopted by the switch, It can be defined as: #(19); #(20); in, Representative in Strategy The expectations below, To accumulate discount rewards, is the discount factor (0≤ ≤1), indicating the importance of future rewards, represent Moment of reward.
[0074] From the attribute parameters of each network and the number of all candidate networks, we know that the number of neurons in the input layer is 6n+2. After the neural network forward propagates, the output of all state-action pairs in the current state is value: , then the number of neurons in the output layer is n. Based on the ε-greedy strategy, according to the output The value selects the target network to switch to. At this time, based on the current network status and the switching decision made , the end user will get an immediate feedback reward value r, and the network state will also change to the next state The switching decision data generated by this interaction process Stored in the training database, used to update the main network, and then enter the next step, and continue to cycle until the BP network converges.
[0075] 2. Target network modeling: The target network has the same structure as the main network, and the initial weights are the same. The difference is that the main network is updated every iteration, while the target network is updated every other period. The parameters of the target network are The input of the target network is , the output is The target network uses the greedy strategy to select the switching action. , is to make The action with the highest value.
[0076] See also Figure 4 , Figure 4 It is a structural block diagram of a deep reinforcement learning neural network model provided by an embodiment of the present invention.
[0077] like Figure 4 As shown, in heterogeneous wireless networks, the algorithm is based on the current network status Select an action (i.e., choose to switch to a specific network). Using BP neural network as the main network, it receives the current state of heterogeneous wireless networks. and actions As input, output , which is the expected return of performing the action in the current state. Every N steps, the parameters of the main network are copied to the target network. The target network, which has the same structure as the main network but has a lower parameter update frequency, is used to calculate the target Value By comparison And to calculate the loss, use the gradient descent method to guide the update of the main network parameters to improve The training database stores the data generated during the interaction between the agent and the environment. , where r is the action to be performed After receiving the instant reward, is the next state.
[0078] 3. Establish training database: Specifically, an experience replay mechanism can be introduced to store historical decision data obtained by the interaction between the terminal and the environment in the database at each decision moment for repeated learning and training of the neural network. During the training process, small batches of training samples are randomly extracted from the database, which reduces the correlation between samples and improves the stability of the algorithm.
[0079] The main network selects the access network through the ε-greedy strategy, while the target network selects the access network through the greedy strategy. In DQN, the ε-greedy strategy is used to balance exploration and exploitation. The process of selecting the current optimal action is exploitation, and the process of selecting other non-optimal actions is exploration. That is, one is randomly selected from all candidate networks with a probability of ε, and one is selected with a probability of 1-ε. The largest value is connected to the network. When the neural network converges, the When the values can be accurately evaluated, the exploration and utilization should be completely inclined to utilization. Therefore, as the number of training times increases, the value of 1-ε can be gradually increased. The greedy strategy is a greedy algorithm that chooses to make The access network with the largest value, that is, .
[0080] 4. Training the main network 4.1 Initialize the neural network parameters of the main network It should be noted that the Evolution Strategy (ES) algorithm is an optimization algorithm based on the idea of evolution in nature. It optimizes mathematical functions through iterative search and is mainly used to solve parameter optimization problems. This algorithm is a black-box optimization algorithm. It is a heuristic search process inspired by natural evolution: in each iteration, the parameter vector is randomly perturbed. Subsequently, the algorithm evaluates its objective function value, i.e., fitness, and then recombine the parameter vectors with the highest objective function value to form the next generation of vector groups (populations), and iterates this process until the target is fully optimized.
[0081] It is worth noting that the evolutionary strategy (ES) is used to pre-train the neural network to obtain a relatively optimal initial neural network parameter. is the optimization goal. The population distribution of each generation follows the mean , the variance is isotropic multivariate Gaussian distribution of , where Represents the parameters of the main network of the kth generation. In order to simplify the optimization process, additive Gaussian noise is applied to the current parameter vector. The Gaussian noise vector is always evaluated in pairs using a mirroring method: and .
[0082] The offspring parameters can be expressed as: #(twenty one); in, , Represents the kth generation Parameters for the children.
[0083] In the evolution strategy (ES), the fitness of each offspring individual is evaluated by placing it in a heterogeneous wireless network environment. The fitness function is expressed as In order to improve the neural network parameters The embodiment of the present invention adopts the stochastic gradient ascent method to maximize the average reward of the group, and its mathematical expression is as follows: #(twenty two); #(twenty three); In formula (23), Represents the learning rate, and n refers to half of the number of offspring. Through repeated iterations, when the accumulated reward value gradually reaches a stable state, the obtained weight parameters will be used to initialize the parameter settings of the main network.
[0084] 4.2 Main network parameter update In order to optimize the switching decision, the user terminal needs to continuously interact with the heterogeneous network environment and adjust the switching strategy according to the feedback rewards provided by the environment, that is, to optimize the weight parameters of the neural network through training, so as to continuously strengthen the terminal's switching decision-making ability. The training of the neural network is essentially the process of minimizing the loss function, which reflects the deviation between the model prediction and the actual label.
[0085] The widely used mean square error model is used to construct the loss function, and the error back propagation and gradient descent method are used to iteratively solve the problem to minimize the loss function. The loss function of the ES-DQN (deep reinforcement learning neural network model) in the embodiment of the present invention is defined as follows: #(twenty four); in, is an estimated value, Output value for the main network, is the discount factor.
[0086] 4.3 Circuit Training See also Figure 5 , Figure 5 FIG. 1 is another flow chart of an ultra-dense heterogeneous wireless network access control method provided by an embodiment of the present invention. Figure 5 As shown in the figure, after building the reinforcement learning model, the BP neural network is used to approximate the state-action-value function in Q learning, the initial parameters of the neural network are pre-trained through the evolutionary strategy (ES), and a training data set is constructed as a training sample of the neural network. Then, the main network is set to approximate the current state-action-value function and the parameters are continuously updated. The target network is set to calculate the Q reality, and the parameters are regularly copied from the main network. A batch of samples are randomly selected from the training database to start cyclic training, and the parameters of the main network are updated to check whether the preset number of training rounds or convergence conditions are reached. If reached, the training ends and the switching decision is executed. Otherwise, the training cycle continues.
[0087] In summary, the embodiment of the present invention firstly constructs a network state space according to the attribute parameters of the candidate network, so as to comprehensively and meticulously describe the state of the ultra-dense heterogeneous wireless network, accurately capture the real-time state of each candidate network, and provide detailed network status information for subsequent access control and switching strategies, which is helpful to better adapt to the complex and changeable characteristics of the network state in the ultra-dense heterogeneous network environment; and, constructs a user satisfaction model based on the psychological curve function, and constructs a reward function based on this, which is an innovative way to consider the user's subjective feelings. The user satisfaction model combines the user's expected benefits and the current benefits of the candidate network, and can more accurately reflect the user's true evaluation of the network service. This reward mechanism enables the access control strategy to better meet the user's needs, optimize network access and switching decisions, and improve the user experience; then, using the network state space, action space and reward function, combined with the back propagation neural network to construct a deep reinforcement learning neural network model, it can effectively learn the optimal strategy for network access and switching. This automatic learning ability can adapt to the complexity and dynamics of ultra-dense heterogeneous networks, reduce manual intervention, and improve the intelligence level of access control; at the same time, the deep reinforcement learning neural network model is trained by the evolutionary strategy algorithm to further optimize the performance of the model. The evolutionary strategy algorithm can search for the optimal solution in a complex parameter space, and it can evaluate and optimize the parameters of the model according to the fitness function. Compared with the traditional training method, this algorithm can better handle the high-dimensional parameters and complex optimization objectives in the ultra-dense heterogeneous network access control model, improve the convergence speed and accuracy of the model, and thus obtain a more effective ultra-dense heterogeneous wireless network access control model; finally, the ultra-dense heterogeneous wireless network access control model is used to execute the network switching strategy, which can achieve smarter and more efficient network switching and improve the user experience in complex network environments.
[0088] On the basis of the above-mentioned method items, the present invention provides corresponding embodiments of the system items.
[0089] See also Figure 6 , Figure 6 The structure diagram of an ultra-dense heterogeneous wireless network access control system provided by an embodiment of the present invention is as follows. The ultra-dense heterogeneous wireless network access control system comprises: A first model element construction module 21, used to construct a network state space according to attribute parameters of the candidate network; A second model element construction module 22, for constructing an action space according to the candidate networks that can be selected for access; The third model element construction module 23 is used to construct a user satisfaction model based on the psychological curve function, combined with the user's expected benefits and the current benefits of the candidate network, and use the user satisfaction model to construct a reward function; A deep reinforcement learning neural network model construction module 24, used to construct a deep reinforcement learning neural network model in combination with a back propagation neural network according to the network state space, the action space and the reward function; A deep reinforcement learning neural network model training module 25 is used to train the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model; The network access control module 26 is used to implement the switching strategy of the ultra-dense heterogeneous wireless network by adopting the ultra-dense heterogeneous wireless network access control model.
[0090] In an optional embodiment, the attribute parameters of the candidate network include user personalized demand parameters and network status information parameters; The user personalized demand parameters include the price, bandwidth and received signal strength of the candidate network; The network status information parameters include the load rate, bit error rate and blocking rate of the candidate network.
[0091] In an optional embodiment, the third model element construction module 23 includes a user satisfaction model construction unit and a reward function construction unit; Wherein, the user satisfaction model building unit is specifically used for: Normalizing the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter, and calculating the current benefit of the candidate network according to the normalized value of each attribute parameter; By analyzing the user's historical network access data, the user's satisfaction with each attribute parameter is calculated to obtain the user's expected benefits; Based on the psychological curve function, a user satisfaction model is constructed according to the expected benefits of the user and the current benefits of the candidate networks.
[0092] Specifically, normalizing the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter, and calculating the current benefit of the candidate network according to the normalized value of each attribute parameter includes: Normalizing the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter; The attribute parameters of the candidate network are divided into three categories: cost, system performance and stability. Among them, the cost attribute parameters include the price of the candidate network, the system performance attribute parameters include the bandwidth and received signal strength of the candidate network, and the stability attribute parameters include the load rate, bit error rate and blocking rate of the candidate network. The fuzzy hierarchical analysis method is used to calculate the weight value of each parameter in the system performance attribute parameter and stability attribute parameter respectively; The normalized value of the price is defined as the current benefit of the candidate network's cost. The current benefit of the candidate network's system performance is calculated based on the normalized values and weight values of the bandwidth and received signal strength. The current benefit of the candidate network's stability is calculated based on the normalized values and weight values of the load rate, bit error rate, and blocking rate.
[0093] In an optional embodiment, the psychological curve function is: ; in, For user satisfaction; Expected benefits for users; is the current revenue of the candidate network; The user satisfaction model is: ; ; ; in, Indicates selection A network, Indicates the status of the network. Indicates the action taken; is the cost satisfaction model, Indicates in status Take action The expected cost benefit after Indicates The network is in state Take action Costs after current benefits; is the system performance satisfaction model, Indicates in status Take action The expected cost benefit after Indicates The network is in state Take action The current benefits of system performance after is the stability satisfaction model, Indicates in status Take action The expected return after stability is Indicates The network is in state Take action The stability of current earnings after.
[0094] Furthermore, the reward function is: ; in, , and They are the weight values of cost attribute parameters, system performance attribute parameters and stability attribute parameters respectively.
[0095] It should be noted that an ultra-dense heterogeneous wireless network access control system provided in an embodiment of the present invention is used to execute all process steps of an ultra-dense heterogeneous wireless network access control method in the above embodiment, and the working principles and beneficial effects of the two correspond one to one, so they will not be repeated.
[0096] The embodiment of the present invention also provides a terminal device, such as Figure 7 , which is a block diagram of a terminal device provided by an embodiment of the present invention. The terminal device includes a processor 31, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 31. When the processor 31 executes the computer program, the ultra-dense heterogeneous wireless network access control method described in any of the above embodiments is implemented.
[0097] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the ultra-dense heterogeneous wireless network access control method as described in any of the above embodiments.
[0098] When the processor 31 executes the computer program, the steps in the above-mentioned ultra-dense heterogeneous wireless network access control method embodiment are implemented, for example: Figure 1 Alternatively, when the processor 31 executes the computer program, the functions of each module in the above-mentioned ultra-dense heterogeneous wireless network access control system embodiment are implemented, for example Figure 6 The functions of each module of the ultra-dense heterogeneous wireless network access control system are shown.
[0099] Preferably, the computer program can be divided into one or more modules / units, which are stored in the memory 32 and executed by the processor 31 to implement the present invention. The one or more modules / units can be a series of computer program instruction segments that can implement specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.
[0100] The processor 31 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor 31 may be any conventional processor. The processor 31 is the control center of the terminal device, and various parts of the terminal device are connected using various interfaces and lines.
[0101] The memory 32 mainly includes a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function, etc., and the data storage area can store related data, etc. In addition, the memory 32 can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, and a flash card (Flash Card), etc., or the memory 32 can also be other volatile solid-state storage devices.
[0102] It should be noted that the above terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 7 The structural block diagram shown is only a structural example of the above-mentioned terminal device and does not constitute a structural limitation of the above-mentioned terminal device. The above-mentioned terminal device may include more or fewer components than shown in the figure, or combine certain components, or different components.
[0103] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for controlling access to an ultra-dense heterogeneous wireless network, characterized in that: include: Constructing a network state space based on the attribute parameters of the candidate network; Construct an action space based on the candidate networks that can be selected for access; Based on the psychological curve function, a user satisfaction model is constructed by combining the user's expected benefit and the current benefit of the candidate network, and a reward function is constructed using the user satisfaction model; According to the network state space, the action space and the reward function, a deep reinforcement learning neural network model is constructed in combination with a back propagation neural network; The deep reinforcement learning neural network model is trained by an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model; The ultra-dense heterogeneous wireless network access control model is adopted to execute the switching strategy of the ultra-dense heterogeneous wireless network.
2. The ultra-dense heterogeneous wireless network access control method according to claim 1, characterized in that: The attribute parameters of the candidate network include user personalized demand parameters and network status information parameters; The user personalized demand parameters include the price, bandwidth and received signal strength of the candidate network; The network status information parameters include the load rate, bit error rate and blocking rate of the candidate network.
3. The ultra-dense heterogeneous wireless network access control method according to claim 2, characterized in that: The method of constructing a user satisfaction model based on the psychological curve function according to the user's expected benefits and the current benefits of the candidate networks includes: Normalizing the attribute parameters of the candidate network respectively to obtain the normalized value of each attribute parameter, and calculating the current benefit of the candidate network according to the normalized value of each attribute parameter; By analyzing the user's historical network access data, the user's satisfaction with each attribute parameter is calculated to obtain the user's expected benefits; Based on the psychological curve function, a user satisfaction model is constructed according to the expected benefits of the user and the current benefits of the candidate networks.
4. The ultra-dense heterogeneous wireless network access control method according to claim 3, characterized in that: The normalizing the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter, and calculating the current benefit of the candidate network according to the normalized value of each attribute parameter, includes: Normalizing the attribute parameters of the candidate networks respectively to obtain the normalized value of each attribute parameter; The attribute parameters of the candidate network are divided into three categories: cost, system performance and stability. Among them, the cost attribute parameters include the price of the candidate network, the system performance attribute parameters include the bandwidth and received signal strength of the candidate network, and the stability attribute parameters include the load rate, bit error rate and blocking rate of the candidate network. The fuzzy hierarchical analysis method is used to calculate the weight value of each parameter in the system performance attribute parameter and stability attribute parameter respectively; The normalized value of the price is defined as the current benefit of the candidate network's cost. The current benefit of the candidate network's system performance is calculated based on the normalized values and weight values of the bandwidth and received signal strength. The current benefit of the candidate network's stability is calculated based on the normalized values and weight values of the load rate, bit error rate, and blocking rate.
5. The ultra-dense heterogeneous wireless network access control method according to claim 4, characterized in that: The psychological curve function is: ; in, For user satisfaction; Expected benefits for users; is the current revenue of the candidate network; The user satisfaction model is: ; ; ; in, Indicates selection A network, Indicates the status of the network. Indicates the action taken; is the cost satisfaction model, Indicates in status Take action The expected benefit after the cost is Indicates The network is in state Take action Costs after current benefits; is the system performance satisfaction model, Indicates in status Take action The expected benefit after the cost is Indicates The network is in state Take action The current benefits of system performance after is the stability satisfaction model, Indicates in status Take action The expected return after stability is Indicates The network is in state Take action The stability of current earnings after.
6. The ultra-dense heterogeneous wireless network access control method according to claim 5, characterized in that: The reward function is: ; in, , and They are the weight values of cost attribute parameters, system performance attribute parameters and stability attribute parameters respectively.
7. The ultra-dense heterogeneous wireless network access control method according to claim 1, characterized in that: The deep reinforcement learning neural network model is trained by an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model, including: Initializing the parameters of the deep reinforcement learning neural network model through an evolutionary strategy algorithm; The training dataset is constructed using historical handover decision data of ultra-dense heterogeneous wireless networks; Setting a main network and a target network for the deep reinforcement learning neural network model; The training data set is used to train the main network and the target network respectively, and the parameters of the main network are updated. When a preset training condition is reached, the training is stopped to obtain an ultra-dense heterogeneous wireless network access control model.
8. An ultra-dense heterogeneous wireless network access control system, characterized in that: include: A first model element construction module is used to construct a network state space according to attribute parameters of the candidate network; A second model element construction module is used to construct an action space based on the candidate networks that can be selected for access; The third model element construction module is used to construct a user satisfaction model based on the psychological curve function, combined with the user's expected benefits and the current benefits of the candidate network, and use the user satisfaction model to construct a reward function; A deep reinforcement learning neural network model construction module, used to construct a deep reinforcement learning neural network model in combination with a back propagation neural network according to the network state space, the action space and the reward function; A deep reinforcement learning neural network model training module, used to train the deep reinforcement learning neural network model through an evolutionary strategy algorithm to obtain an ultra-dense heterogeneous wireless network access control model; The network access control module is used to implement a switching strategy of the ultra-dense heterogeneous wireless network by adopting the ultra-dense heterogeneous wireless network access control model.
9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the ultra-dense heterogeneous wireless network access control method as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the ultra-dense heterogeneous wireless network access control method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Deep reinforcement learning-based heterogeneous cellular network joint optimization method
CN108848561A
Network selection method based on neural network in ultra-dense heterogeneous wireless network
CN113242584A
5G ultra-dense network multi-user access selection method based on deep reinforcement learning
CN114449536A
Edge computing unloading method based on dynamic user satisfaction in ultra-dense network
CN114641076A
Automated reinforcement-learning-based application manager that learns and improves a reward function
US20200065157A1