Ultra-dense heterogeneous network access method, apparatus and device, and storage medium
Through the deep Q network combining the reward value calculation model of utility functions and preference vectors, the optimal access strategy for super-intensive heterogeneous networks is dynamically selected, which solves the problem of inability to adapt to network performance changes in the existing technology and improves network access quality and service quality.
Patent Information
- Application Number
- CN202510584173.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-18
AI Technical Summary
The existing ultra-intensive heterogeneous network access strategy cannot dynamically adapt to network performance changes, fail to fully consider different service types and user needs, resulting in a decline in service quality and fail to give full play to the advantages of network architecture.
The reward value calculation model based on the deep Q network is adopted, combined with the utility function and the preference vector, and by calculating the internal product of the target preference vector and the utility matrix, the optimal network access strategy is dynamically selected, and factors at the network, user and service are considered.
It improves network access quality, can better adapt to complex needs in highly dynamic and highly complex network environments, and make full use of the architectural advantages of ultra-intensive heterogeneous networks.
Smart Images

Figure CN120343585A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mobile communication technologies, and in particular, to an ultra-dense heterogeneous network access method, apparatus, device, and storage medium. Background Art
[0002] With the continuous evolution of mobile communication technologies, the network architecture has become increasingly complex, and the ultra-dense heterogeneous network is one of the cutting-edge network forms. As the name implies, the ultra-dense heterogeneous network is characterized by the ultra-dense deployment of network nodes and the heterogeneity of network types. For example, in the ultra-dense heterogeneous network, an overlapping network composed of LTE (Long Term Evolution), WLAN (Wireless Local Area Network), and WiMAX (World Interoperability for Microwave Access) may be provided for users at the same time. Moreover, in this network paradigm, multiple types of base stations such as traditional macro base stations, micro base stations, and pico base stations coexist to jointly provide seamless communication services for users.
[0003] How to provide the optimal network access strategy for users is an urgent problem to be solved in the development of ultra-dense heterogeneous networks. The existing network selection strategies are mainly static network selection strategies. For example, fixedly select a certain type of network (such as WLAN), or fixedly select the network with the fastest speed or the lowest price. However, the existing network selection strategies do not consider the changes in network performance in a dynamic network environment, nor do they consider the actual network requirements of different service types or different users, which reduces the quality of service (QoS). Moreover, the single selection criterion of the static network selection strategy cannot fully utilize the architecture advantages of the ultra-dense heterogeneous network. Summary of the Invention
[0004] The objective of the embodiments of the present invention is to provide an ultra-dense heterogeneous network access method, apparatus, device, and storage medium, which can better adapt to the complex requirements in the ultra-dense heterogeneous network and provide high-quality network access in a highly dynamic and highly complex network environment.
[0005] To achieve the above objective, the embodiments of the present invention provide an ultra-dense heterogeneous network access method, including: Input the target service type and the network parameters of candidate networks; wherein, the types of the network parameters are more than 1; Based on the network parameters, call the first reward value calculation model or the second reward value calculation model to calculate the reward values of the candidate networks; Select the candidate network with the highest reward value as the optimal network, and connect the user terminal to the optimal network; Among them, the first reward value calculation model calculates the reward value through the following steps: Based on the target service type, calculate the target preference vector, and select the target utility function for each of the network parameters from the utility function set; Call the target utility function to calculate the utility values of each of the network parameters, and construct a utility matrix; Calculate the inner product of the target preference vector and the utility matrix to obtain the reward values of each of the candidate networks.
[0006] As an improvement to the above solution, the utility function set includes S-type utility functions, exponential utility functions, logarithmic utility functions, linear utility functions, and linear piecewise utility functions; where: The S-type utility function is shown as follows: The exponential utility function is shown as follows: The logarithmic utility function is shown as follows: u(x) = d + e' * ln(x + f) The linear utility function is shown as follows: u(x) = gx + h The linear piecewise utility function is shown as follows: In the above formulas, u(x) represents the first utility value; x represents the network parameter; e represents the natural constant; a″, b, c, e′, f, g, and h represent coefficients; i′ and j′ represent thresholds.
[0007] As an improvement to the above solution, the step of calling the target utility function to calculate the utility values of each of the network parameters specifically includes: Input each of the network parameters into the corresponding target utility function to obtain the initial utility values of each of the network parameters; Normalize the initial utility values of each of the network parameters to obtain the utility values of each of the network parameters for constructing the utility matrix.
[0008] As an improvement to the above solution, the preference vector is calculated in the following way: Based on the target service type, compare the importance of any two of the network parameters pairwise to obtain a number of importance ratios; Summarize each of the importance ratios to obtain a fuzzy consistent matrix; Based on the fuzzy consistent matrix, a preference weight calculation formula is used to calculate the preference weights of the network parameters, so as to construct the preference vector.
[0009] As an improvement to the above solution, the second reward value calculation model is a first deep Q-network, and the first deep Q-network is pre-trained in the following manner: Initialize the weight parameters of the first deep Q-network and the weight parameters of the second deep Q-network so that the weight parameters satisfy the normal distribution; where the structures of the second deep Q-network and the first deep Q-network are the same; Iteratively train the first deep Q-network and the second deep Q-network respectively to update the weight parameters of the first deep Q-network and the weight parameters of the second deep Q-network respectively; where after the first deep Q-network completes the network selection at each time step, the current network parameters, the target access network, the reward value, and the future network parameters are stored as a set of experience samples; the second deep Q-network is trained using the experience samples; During the iterative training process, when the first parameter update condition is satisfied, gradient descent update is also performed on the first deep Q-network based on the second deep Q-network to update the weight parameters of the first deep Q-network; During the iterative training process, the weight parameters of the second deep Q-network are also updated according to the current weight parameters of the first deep Q-network based on a preset update frequency.
[0010] As an improvement to the above solution, the iterative training of the first deep Q-network and the second deep Q-network respectively includes the following steps for iterative training of the first deep Q-network: At each time step, input the current network parameters; Determine the network selection strategy based on the exploration rate. When the network selection strategy is random exploration, randomly select one of the candidate networks as the target access network. When the network selection strategy is model selection, calculate the first reward value of each candidate network based on the current network parameters, and select the candidate network with the highest first reward value as the target access network; where the exploration rate has a negative correlation with the time step; Based on the target access network, use the state transition function to calculate the future network parameters of each candidate network; Based on the current network parameters, use the first reward value calculation model to calculate the label reward value, and update the weight parameters of the first deep Q-network by calculating the loss value based on the first reward value and the label reward value; Take the future network parameters as the current network parameters for the next time step, and return to the step of network policy selection for training in the next time step until the preset iterative training end condition is met.
[0011] As an improvement to the above solution, the gradient descent update of the first deep Q network based on the second deep Q network includes: Input the current network parameters into the first deep Q network to obtain the first reward values of the candidate networks and the target access network; Based on the target access network, calculate the future network parameters using the state transition function; Based on the future network parameters and the target access network, calculate the second reward values of the candidate networks in the next time step using the second deep Q network; Based on the current network parameters, calculate the label reward value using the first reward value calculation model; Based on the label reward value, the second reward value, and the discount factor, calculate the expected cumulative reward value; Calculate the error between the expected cumulative reward value and the first reward value, and calculate the gradient according to the error; Perform gradient descent update on the first deep Q network based on the gradient.
[0012] To achieve the above object, an ultra-dense heterogeneous network access device is further provided in an embodiment of the present invention, including: a parameter input module for inputting the target service type and the network parameters of the candidate networks; wherein, the type of the network parameters is greater than 1; A reward value calculation module for calculating the reward values of the candidate networks by calling the first reward value calculation model or the second reward value calculation model based on the network parameters; A network access module for selecting the candidate network with the highest reward value as the optimal network and connecting the user terminal to the optimal network; Wherein, the first reward value calculation model calculates the reward value through the following steps: Based on the target service type, calculate the target preference vector, and select the target utility function from the utility function set for each of the network parameters; Call the target utility function to calculate the utility values of the network parameters and construct a utility matrix; Calculate the inner product of the target preference vector and the utility matrix to obtain the reward values of the candidate networks.
[0013] To achieve the above object, an embodiment of the present invention further provides a hyper-dense heterogeneous network access device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the hyper-dense heterogeneous network access method described in any of the above embodiments is implemented.
[0014] To achieve the above object, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. When the computer program runs, the device where the computer-readable storage medium is located is controlled to execute the hyper-dense heterogeneous network access method described in any of the above embodiments.
[0015] Compared with the prior art, the hyper-dense heterogeneous network access method, device, equipment, and storage medium provided by the embodiments of the present invention input the target service type and network parameters of candidate networks; wherein, the types of the network parameters are greater than 1; based on the network parameters, a first reward value calculation model or a second reward value calculation model is called to calculate the reward values of the candidate networks; the candidate network with the highest reward value is selected as the optimal network, and the user terminal is connected to the optimal network; wherein, the first reward value calculation model calculates the reward value through the following steps: based on the target service type, calculate a target preference vector, and select a target utility function from the utility function set for each of the network parameters; call the target utility function to calculate the utility values of the network parameters, and construct a utility matrix; calculate the inner product of the target preference vector and the utility matrix to obtain the reward values of the candidate networks. By comprehensively considering factors at multiple levels such as the network, users, and services, the embodiments of the present invention can improve the network access quality, thus better adapting to the complex requirements in the hyper-dense heterogeneous network and having more advantages in a highly dynamic and complex network environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flowchart of a hyper-dense heterogeneous network access method provided by an embodiment of the present invention; Figure 2 is a training flowchart of a deep Q network provided by an embodiment of the present invention; Figure 3 is a schematic structural diagram of a hyper-dense heterogeneous network access device provided by an embodiment of the present invention; Figure 4 is a schematic structural diagram of a hyper-dense heterogeneous network access equipment provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] In the description of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0019] In the description of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0020] See Figure 1 , which is a flowchart of an ultra-dense heterogeneous network access method provided by an embodiment of the present invention, including steps S1 to S3: S1. Input the target service type and the network parameters of the candidate networks; wherein, the types of the network parameters are greater than 1; S2. Based on the network parameters, call the first reward value calculation model or the second reward value calculation model to calculate the reward values of the candidate networks; S3. Select the candidate network with the highest reward value as the optimal network and connect the user terminal to the optimal network; wherein, the first reward value calculation model calculates the reward value through steps S211 to S213: S211. Based on the target service type, calculate the target preference vector and select the target utility function for each network parameter from the utility function set; S212. Call the target utility function to calculate the utility values of the network parameters and construct a utility matrix; S213. Calculate the inner product of the target preference vector and the utility matrix to obtain the reward values of the candidate networks.
[0021] Exemplarily, in step S1, the service types include voice service, video service, and data service; the network parameters include bandwidth, delay, jitter, loss rate, and price; the candidate networks include UMTS (Universal Mobile Telecommunications System), LTE, WLAN, and WiMAX; for the convenience of explanation, in the following embodiments, the network parameters and network types are both illustrated by the examples in this embodiment.
[0022] Further, in step S1, the network parameters of the candidate network UMTS at time t can be represented as a vector shown in Equation (1): Wherein, represents the network parameters of the candidate network UMTS at time t; Bandwidth represents bandwidth; Delay represents delay; Jitter represents jitter; Loss represents loss rate; Price represents price.
[0023] Further, by summarizing the network parameters of different candidate networks at time t, the state matrix (state space) at time t can be obtained, as shown in Equation (2): Wherein, s t represents the network state at time t; represents the network parameters of the candidate network UMTS at time t; represents the network parameters of the candidate network LTE at time t; represents the network parameters of the candidate network WLAN at time t; represents the network parameters of the candidate network WiMAX at time t.
[0024] It can be understood that each "element" in the state matrix s t in Equation (2) (such as ) represents the multi-dimensional network information provided by a candidate network, and the combination of these network information constitutes the state space in the radio access control problem of the ultra-dense heterogeneous network, which is used to describe the state of the ultra-dense heterogeneous network at each moment.
[0025] Furthermore, in step S2, the embodiments of the present invention provide two methods for calculating the reward value of the candidate network. One is to calculate using the first reward value calculation model, and the other is to calculate using the second reward value calculation model. In practical applications, only one of them needs to be selected. Among them, the calculation process of the first reward value calculation model includes the above steps S211 to S213; the second reward value calculation model is a Deep Q-Network (DQN), and the second reward value calculation model can achieve the same or similar input-output relationship as the first reward value calculation model, that is, after the second reward value calculation model is trained, when the same network parameters are input, it can calculate the same or similar reward value as the first reward value calculation model. It should be noted that by using the Deep Q-Network to construct the second reward value calculation model, the embodiments of the present invention can improve the reward value calculation efficiency, and thus improve the network selection efficiency.
[0026] Furthermore, since the network requirements of different service types are different, that is, different service types attach different importance to different network parameters. Therefore, referring to steps S211 to S213, the embodiments of the present invention not only consider the network parameters of each candidate network at the current moment when calculating the reward value, but also consider the influence of the target service type and user preferences. Specifically, in step S211, the target utility function is selected according to the service type to convert the current value of the network parameters into a utility value, and in step S211, the target preference vector is also calculated according to the service type; further, in step S213, the reward value of each candidate network is calculated according to the target preference vector and the utility matrix.
[0027] Compared with the prior art, the embodiments of the present invention determine the target preference vector and the target utility function based on the service type, calculate the utility matrix according to the target utility function and the current network parameters, and then calculate the reward value of each candidate network based on the utility matrix and the preference vector, which can select the network that best meets the user's expected requirements from the ultra-dense heterogeneous network as the access network, improve the network access quality, and make full use of the architecture advantages of the ultra-dense heterogeneous network.
[0028] Furthermore, in step S3, the optimal network is the candidate network with the highest reward value. From the above example, it can be seen that the candidate networks include UMTS, LTE, WLAN, and WiMAX. Then, accessing the optimal network can be represented as the action vector a t , for example, the action vector a t = [1, 0, 0, 0] means selecting to access the UMTS network at time t, a t = [0, 1, 0, 0] means selecting to access the LTE network at time t, and so on; further, the network access actions of the user terminal at all times can form the action space At, as shown in Equation (3): A t = {a1, a2, a3, ..., a t} (3) where A t represents the action space; a t represents the action vector at time t.
[0029] Compared with the prior art, the embodiments of the present invention can improve the network access quality by comprehensively considering factors at multiple levels such as the network, users, and services, so as to better adapt to the complex requirements in the ultra-dense heterogeneous network and have more advantages in the high-dynamic and high-complex network environment.
[0030] As an optional implementation manner, the utility function set includes an S-shaped utility function, an exponential utility function, a logarithmic utility function, a linear utility function, and a linear piecewise utility function; where: The S-shaped utility function is shown as follows: The exponential utility function is shown as follows: The logarithmic utility function is shown as follows: u(x) = d + e' * ln(x + f) (6) The linear utility function is shown as follows: u(x) = gx + h (7) The linear piecewise utility function is shown as follows: In the above formulas (4) to (8), u(x) represents the first utility value; x represents the network parameter; e represents the natural constant; a″, b, c, e′, f, g, and h represent coefficients; i′ and j′ represent thresholds.
[0031] It should be noted that different service types have different requirements for different network parameters. Therefore, the embodiments of the present invention pre-construct a utility function set. It can be seen from formulas (4) to (8) that the utility functions of the embodiments of the present invention include S-shaped functions (Sigmoid functions), exponential functions (Exponential functions), logarithmic functions (Logarithm functions), linear functions (Linear functions), and linear piecewise functions (Linear piecewise functions). By constructing a utility function set containing various types of utility functions, the embodiments of the present invention can comprehensively describe the requirements of different service types and different users for different network parameters.
[0032] Exemplarily, when the service type is a voice service, the utility function type for bandwidth and latency can be selected as an S-shaped utility function, the utility function type for packet loss rate and price can be selected as a linear piecewise utility function, and the utility function type for jitter can be selected as an exponential utility function; further, after selecting the utility function type, the coefficients of the utility function can also be adjusted to obtain the final target utility function.
[0033] As an optional implementation manner, the calculating the utility values of the network parameters by invoking the target utility function specifically includes: Inputting the network parameters into the corresponding target utility function to obtain the initial utility values of the network parameters; Normalizing the initial utility values of the network parameters to obtain the utility values of the network parameters for constructing the utility matrix.
[0034] It should be noted that network parameters can be divided into benefit-type parameters and non-benefit-type parameters. Among them, the characteristic of benefit-type parameters is that the larger the value of the network parameter, the higher the user satisfaction; the characteristic of non-benefit-type parameters is that the larger the value of the network parameter, the lower the user satisfaction. Exemplarily, bandwidth is a benefit-type parameter, and latency, jitter, packet loss rate, and price are non-benefit-type parameters. Then, by selecting appropriate utility functions for each network parameter and setting corresponding coefficients, benefit-type parameters and non-benefit-type parameters can be distinguished, and the conversion from the numerical value of the network parameter to the utility value can be realized.
[0035] Further, in order to widen the gap between the utility values of the same type of network parameters during comparison and discrimination, it is also necessary to normalize the initial utility values of each type of network parameter, that is, to normalize the initial utility values of bandwidth, latency, jitter, packet loss rate, and price respectively. Exemplarily, taking the normalization of the initial utility value of bandwidth at time t as an example, it is necessary to summarize the initial utility values of bandwidth of different candidate networks at time t, and select the maximum value and the minimum value among them. Then, use formula (9) to normalize the initial utility values of bandwidth of each candidate network to obtain the utility values of bandwidth of each candidate network: where, utility value represents the utility value; value represents the initial utility value; value min represents the minimum value in the initial utility value; value max represents the maximum value in the initial utility value.
[0036] Further, summarize each utility value to obtain the utility matrix.
[0037] As an optional implementation manner, the preference vector is calculated by the following method: Based on the target service type, pairwise comparison is performed on the importance of any two of the network parameters to obtain a number of importance ratios; The importance ratios are summarized to obtain a fuzzy consistent matrix; Based on the fuzzy consistent matrix, a preference weight calculation formula is used to calculate the preference weights of the network parameters to construct the preference vector.
[0038] It should be noted that the preference weights in the embodiments of the present invention are used to represent the weights of different network parameters; and since the importance of different network parameters to different service types is different, the embodiments of the present invention construct different preference vectors for different service types.
[0039] Specifically, any two network parameters are compared to determine the importance ratio between them, where the specific meaning of the importance ratio is shown in Table 1; then, the importance ratios are summarized to construct a fuzzy consistent matrix. It should be noted that in some embodiments, a consistency test is also performed according to formula (10) before calculating the preference vector: where r ij represents the value in the i-th row and j-th column of the fuzzy consistent matrix; r ii represents the value in the i-th row and i-th column of the fuzzy consistent matrix; r ji represents the value in the j-th row and i-th column of the fuzzy consistent matrix; r ik represents the value in the i-th row and k-th column of the fuzzy consistent matrix; r jk represents the value in the j-th row and k-th column of the fuzzy consistent matrix; n represents the total number of network parameters.
[0040] Furthermore, the preference weight calculation formula is shown in formula (11), that is, the preference vector is calculated through formula (11): where represents the preference weight of the i-th network parameter when the service type is sb; r ij represents the value in the i-th row and j-th column of the fuzzy consistent matrix.
[0041] Furthermore, the preference weights of the network parameters are summarized to obtain a preference vector.
[0042] Exemplarily, in some embodiments, the preference vectors can also be calculated in advance for different service types and summarized into a preference matrix (preference vector of selected network) as shown in formula (12), then, in actual application, the preference vector can be selected from the preference matrix according to the target service type: Among them, W represents the preference matrix; represents the preference weight of the first network parameter when the service type is sb; represents the preference weight of the i-th network parameter when the service type is sb; represents the preference weight of the first network parameter when the service type is sb'; represents the preference weight of the i-th network parameter when the service type is sb';
[0043] Exemplarily, in some embodiments, the preference vector can also be set by the user, which is not limited herein.
[0044] Furthermore, calculate the inner product of the preference vector and the utility matrix to obtain the reward value as shown in Equation (13): r t = p'·U' (13) Among them, r t represents the reward value at time t; p' represents the preference vector; U' represents the utility matrix.
[0045] Table 1 Meanings of importance ratios Among them, x in Table 1 i represents the i-th network parameter; x j represents the j-th network parameter; r ij represents the importance ratio between the i-th network parameter and the j-th network parameter, that is, the value in the i-th row and j-th column of the fuzzy consistency matrix.
[0046] As one alternative embodiment, the second reward value calculation model is a first deep Q-network, and the first deep Q-network is pre-trained in the following manner: Initialize the weight parameters of the first deep Q-network and the weight parameters of the second deep Q-network so that the weight parameters satisfy the normal distribution; among them, the second deep Q-network has the same structure as the first deep Q-network; Iteratively train the first deep Q-network and the second deep Q-network respectively to update the weight parameters of the first deep Q-network and the weight parameters of the second deep Q-network respectively; among them, after the first deep Q-network completes the network selection at each time step, store the current network parameters, the target access network, the reward value, and the future network parameters as a set of experience samples; the second deep Q-network is trained using the experience samples; During the iterative training process, when the first parameter update condition is satisfied, also perform gradient descent update on the first deep Q-network based on the second deep Q-network to update the weight parameters of the first deep Q-network; During the iterative training process, based on a preset update frequency, the weight parameters of the second deep Q-network are updated according to the current weight parameters of the first deep Q-network.
[0047] It should be noted that the deep Q-network can learn complex network state-network selection relationships through a neural network, so as to support the personalized service requirements of users and optimize the network access strategy. Therefore, in some embodiments of the present invention, a first deep Q-network is also constructed and trained to enable it to replace the first reward value calculation model to calculate the reward value of candidate networks. It should be noted that in some embodiments, the first deep Q-network can be used to output the reward values of different candidate networks, while in some embodiments, the first deep Q-network can also directly output the candidate network with the highest reward value, that is, the optimal network, which is not limited herein.
[0048] Furthermore, in order to reduce the bias transmission caused by "bootstrapping", an embodiment of the present invention also constructs a second deep Q-network to assist the first deep Q-network in training, where the second deep Q-network has the same structure as the first deep Q-network. It should be noted that the first deep Q-network can also be called a value network, and the second deep Q-network can also be called a target network. In the first deep Q-network, the Q value of the network parameter-network selection pair is approximated through a neural network, where the Q value represents the reward value that can be obtained by performing a certain network selection action in a specific network state. The first deep Q-network updates its weights through interaction with the first reward value calculation model and the second deep Q-network. The second deep Q-network is used to stabilize the training process. The second deep Q-network has the same structure as the first deep Q-network, but the update speed of the weight parameters of the second deep Q-network is restricted, so as to provide a relatively stable target Q value (expected cumulative reward value) during the training process.
[0049] Furthermore, the first deep Q-network and the second deep Q-network have the same structure. Based on the application scenario of network selection, the first deep Q-network adopts a fully connected network structure, including an input layer, an output layer, and a feedforward layer, and the activation function is the ReLu activation function. Further, assuming that there are N candidate networks, and n network parameters (such as channel conditions, delay, and jitter, etc.) of the candidate networks are recorded in the state space, then the first deep Q-network will use these n network parameters as inputs, and the number of neurons in the output layer is N.
[0050] Specifically, the first deep Q-network and the second deep Q-network are trained independently, and the training data is different. Specifically, after the network selection at each time step is completed, the first deep Q-network stores the current network parameters, the target access network, the reward value, and the future network parameters as a set of experience samples, while the second deep Q-network is trained using the experience samples. Exemplarily, the experience samples can be stored in an experience replay module, and the second deep Q-network randomly extracts experience samples from the experience replay module for training during training.
[0051] It is worth noting that the purpose of the Experience Replay module is to improve the sample utilization rate and stability, and prevent over-reliance on the current network state and network selection actions during the training process, thus resulting in unstable learning results. The Experience Replay module stores the current network parameters, the target access network, the labeled reward value, and the future network parameters at each time step. Moreover, during training, randomly extracted stored experience samples are used for updating to reduce the correlation between data and improve the training efficiency and the generalization ability of the model.
[0052] Furthermore, the hyperparameters related to the experience replay module are the buffer size and the sampling frequency. The size of the buffer of the experience replay module is generally set to a fixed value. When the buffer is full, the earlier experience samples will be overwritten by new experience samples. A larger buffer can store more experience samples to improve the sample diversity. For relatively simple tasks (such as a small number of network parameters and actions), a smaller buffer can be set, for example, several thousand experience samples; for complex tasks (such as large heterogeneous network access scenarios), more experience samples are needed to cover various network states, so the buffer size usually ranges from tens of thousands to millions of experience samples. It can be understood that during training, a small batch of samples is randomly extracted from the experience buffer for training. The sampling frequency represents the number of small batches of samples randomly extracted from the buffer each time during training. When the number of stored experience samples reaches a certain threshold, the experience samples are extracted from the queue. In the embodiments of the present invention, a uniform sampling method is used for data extraction. Uniform sampling means randomly selecting samples from the experience replay module for training, where the probability of each experience sample being selected is equal. This method can ensure the diversity of experience samples, prevent the model from overfitting certain specific experience samples, and can break the temporal correlation.
[0053] Further, during the independent training process, the first deep Q-network updates its weight parameters according to the difference between the first reward value (the reward value calculated by the first deep Q-network) and the labeled reward value (the reward value calculated by the first reward value calculation model). Further, in addition to the independent training, the first deep Q-network and the second deep Q-network also interact to update the weights. For example, when the first parameter update condition is met (such as the number of experience samples in the experience replay module meets the preset threshold), gradient descent update is performed on the first deep Q-network based on the second deep Q-network to update the weight parameters of the first deep Q-network; and, based on the preset update frequency, the first deep Q-network updates the weight parameters of the second deep Q-network according to the current weight parameters. For example, the current weight parameters of the first deep Q-network are copied to the second deep Q-network, or the current weight parameters of the first deep Q-network and the second deep Q-network are weighted and summed as the updated weight parameters of the second deep Q-network, which is not limited here.
[0054] It should be noted that during the actual training process, it is also necessary to first initialize the weight parameters of the two deep Q-networks with a normal distribution. By initializing the weight parameters as random values with small variances, the problem of gradient explosion or disappearance can be avoided. And, by appropriately initializing the weight parameters of the hidden layer, the deep Q-network can reasonably explore the solution space at the beginning, improving the subsequent learning effect. Further, when initializing the weight parameters, the discount factor, update frequency, and exploration rate can also be initialized.
[0055] As an optional implementation manner, the iterative training of the first deep Q-network and the second deep Q-network respectively includes the following steps for iterative training of the first deep Q-network: At each time step, input the current network parameters; Determine the network selection strategy based on the exploration rate. When the network selection strategy is random exploration, randomly select one of the candidate networks as the target access network. When the network selection strategy is model selection, calculate the first reward value of each candidate network based on the current network parameters, and select the candidate network with the highest first reward value as the target access network; where the exploration rate has a negative correlation with the time step; Based on the target access network, use the state transition function to calculate the future network parameters of each candidate network; Based on the current network parameters, use the first reward value calculation model to calculate the labeled reward value, and calculate the loss value according to the first reward value and the labeled reward value to update the weight parameters of the first deep Q-network; Use the future network parameters as the current network parameters for the next time step, and return the step of network policy selection for training in the next time step until the preset iteration training end condition is met.
[0056] It can be understood that the behavior policy function is used to control the interaction between the user and the environment. That is, the user will select actions according to the policy. Among them, the most common policy is the conditional probability distribution π(a|s), that is, the probability of taking the network selection action a in the network state s (network parameters), as shown in Equation (14): π(a|s) = P(A t = a|S t = s) (14) Among them, π(a|s) represents the conditional probability distribution; a represents the network selection action; A t represents the action space; s represents the network state; S t represents the state space; P(A t = a|S t = s) represents the probability of taking the network selection action a in the network state s.
[0057] At this time, the action with a larger probability value has a higher probability of being selected by the user. In the embodiments of the present invention, the E-greedy strategy is directly adopted, that is, with probability randomly select a candidate network, and with probability ε select the candidate network that maximizes the network value function Q(s, a), where where π(a|s) represents the probability of selecting action a under the given state s, where π is usually called the policy; A(s) represents the set of all optional actions in the action space under state s; |A(s)| represents the size of the action space, that is, the total number of optional actions; ε represents the exploration rate; a represents the network selection action; s represents the network state; argmax represents the argmax function; Q(s, a) represents the network value function. It should be noted that Q and Q with superscripts and / or subscripts in the following π 、Q ★ etc. refer to the network value function and are indistinguishable.
[0058] It should be noted that the exploration rate determines the proportion of network access actions of the first deep Q-network that are based on random exploration in the initial stage of training. The exploration rate plays a role in balancing exploration (random exploration) and exploitation (model selection). In the network switching scenario, it is often necessary to weigh between these two strategies. Among them, exploitation refers to "selecting the currently known optimal action (i.e., switching to the candidate network with the best known performance)", which helps to optimize the current performance and reduce unnecessary exploration; exploration refers to "selecting a random action with a certain probability, that is, randomly selecting a candidate network for switching", so as to explore new possibilities, which helps to avoid falling into local optimal solutions. Since in a wireless network, factors such as network load, number of users, and signal strength will change dynamically, if only relying on the current optimal strategy, it may miss better future choices. Therefore, it is necessary to adopt a combination of exploration and exploitation.
[0059] Furthermore, a linear decay strategy can be used to dynamically adjust the exploration rate, that is, as training progresses, the exploration rate gradually decreases. For example, it linearly decreases from an initial value until a certain fixed value. This is because in the initial stage, the first deep Q-network has not been fully trained and more exploration is needed to collect data in the environment, so a higher exploration rate is set; while as the understanding of the environment gradually deepens, the exploration rate can be gradually reduced to utilize the learned strategies.
[0060] Furthermore, to evaluate the reward value of the candidate network for each vertical handover decision, the Bellman equation of the network value function is introduced again: Q π (s, a) = R(s, a) + γ∑ s'∈S P(s’|s, a)∑ a’∈A π(a’|s’)Q π (s’, a’) (16) where Q π (s, a) represents the expected cumulative reward value that the user can obtain after executing action a from state s under policy π; R(s, a) represents the label reward value calculated by the first reward value calculation model after executing action a in state s; ∑ s'∈S P(s’|s, a) represents the probability of transitioning to the next state s' after executing action a in state s; γ represents the discount factor; P(s’|s, a) represents the transition probability function in the Markov Decision Process (MDP), satisfying ∑ s'∈s P(s’|s, a) = 1; π(a|s) represents the probability of executing action a given state s, and π is usually called the policy; Qπ (s’, a’) represents the reward value for executing action a’ in state s’ under policy π; s represents the network state; a represents the network-selected action; s’ represents the state that the network environment may transfer to after the user executes action a in the current state s; S represents the state space, which is the set of all possible states; A represents the action space, which contains all the network-selected actions that can be executed in each state; a’ represents a specific network-selected action selected in state s’.
[0061] Equation (16) indicates that the reward value for taking network-selected action a in network state s is equal to the immediate reward plus the discounted sum of the reward values of future state-action pairs, where the discount rate γ ∈ [0, 1]. In the model, the Bellman equation of the optimal value function is adopted, so the equation will become: Q π (s, a) = R(s, a) + γ ∑ s'∈S P(s’|s, a) max a'∈A Q π (s’, a’) (17) Among them, Q π (s, a) represents the expected cumulative reward value that the user can obtain after executing action a from state s under policy π; R(s, a) represents the label reward value calculated by the first reward value calculation model after executing action a in state s; γ represents the discount factor; P(s’|s, a) represents the transition probability function in the Markov decision process; Q π (s′, a’) represents the reward value for executing action a’ in state s’ under policy π; s represents the network state; a represents the network-selected action; s’ represents the state that the network environment may transfer to after the user executes action a in the current state s; S represents the state space, which is the set of all possible states; A represents the action space, which contains all the actions that can be executed in each state; a’ represents a specific action selected in state s’.
[0062] Equation (17) indicates that under the optimal policy π, the network selection reward value of the user in each network state is achieved by maximizing the combination of the immediate reward and the future expected reward. The network value function Q π (s, a) will converge to the optimal after model iteration.
[0063] Further, the training end condition may be reaching the maximum number of steps, entering a special state, or other conditions, which are not limited herein. Exemplarily, it can be set to terminate the current training round after a certain number of network access switches or a certain number of network selections. Further, in order to better evaluate the evaluation metrics of the network access selection strategy on the user experience in each training round, the termination condition is not set based on the total reward value of the entire training round, but on the average reward value obtained by dividing the total reward value by the number of network switches as a reference to set the termination condition.
[0064] As an optional implementation manner, the gradient descent update of the first deep Q network based on the second deep Q network includes: Input the current network parameters into the first deep Q network to obtain the first reward value and the target access network of each candidate network; Based on the target access network, calculate the future network parameters using the state transition function; Based on the future network parameters and the target access network, calculate the second reward value of each candidate network at the next time step using the second deep Q network; Based on the current network parameters, calculate the label reward value using the first reward value calculation model; Based on the label reward value, the second reward value, and the discount factor, calculate the expected cumulative reward value; Calculate the error between the expected cumulative reward value and the first reward value, and calculate the gradient based on the error; Perform gradient descent update on the first deep Q network based on the gradient.
[0065] It should be noted that the embodiments of the present invention use discounted return to discount the future reward value. Among them, the definition of discounted return is as follows: U t =r t +γ·r t+1 +γ 2 ·r t+2 +γ 3 ·r t+3 +…(18) Wherein, U t represents the discounted return (return function); γ represents the discount factor; r t represents the reward value at time t; r t+1 represents the reward value at time t + 1; r t+2 represents the reward value at time t + 2; r t+3 represents the reward value at time t + 3.
[0066] It should be noted that the discount factor controls the impact of future reward values on the current decision (network access selection), and determines whether the algorithm pays more attention to recent reward values or future reward values during the optimization process. When the discount factor is high, it indicates that the importance of future reward values is high, while when the discount factor is low, it means that more attention is paid to the current reward value. Further, at each time step, the future reward value is multiplied by γ, so that the reward value after a certain number of steps becomes smaller. By decreasing the future reward value, the discount factor can prevent the model from overemphasizing distant reward values and focus on more certain short-term goals.
[0067] Furthermore, the value range of the discount factor is between [0, 1]. The closer it is to 1, the greater the impact of future reward values. For the network switching problem, a larger discount factor can ensure long-term benefits, while a smaller discount factor focuses on short-term interests. Exemplarily, for scenarios that focus on immediate network performance, such as scenarios where users have high immediate experience requirements for parameters such as bandwidth and latency, a smaller discount factor can be set to make the first deep Q-network more inclined to switch to the currently optimal network. Another example is that for scenarios that focus on long-term network stability, frequent network switching may lead to greater power consumption or fluctuations in user experience, so a higher discount factor can be set. That is, a higher discount factor prompts the model to consider future stability more and avoid frequent switching.
[0068] Furthermore, the calculation processes of the first deep Q-network and the second deep Q-network are described in detail below.
[0069] Assume U t is the value of the reward function at time t, defined as: where, U t represents the value of the reward function at time t; r k' represents the reward value at time k′. In the present invention, lowercase r and uppercase R refer to the same meaning; γ represents the discount factor; n′ represents the upper limit of the number of time steps.
[0070] Similarly, let U t be the value of the reward function at time t + 1 as: where, U t+1 represents the value of the reward function at time t + 1; r k' represents the reward value at time k′; γ represents the discount factor; n′ represents the upper limit of the number of time steps.
[0071] From the definitions of U t and U t+1 the formula (21) can be deduced: Among them, U t represents the value of the return function at time t; r t represents the reward value at time t; r k' represents the reward value at time k'; γ represents the discount factor; n' represents the upper limit of the time step.
[0072] Then, the optimal network value function can be expressed as shown in Equation (22): Q ★ (s t , a t ) = maxπ E[U t |S t = s t , A t = a t (22) Among them, Q ★ (s t , a t ) represents the maximum expected value of the expected cumulative reward that the user can obtain when choosing action a t in state s t and acting according to the optimal policy π; π represents the rule or distribution for the user to choose actions in each state s t ; E[U t |S t = s t , A t = a t represents the expected value of the expected cumulative reward value U t in state s t and action a t ; S t represents the state space at time t; At represents the action space at time t; s t represents the network state at time t; a t represents the action vector at time t, that is, the network selected at time t.
[0073] Furthermore, Equation (22) can be written as shown in Equation (23): Among them, Q ★ (s t , a t ) represents the maximum expected value of the expected cumulative reward that the user can obtain when choosing action a t in state s t and acting according to the optimal policy π; s t represents the network state at time t; a t represents the network selection action at time t; St+1 Indicates that the user is in the current state S t = s t and performs action A t = a t The next moment state to which the environment transfers after that; p(·|s t , a t ) represents the conditional probability distribution of transferring to the next state s t after performing action a t in state s t+1 ; r t represents the label reward value at time t; γ represents the discount factor; max A ∈ A Q ★ (S t+1 ) represents the maximum value of the corresponding Q t+1 values for all possible actions in the next state S ★ , that is, it represents the optimal action reward value; S t represents the state space at time t; A t represents the action random variable at time t; A represents the action space, which contains the set of all optional actions, that is, the set of access networks that the user can choose from.
[0074] It should be noted that in formula (23), the script letter is used to represent the action space to represent the set of all possible actions, the capital letter A t is used to represent the action random variable at time t, and the lowercase letter a t is used to represent the specific value (action instance) of A t .
[0075] It should be noted that in formula (23), Q ★ (s t , a t ) is the expectation of U t , max A∈A Q ★ (S t+1 , A) is the expectation of U t+1 , that is, the right side of formula (23) is the expectation, so a Monte Carlo approximation is made for the expectation.
[0076] Furthermore, when the first deep Q-network performs action a t , that is, after selecting a certain candidate network, a new network state s t+1 is calculated through the state transition function p″(s t |s t , a t+1, that is, new network parameters (future network parameters) are obtained, and thus a four-tuple can be obtained as (s t , a t , r t , s t+1 ) Therefore, the result of formula (24) can be calculated as follows: r t + γ·max A∈A Q ★ (S t+1 , A) (24) where r t represents the label reward value at time t; γ represents the discount factor; max A∈A Q ★ (S t+1 , A) represents the maximum value of the corresponding Q t+1 values among all possible actions in the next state S ★ , that is, it represents the optimal action reward value.
[0077] Formula (24) can be regarded as the Monte Carlo approximation of formula (23). Therefore, from formula (23) and formula (24), formula (25) can be obtained: Q ★ (s t , a t ) ≈ r t + γ·max a∈A Q ★ (s t+1 , a) (25) where Q ★ (s t , a t ) represents the maximum expected value of the expected cumulative reward that the user can obtain when choosing action a t in state s t and acting according to the optimal policy π; r t represents the reward value at time t; γ represents the discount factor; max a∈A Q ★ (s t+1 , a) represents the maximum value of the corresponding Q t+1 values among all possible actions in the next state S ★ , that is, it represents the optimal action reward value.
[0078] The optimal network value function Q ★ (s, a) in formula (25) can be replaced by the neural network Q ★ (s, a; ω) to obtain formula (26): Q(st ,a; ω) ≈ r t + γ·mgx a∈A Q(s t+1 ,a; ω) (26) Wherein, Q(s t ,a; ω) represents the value function approximately obtained by the neural network model in the given state s t and action a; ω represents the weight parameters of the neural network. By continuously updating ω, Q(s t ,a; ω) is closer to the actual expected cumulative reward value, so as to provide an accurate network selection basis for users; r t represents the label reward value at time t; γ represents the discount factor; max a∈A Q(s t+1 ,a; ω) represents the Q t+1 value corresponding to all possible actions in the next state S ★ value, that is, it represents the optimal action reward value.
[0079] It should be noted that, in order to distinguish the reward values calculated by different deep Q networks or models, in the embodiments of the present invention, the reward value calculated by the first deep Q network is called the "first reward value", the reward value calculated by the second deep Q network is called the "second reward value", and the reward value calculated by the first reward value calculation model is called the "label reward value"; further, use to represent the first reward value calculated by the first deep Q network at time t, and use to represent the expected cumulative reward value at time t calculated based on the second deep Q network, as shown in formulas (27) and (28) respectively: In formulas (27) and (28), represents the first reward value at time t; s t represents the network state at time t; a t represents the network selection action at time t; ω represents the weight parameters of the first deep Q network; ω* represents the weight parameters of the second deep Q network; Q(s t ,a t ; ω) represents the network value function approximately obtained by the neural network model under the weight parameters ω, the given state s t and action a t ; represents the defined relationship; represents the expected reward value at time t; r t represents the label reward value at time t; γ represents the discount factor; represents in the next state S t+1Under all possible actions The corresponding Q in ★ The maximum value, that is, it represents the optimal action reward value; Represents the action space; s t+1 Represents the network state at time t+1.
[0080] It should be noted that in formulas (27) and (28), Q(s t , a t ; ω) is the reward value calculated by the first deep Q network, and Q(s t+1 , a; ω*) is the reward value calculated by the second deep Q network, and the expected reward value Is based on the labeled reward value r t And (s t+1 , a; ω*) is calculated.
[0081] It should be noted that since Is based on the actually observed labeled reward value rt (calculated by the first reward value calculation model) and the prediction value of the second deep Q network, therefore, Can more accurately reflect the current network situation, and since Is the reward value calculated by the first deep Q network based on the weight parameters during the training process, there is a certain error. Therefore, it is necessary to adjust the Of the first deep Q network to make it gradually approach Define the loss function: Among them, L(ω) represents the loss function; Q(s t , a t ; ω) represents the network value function approximately obtained by the neural network model under the weight parameter ω, given state s t And action a t Under, ω represents the weight parameter of the first deep Q network; s t Represents the network state at time t; a t Represents the network selected action at time t; Represents the expected cumulative reward value at time t.
[0082] Furthermore, the gradient of L(ω) with respect to ω is shown in formula (30): Among them, Represents the gradient of the loss function L(ω) with respect to ω, which is used to adjust the network parameter ω to make Closer to Represents the first reward value at time t; Represents the expected cumulative reward value at time t; Represents the gradient vector of Q(s t , a t ; ω) with respect to ω. When deviates from , guide ω to be adjusted to reduce the deviation; Q(s t , a t ; ω) represents the network value function approximated by the neural network model under the condition that the weight parameter is ω, the given state is s t and the action is a t ; ω represents the weight parameter of the first deep Q network; s t represents the network state at time t; a t represents the network selected action at time t.
[0083] It should be noted that during the training process, in order to reduce the deviation transmission caused by "bootstrapping" and alleviate the influence of the overestimation problem of the deep Q network, the embodiments of the present invention use the second deep Q network to provide the Specifically, let ω represent the weight parameter of the first deep Q network, ω′ represent the updated weight parameter of the first deep Q network, ω* represent the weight parameter of the second deep Q network, and ω′* represent the updated weight parameter of the second deep Q network; each time during training, a quadruple (s t , a t , r t , s t+1 ) will be randomly taken out from the experience replay module, that is, a set of experience samples will be taken out; first, perform forward propagation on the first deep Q network to obtain the first reward value at time t As shown in Equation (31): Among them, represents the first reward value at time t; s t represents the state at time t; a t represents the network selected action at time t; ω represents the weight parameter of the first deep Q network; Q(s t , a t ; ω) represents the network value function approximated by the neural network model under the condition that the weight parameter is ω, the given state is s t and the action is a t .
[0084] Furthermore, in the selection of the target access network, based on the network state s t , select a target access network a * to maximize the output of the first deep Q network: where a * represents the optimal access network; argmax represents the argmax function; represents the candidate network in the candidate network set A that maximizes the first reward value when the weight parameter is ω and the network state is s t ; s t represents the network state at time t; ω represents the weight parameter of the first deep Q-network; a represents a certain network in the candidate network set; A represents the candidate network set.
[0085] Furthermore, based on the target access network a * and the network state s at time t + 1 t+1 use the second deep Q-network to calculate the second reward value where represents the second reward value at time t + 1; s t+1 represents the network state at time t + 1; a * represents the target access network; ω* represents the weight parameter of the second deep Q-network.
[0086] Then, calculate the expected cumulative value and the error δ t , as shown in equations (34) and (35): In equations (34) and (35), represents the expected cumulative reward value at time t; r t represents the label reward value at time t; γ represents the discount factor; represents the second reward value at time t + 1; δ t represents the error at time t; represents the first reward value at time t.
[0087] Furthermore, perform backpropagation on the first deep Q-network. Specifically, obtain the gradient according to equation (30) and then perform gradient descent to update the weight parameter of the first deep Q-network: where ω′ represents the updated weight parameter of the first deep Q-network; ω represents the weight parameter of the first deep Q-network before update; α represents the learning rate, which is used to control the step size of the model parameters in each update, and, a needs to be set manually, and its value is generally between 0.001 - 0.00010; δ t represents the error at time t; represents; s tRepresents the network state at time t; a t Represents the network selection action at time t.
[0088] Furthermore, the weight parameters of the second deep Q-network are updated using a soft update method, also known as Exponential Moving Average (EMA). Specifically, a momentum τ is introduced, and τ ∈ (0, 1) is a hyperparameter that needs to be manually adjusted. The weighted average of the weight parameters of the old second deep Q-network and the new first deep Q-network is calculated and then assigned to the second deep Q-network: ω′*←τ·ω′+(1 - τ)·ω* (37) where ω′* represents the updated weight parameters of the second deep Q-network; ω′ represents the updated weight parameters of the first deep Q-network; τ represents the momentum; and ω* represents the weight parameters of the second deep Q-network before update.
[0089] As can be seen from the above, in each training process of gradient descent update, the output of the first deep Q-network is used to calculate the loss function, and the expected cumulative reward value is calculated through the second deep Q-network. The purpose of this design is to reduce the fluctuation of Q-values and prevent unstable phenomena (such as divergence) during the training process. In actual operation, the weight parameters of the second deep Q-network are not updated step by step following the first deep Q-network, but are updated based on the update frequency, that is, the weight parameters are copied from the first deep Q-network every certain number of steps (for example, every 1000 steps), so as to maintain a certain lag. When the update frequency of the second deep Q-network is too low, it may lead to inaccurate estimation of Q-values; while if the update frequency is too high, it may cause the second deep Q-network to be too close to the first deep Q-network, resulting in bootstrapping, thus causing bias and losing the stabilizing effect.
[0090] It can be understood that the update frequency is used to stabilize the training process and prevent training instability. The first deep Q-network updates its own parameters through experience replay in each round of training, while the second deep Q-network remains relatively stable and is only updated at certain specific moments. Therefore, the update frequency determines the synchronization frequency between the first deep Q-network and the second deep network. The update frequency is usually represented by a fixed number of steps N′, that is, the weight parameters of the second deep Q-network will be updated every N′ rounds.
[0091] Furthermore, a higher update frequency allows the second deep Q-network to adapt to network environment changes faster. However, overly frequent updates may lead to poor learning effects of the first deep Q-network, introduce instability, and cause policy oscillations. A lower update frequency can maintain the stability of the network switching policy. However, the model may be slow to respond to network environment changes. Although the policy is more stable, the difference between the first deep Q-network and the second deep Q-network may be too large, resulting in a slower convergence of the training process or an overly conservative learned policy. Specifically, the update frequency will be set according to the number of training rounds and the actual number of steps generated.
[0092] To better understand the training process of the deep Q-network, a specific example is given below for illustration. Refer to Figure 2 , which is the training flow chart of the deep Q-network provided by an embodiment of the present invention. Herein, the "Q-network" is the first deep Q-network, and the "target network" is the second deep Q-network. And by Figure 2 It can be seen that the training algorithm can be divided into several modules, including an "initialization" module, a "policy update" module, an "experience replay" module, and a "network environment feedback" module. Then, the training process of the deep Q-network mainly includes the following steps 1 to 5: Step 1: The algorithm uses the "initialization" module to initialize the weight parameters of the Q-network and the target network, sets the discount factor, update frequency, and exploration rate, and receives network state parameters (network parameters) from the user terminal, and converts them into state space information and action space information.
[0093] Step 2: The network state parameters are input into the "policy update" module. The algorithm randomly selects a candidate network with a probability of , or selects the candidate network with the largest first reward value (optimal candidate network) according to the first reward value calculated by the Q-network with a probability of . And at the initial stage of iteration, random actions will be selected with a higher probability to ensure full exploration of the access effects of each network.
[0094] Step 3: The "network environment feedback" module receives the network selection result and relevant network state parameters, and outputs the updated network state information (new network state / future network parameters) and the reward value of this network selection (candidate network reward value). And the network state information will be returned to the "initialization" module for state update.
[0095] Step 4: The "experience replay" module stores the network selection result, relevant network state parameters, and the reward value of the network selection. And the stored experience samples are continuously replaced and updated after reaching the upper limit.
[0096] Step 5: After the experience samples meet the training conditions, the target network randomly samples experience samples from the "network environment feedback" module for learning, that is, sampling network state reward values and network state parameters for learning, and outputs a network value function to update the Q network. The Q network periodically transmits the latest network gradients learned to the target network, where ω represents the weight parameters of the Q network, ω T represents the weight parameters of the target network. Moreover, it is necessary to ensure that each Q-value update is based on the actual network access feedback to gradually optimize the action strategy.
[0097] Finally, through continuous iteration, the Q network will gradually converge and learn the optimal network access strategy.
[0098] Compared with the prior art, in the embodiment of the present invention, two independent deep Q networks are respectively used for network selection and network value evaluation. By adopting a separation strategy, the deviation of a single network in selecting and evaluating actions can be effectively reduced, and the problem of overestimating the Q value in the traditional DQN algorithm is successfully solved, improving the decision-making accuracy of the model.
[0099] Compared with the prior art, the ultra-dense heterogeneous network access method provided by the embodiment of the present invention inputs the target service type and network parameters of candidate networks; wherein, the types of the network parameters are greater than 1; based on the network parameters, a first reward value calculation model or a second reward value calculation model is called to calculate the reward values of the candidate networks; the candidate network with the highest reward value is selected as the optimal network, and the user terminal is connected to the optimal network; wherein, the first reward value calculation model calculates the reward value through the following steps: based on the target service type, calculate a target preference vector, and select a target utility function from a set of utility functions for each of the network parameters; call the target utility function to calculate the utility values of the network parameters and construct a utility matrix; calculate the inner product of the target preference vector and the utility matrix to obtain the reward values of the candidate networks. By comprehensively considering factors at multiple levels such as the network, user, and service, the embodiment of the present invention can improve the network access quality, thereby better adapting to the complex requirements in the ultra-dense heterogeneous network and having more advantages in a highly dynamic and highly complex network environment.
[0100] See Figure 3 , the embodiment of the present invention further provides an ultra-dense heterogeneous network access device 10, including: A parameter input module 11 for inputting the target service type and network parameters of candidate networks; wherein, the types of the network parameters are greater than 1; A reward value calculation module 12 for calling a first reward value calculation model or a second reward value calculation model to calculate the reward values of the candidate networks based on the network parameters; A network access module 13 is configured to select the candidate network with the highest reward value as the optimal network and connect the user terminal to the optimal network; Among them, the first reward value calculation model calculates the reward value through the following steps: Based on the target service type, calculate a target preference vector and select a target utility function from the utility function set for each of the network parameters; Call the target utility function to calculate the utility values of each of the network parameters and construct a utility matrix; Calculate the inner product of the target preference vector and the utility matrix to obtain the reward values of each of the candidate networks.
[0101] The ultra-dense heterogeneous network access device provided by the embodiments of the present invention can implement all the process steps of the ultra-dense heterogeneous network access method described in the above embodiments. The functions and the achieved technical effects of each module and unit in the device respectively correspond to the functions and the achieved technical effects of the ultra-dense heterogeneous network access method described in the above embodiments, and the specific implementation manners will not be elaborated herein.
[0102] See Figure 4 , the embodiments of the present invention further provide an ultra-dense heterogeneous network access device 20, including a processor 21, a memory 22, and a computer program stored in the memory 22 and configured to be executed by the processor 21. When the processor 21 executes the computer program, it implements the steps in the embodiments of the ultra-dense heterogeneous network access method as described above, such as Figure 1 the steps S1 to S3 described in
[0103] The ultra-dense heterogeneous network access device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The ultra-dense heterogeneous network access device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the ultra-dense heterogeneous network access device, and does not constitute a limitation on the ultra-dense heterogeneous network access device. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the ultra-dense heterogeneous network access device may further include input / output devices, network access devices, a bus, etc.
[0104] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the ultra-dense heterogeneous network access device, and connects all parts of the ultra-dense heterogeneous network access device through various interfaces and lines.
[0105] The memory may be used to store the computer programs and / or modules. The processor realizes various functions of the ultra-dense heterogeneous network access device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the controller, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0106] Among them, if the modules integrated in the ultra-dense heterogeneous network access device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate forms, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0107] Compared with the prior art, the ultra-dense heterogeneous network access device, equipment and storage medium provided by the embodiments of the present invention input the target service type and network parameters of the candidate network; wherein, the type of the network parameters is greater than 1; based on the network parameters, call the first reward value calculation model or the second reward value calculation model to calculate the reward values of each candidate network; select the candidate network with the highest reward value as the optimal network, and connect the user terminal to the optimal network; wherein, the first reward value calculation model calculates the reward value through the following steps: based on the target service type, calculate the target preference vector, and select the target utility function from the utility function set for each network parameter; call the target utility function to calculate the utility values of each network parameter, and construct a utility matrix; calculate the inner product of the target preference vector and the utility matrix to obtain the reward values of each candidate network. The embodiments of the present invention can improve the network access quality by comprehensively considering factors at multiple levels such as the network, user and service, so as to better adapt to the complex requirements in the ultra-dense heterogeneous network and have more advantages in a high-dynamic and high-complex network environment.
[0108] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and retouches can be made, and these improvements and retouches are also regarded as the protection scope of the present invention.
Claims
1. A method for accessing a super-dense heterogeneous network, characterized in that Including: Input the target service type and the network parameters of the candidate networks; wherein, the types of the network parameters are more than 1; Based on the network parameters, call the first reward value calculation model or the second reward value calculation model to calculate the reward values of the candidate networks; Select the candidate network with the highest reward value as the optimal network, and connect the user terminal to the optimal network; Wherein, the first reward value calculation model calculates the reward value through the following steps: Based on the target service type, calculate the target preference vector, and select the target utility function for each of the network parameters from the utility function set; Call the target utility function to calculate the utility values of the network parameters, and construct a utility matrix; Calculate the inner product of the target preference vector and the utility matrix to obtain the reward values of the candidate networks.
2. The ultra-dense heterogeneous network access method according to claim 1, characterized in that The utility function set includes S-shaped utility function, exponential utility function, logarithmic utility function, linear utility function and linear piecewise utility function; wherein: The S-shaped utility function is shown as the following formula: The exponential utility function is shown as the following formula: The logarithmic utility function is shown as the following formula: u(x) = d + e' * ln(x + f) The linear utility function is shown as the following formula: u(x) = gx + h The linear piecewise utility function is shown as the following formula: In the above formulas, u(x) represents the first utility value; x represents the network parameter; e represents the natural constant; a″, b, c, e', f, g and h represent coefficients; i' and j' represent thresholds.
3. The ultra-dense heterogeneous network access method according to claim 2, wherein The calling the target utility function to calculate the utility values of the network parameters specifically includes: Input each of the network parameters into the corresponding target utility function to obtain the initial utility values of the network parameters; Normalize the initial utility values of the network parameters to obtain the utility values of the network parameters for constructing the utility matrix.
4. The ultra-dense heterogeneous network access method according to claim 1, wherein The preference vector is calculated by the following method: Based on the target service type, compare the importance of any two of the network parameters pairwise to obtain a number of importance ratios; Summarize the importance ratios to obtain a fuzzy consistent matrix; Based on the fuzzy consistent matrix, use the weight calculation formula to calculate the preference weights of the network parameters to construct the preference vector.
5. The ultra-dense heterogeneous network access method according to claim 1, wherein The second reward value calculation model is the first deep Q network, and the first deep Q network is pre-trained by the following method: Initialize the weight parameters of the first deep Q network and the weight parameters of the second deep Q network so that the weight parameters satisfy the normal distribution; wherein, the structures of the second deep Q network and the first deep Q network are the same; Perform iterative training on the first deep Q network and the second deep Q network respectively to update the weight parameters of the first deep Q network and the weight parameters of the second deep Q network respectively; wherein, after the first deep Q network completes the network selection at each time step, store the current network parameters, the target access network, the reward value and the future network parameters as a set of experience samples; the second deep Q network is trained using the experience samples; During the iterative training process, when the first parameter update condition is satisfied, gradient descent update is also performed on the first deep Q-network based on the second deep Q-network to update the weight parameters of the first deep Q-network; During the iterative training process, the weight parameters of the second deep Q-network are also updated based on a preset update frequency according to the current weight parameters of the first deep Q-network.
6. The ultra-dense heterogeneous network access method according to claim 5, wherein The iterative training of the first deep Q-network and the second deep Q-network respectively includes performing iterative training on the first deep Q-network by the following steps: At each time step, input the current network parameters; Determine the network selection strategy based on the exploration rate. When the network selection strategy is random exploration, randomly select one of the candidate networks as the target access network. When the network selection strategy is model selection, calculate the first reward value of each candidate network based on the current network parameters, and select the candidate network with the highest first reward value as the target access network; wherein, the exploration rate has a negative correlation with the time step; Based on the target access network, calculate the future network parameters of each candidate network by using the state transition function; Based on the current network parameters, calculate the label reward value by using the first reward value calculation model, and update the weight parameters of the first deep Q-network by calculating the loss value according to the first reward value and the label reward value; Take the future network parameters as the current network parameters of the next time step, and return to the step of network policy selection to perform the training of the next time step until the preset iterative training end condition is satisfied.
7. The ultra-dense heterogeneous network access method according to claim 5, wherein The gradient descent update of the first deep Q-network based on the second deep Q-network includes: Input the current network parameters into the first deep Q-network to obtain the first reward value of each candidate network and the target access network; Based on the target access network, calculate the future network parameters by using the state transition function; Based on the future network parameters and the target access network, calculate the second reward value of each candidate network at the next time step by using the second deep Q-network; Based on the current network parameters, calculate the label reward value by using the first reward value calculation model; Based on the label reward value, the second reward value and the discount factor, calculate the expected cumulative reward value; Calculate the error between the expected cumulative reward value and the first reward value, and calculate the gradient according to the error; Perform gradient descent update on the first deep Q-network based on the gradient.
8. A super-dense heterogeneous network access device, characterized in that, Includes: A parameter input module for inputting the target service type and the network parameters of the candidate networks; wherein, the type of the network parameters is greater than 1; A reward value calculation module for calculating the reward value of each candidate network by calling the first reward value calculation model or the second reward value calculation model based on the network parameters; A network access module for selecting the candidate network with the highest reward value as the optimal network and connecting the user terminal to the optimal network; Wherein, the first reward value calculation model calculates the reward value through the following steps: Based on the target service type, calculate a target preference vector, and select a target utility function for each of the network parameters from a set of utility functions; Call the target utility function to calculate the utility values of each of the network parameters, and construct a utility matrix; Calculate the inner product of the target preference vector and the utility matrix to obtain the reward values of each of the candidate networks.
9. A super-dense heterogeneous network access device, characterized in that It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the ultra-dense heterogeneous network access method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program. Wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the ultra-dense heterogeneous network access method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Handwriting skeleton refinement method, system and equipment based on deep reinforcement learning, and medium
CN117011856A
Load frequency control method based on deep reinforcement learning and confidence boundary
CN118889439A
Cited By
Ultra-dense heterogeneous wireless network access control method and device fusing prediction and reinforcement learning, equipment and storage medium
CN120751463A