A Fast Spectrum Allocation Method and System for Ultra-Dense Networks Based on Transfer Learning
By adopting a fast spectrum allocation method based on transfer learning in ultra-dense networks, the resource allocation delay problem caused by dynamic changes in the number of users is solved, and efficient and sensitive resource management is achieved.
Patent Information
- Application Number
- CN202411119868.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-08-15
AI Technical Summary
In super-dense networks, due to the dynamic changes in the number of users, traditional resource allocation methods based on static assumptions are difficult to respond quickly and manage resources effectively, resulting in resource allocation delays and inefficiency.
The fast spectrum allocation method based on transfer learning is adopted to construct a maximum throughput problem model, set the state, action and reward functions of the user equipment, and calculate the similarity between new users and source scene users by using Euclidean distance, cosine similarity and Pearson correlation coefficients, thereby migrating the old user strategy with the highest similarity to new users.
It realizes rapid response and efficient resource allocation under the dynamic changes in the number of users, significantly improving system throughput and resource allocation efficiency.
Smart Images

Figure CN119095170B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ultra-dense network communication, and particularly to a method and system for fast spectrum allocation in an ultra-dense network based on transfer learning. Background Art
[0002] In the research of ultra-dense network (UDN), the efficiency and flexibility of resource allocation are particularly important. UDN aims to meet the growing mobile communication demand and improve the quality of service by densely deploying a large number of small base stations in a specific area. This network layout can not only improve the spectrum utilization efficiency and reduce communication latency, but also significantly improve the user experience. The architecture of UDN can provide higher transmission rates and more reliable quality of service, thus greatly enhancing the user experience. However, the dense deployment of UDN also brings new technical challenges, especially in resource management and allocation.
[0003] Most traditional resource allocation strategies are based on static network state and environmental assumptions. Many existing research methods also mainly focus on resource allocation in specific and fixed scenarios. For example, a research proposed an improved cluster-assisted resource allocation (CARA) scheme to reduce interference between small cells in UDN, but it considered the case where the static positions of users remained unchanged. Another research proposed a local anchor point-based VC handover scheme for meaningless mobility in UDN under the condition of only considering static virtual cells (VCs), where user equipment (UE) within a VC could select multiple service access points according to different channel conditions. Similarly, other researches also explored the energy efficiency optimization problem in dense networks under the condition of fixed user positions.
[0004] However, in the actual environment, the network state and user behavior patterns are dynamically changing and uncertain. In this case, methods based on static assumptions often cannot effectively cope with changes in scenarios or environments. Especially when it comes to complex learning models, retraining the model to adapt to the new environment usually requires a large amount of time and computing resources. For example, methods using reinforcement learning models and deep learning models both need to continuously iterate and learn to obtain the optimal strategy. Whenever a new user joins the network or the external conditions of the network change, it is necessary to retrain to find the optimal strategy. This method is not only inefficient but also difficult to apply in dynamic scenarios that require quick responses.
[0005] The core challenge in resource allocation in UDN is not only to improve spectrum utilization efficiency and reduce network latency, but more critically, to be able to sensitively adapt to changing network conditions and user behavior patterns. This requires the resource management system to be not only efficient, but also able to quickly respond to changes in various scenarios. In view of this, a more flexible and adaptive strategy is needed to cope with this changing and uncertain environment. Summary of the Invention
[0006] To achieve the object of the present invention, the present application provides a fast spectrum allocation method for ultra-dense networks based on transfer learning, including:
[0007] Step S1: Based on the application scenario of dense base station deployment in ultra-dense networks, construct a maximum throughput problem model under spectrum reuse and co-channel interference conditions;
[0008] Step S2: Set the state function, action function and reward function of each user equipment, and train the maximum throughput problem model;
[0009] Step S3: After the number of users changes dynamically, calculate the similarity between the new user and the users in the source scenario according to the Euclidean distance, cosine similarity and Pearson correlation coefficient;
[0010] Step S4: Transfer the strategy of the source scenario user with the highest similarity to the new user according to the transfer learning algorithm.
[0011] In some specific embodiments, in step S1, the total system throughput is maximized under the constraints of user association and sub-channel allocation:
[0012]
[0013] In the formula, is the transmission power of base station i on sub-channel n, is the signal-to-interference-plus-noise ratio of user k and base station i on signal n at time t, C1 and C2 are binary variables, respectively representing the association relationship between the user and the base station and the channel; C3 and C4 respectively represent that a user can only select one base station and one channel within time t; C5 restricts that the transmission power of the base station cannot exceed its maximum power.
[0014] In some specific embodiments, in step S2, the user at position t is represented as a two-dimensional vector (x k (t), y k (t)), each agent in the network needs to know the information of the entire network to determine the action selection, and at the same time each agent needs to determine the optimal strategy according to its own position. Therefore, the state of user k consists of a channel selection matrix, a power selection matrix and a position state:
[0015] sk \(\pi(t)=\{C(t),B(t),d\ k (t)\}
[0016] where \(s k (t)\) is the state observed by user \(k\) at time slot \(t\), \(C(t)\) is the set of channel selections of all users, and \(B(t)\) is the set of associated base station selections of all users.
[0017] In some specific embodiments, according to the base station location and signal strength in the dynamic network environment, the optimal associated base station and channel of the user are determined, expressed as:
[0018] a k (t)=\{b k (t),c k (t)\}
[0019] where \(a k (t)\) is the action selection of user \(k\) at time \(t\); \(b k (t)\) is the base station selection of user \(k\) at time \(t\); \(c k (t)\) is the channel selection of user \(k\) at time \(t\).
[0020] In some specific embodiments, according to the system throughput, the agent is incentivized to take actions to improve the throughput, and the incentive function is expressed as:
[0021]
[0022] where \(B\) is the set of base station indices, \(U\) is the set of users, \(n\) is the selected channel, \(r i is the service radius of base station \(i\), and \(d\) is the service range of the user connected to the associated base station.
[0023] In some specific embodiments, in step S3, the movement path and channel state of each user are determined according to the following formula:
[0024] U i =\{Path i ,Ch i \}
[0025] where \(U i is the information of each user; \(Path i represents the movement path information of the \(i\)-th user; \(Ch i represents the channel state information of the \(i\)-th user.
[0026] In some of these specific embodiments, in step S3, using the linear interpolation method, the movement path of each user from its starting position to the ending position is divided into m sampling points, and the spatial similarity and the movement direction similarity when the new user and the source scenario user move are calculated respectively according to the Euclidean distance and the cosine similarity.
[0027] In some of these specific embodiments, in step S3, the Pearson correlation coefficient is used to measure the similarity between two users in terms of channel state and the linear correlation of channel gains. For each sub-channel, the Pearson correlation coefficient between the channel gains of the two users is calculated, and the average value is taken as the overall channel similarity.
[0028] In some of these specific embodiments, two weights are designed to calculate the comprehensive similarity of the spatial similarity and the movement direction similarity and the similarity determined according to the relative importance of the path and the channel. To achieve the same invention purpose, the present application also provides a super-dense network fast spectrum allocation system based on transfer learning, including:
[0029] Model construction module: used to construct a maximum throughput problem model under spectrum reuse and co-channel interference conditions based on the application scenario of dense base station deployment in a super-dense network;
[0030] Function determination module: used to set the state function, action function and reward function of each user equipment, and train the maximum throughput problem model;
[0031] Similarity calculation module: used to calculate the similarity between the new user and the source scenario user according to the Euclidean distance, cosine similarity and Pearson correlation coefficient after the number of users changes dynamically;
[0032] Policy migration module: used to migrate the policy of the source scenario user with the highest similarity to the new user according to the transfer learning algorithm.
[0033] Advantages of the above technical solutions:
[0034] The present application constructs a scenario model for the dynamic change of the number of users in a super-dense network and proposes a resource allocation algorithm based on transfer learning for the problem of resource allocation delay caused by the dynamic change of the number of users under a super-dense network.
[0035] The present application designs a comprehensive similarity calculation to analyze the correlation between new and old users. By finding the old user with the most similar characteristics to the newly added user and migrating its optimized policy to the new user as the initial policy for further training, the adaptation speed of the neural network to user changes is accelerated.
[0036] The simulation experiment results of this application show that in the scenario where the number of users is increasing continuously, the resource allocation algorithm adopting the transfer learning strategy demonstrates higher efficiency and better effect in terms of the speed of learning the optimal strategy and the optimization of system throughput compared with the traditional deep reinforcement learning algorithm. Brief Description of the Drawings
[0037] To more clearly illustrate the technical solutions in the embodiments of this application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0038] Figure 1 Flowchart of a fast spectrum allocation method for ultra-dense networks based on transfer learning provided by an embodiment of the present invention;
[0039] Figure 2 Scenario diagram of the change in the number of users in an ultra-dense network provided by an embodiment of the present invention;
[0040] Figure 3 Flowchart of the transfer learning algorithm provided by an embodiment of the present invention;
[0041] Figure 4 Comparison diagram of model utilities under different learning rates provided by an embodiment of the present invention;
[0042] Figure 5 Comparison diagram of model utilities under different neuron scales provided by an embodiment of the present invention;
[0043] Figure 6 Comparison diagram of utilities under different learning models with six users provided by an embodiment of the present invention;
[0044] Figure 7 Comparison diagram of utilities under different learning models with seven users provided by an embodiment of the present invention;
[0045] Figure 8 Comparison diagram of system throughput of different learning models in the middle stage of training provided by an embodiment of the present invention;
[0046] Figure 9 Comparison diagram of system throughput of different learning models in the late stage of training provided by an embodiment of the present invention;
[0047] Figure 10 Structure diagram of a fast spectrum allocation system for ultra-dense networks based on transfer learning provided by an embodiment of the present invention. Detailed Embodiments
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0049] Examples of the embodiments are shown in the accompanying drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.
[0050] Embodiment 1
[0051] An embodiment of the present invention provides a fast spectrum allocation method for a hyper-dense network based on transfer learning. Referring to Figure 1 as shown, it includes:
[0052] Step S1: Based on the application scenario of dense base station deployment in a hyper-dense network, construct a maximum throughput problem model under spectrum reuse and co-channel interference conditions.
[0053] In a specific embodiment of the present invention, referring to Figure 2 as shown, consider a scenario where the number of users in a hyper-dense heterogeneous network changes dynamically. Each user moves randomly on its moving path. As Figure 1 shown, the source scenario includes a macro base station (MBS) and F femto base stations. The index set of the base stations is represented as B = {0, 1, 2, 3... F}, where 0 represents the macro base station index. The total user set of the network can be represented as U = {0, 1, 2, 3... K - 1}.
[0054] The macro base station and each small base station share the same spectrum resource. At the same time, each user equipment (UE) can only be associated with one base station. Both macro user equipment (MUE) and small base station user equipment (SUE) can use the same spectrum resource. Considering that the base station can have different UE sets, the association vector of user k is represented by When the k-th user selects to associate with base station i, Otherwise, Assume that each UE selects at most one BS at the same time. That is:
[0055]
[0056] The spectrum resource of each base station is allocated to the UE associated with it. For the k-th user, the binary channel allocation vector is represented by If the k-th UE selects channel n at time t, then Otherwise Assume that each UE can select at most one channel at any time t, that is:
[0057]
[0058] Since the base station group in the scenario shares the same spectrum resource, users will be interfered by co-channel users. Therefore, in time slot t, the signal-to-interference-plus-noise ratio (SINR) of the k-th UE on sub-channel n can be expressed as:
[0059]
[0060] In the formula, is the transmission power of base station i on sub-channel n, which is evenly distributed by the associated base station according to the number of users associated with the base station; σ 2 is the Gaussian white noise; represents the channel gain of base station k serving the i-th user on channel n at time t. It is a signal transmission loss close to reality and consists of two parts: path loss and Rayleigh fading, and can be expressed as:
[0061]
[0062] In the formula, L is the modified path loss constant; d k (t) represents the straight-line distance from base station f k to its served user k at time t; α represents the path loss exponent; represents the Rayleigh fading of base station f k and its served user on sub-channel n at time t. Therefore, through the Shannon formula, the total capacity of user k at time t can be obtained as:
[0063]
[0064] The total channel capacity of the system at time t can be expressed as:
[0065]
[0066] Specifically, considering the maximization of the total system throughput under the constraints of user association and sub-channel allocation, the problem can be modeled as:
[0067]
[0068] In the formula, is the transmission power of base station i on sub-channel n, is the signal-to-interference-plus-noise ratio of user k and base station i on signal n at time t. C1 and C2 are binary variables, representing the association relationship between the user and the base station and the channel respectively; C3 and C4 represent that a user can only select one base station and one channel at time t; C5 restricts the transmit power of the base station not to exceed its maximum power.
[0069] Step S2: Set the state function, action function and reward function of each user equipment, and train the maximum throughput problem model.
[0070] In a specific embodiment of the present invention, in the UDN scenario, all participants, i.e., UEs, are selfish and hope to obtain the maximum long-term reward by selecting the best base station and channel. At any time t, the reward of the participant is affected by the current state of the network environment and the behaviors of other participants. At the next moment, it will be converted into a completely new random state, which is affected by the previous states of all participants and the selected actions. Therefore, we can regard this optimization problem as a competitive game. By using the same reward for all agents, it is converted into a completely cooperative game to ensure the performance of the global network.
[0071] In a specific embodiment of the present invention, state: Let s k (t) represent the state observed by user k in the t-th time slot, used to characterize the UDN environment. Public information (such as the action selection information of each FUE) can be obtained from the MBS at the beginning of each time slot. The position of the user at time t can be represented as a two-dimensional vector (x k (t), y k (t)), denoted as Convert the user position into a one-dimensional state for storage. Each agent in the network needs to know the information of the entire network to determine the action selection, and at the same time each agent needs to determine the optimal strategy according to its own position. Therefore, the state of user k consists of a channel selection matrix, a power selection matrix and a position state:
[0072] s k (t) = {C(t), B(t), d k (t)} (8)
[0073] where s k (t) is the state observed by user k in the t-th time slot, C(t) is the set of all user channel selections, and B(t) is the set of all user associated base station selections.
[0074] In a specific embodiment of the present invention, according to the base station position and signal strength in the dynamic network environment, determine the best associated base station and channel of the user.
[0075] Action: In the UDN, the key to the user's action design lies in intelligently selecting the best associated base station and channel. This process requires the user to comprehensively consider various factors such as the base station location and signal strength in a dynamic network environment to ensure the accurate selection of the associated base station. Using an intelligent resource allocation strategy, the user can obtain a more flexible and reliable connection in the UDN, thereby significantly improving the overall performance of the network. Therefore, it can be expressed as:
[0076] a k (t) = {b k (t), c k (t)} (9)
[0077] where a k (t) is the action selection of user k at time t; b k (t) is the base station selection of user k at time t; c k (t) is the channel selection of user k at time t.
[0078] In a specific embodiment of the present invention, according to the system throughput, the agent is incentivized to take actions to improve the throughput.
[0079] Reward: The reward function is based on the system throughput, incentivizing the agent to take actions to improve the throughput. At the same time, all agents share the same reward, forming a fully cooperative game to ensure the overall network performance. In particular, when the user connection exceeds the service range of the associated base station, the reward is zero, ensuring that the user takes into account the coverage range when selecting the base station. Specifically, it can be expressed as follows:
[0080]
[0081] where B is the set of base station indices, U is the set of users, n is the selected channel, r i is the service radius of base station i, and d is the service range of the user connection to the associated base station.
[0082] Step S3: After the number of users changes dynamically, calculate the similarity between the new user and the source scenario users according to the Euclidean distance, cosine similarity, and Pearson correlation coefficient.
[0083] In a specific embodiment of the present invention, in step S3, the movement path and channel state of each user are determined according to the following formula:
[0084] U i = {Path i , Ch i}
[0085] where U i is the information of each user; Path i represents the movement path information of the i-th user; Chi Represents the channel state information of the i-th user.
[0086] In a specific embodiment of the present invention, in step S3, using the linear interpolation method, the movement path of each user from its starting position to the ending position is divided into m sampling points, and the spatial similarity and the movement direction similarity when the new user and the source scenario user move are calculated respectively according to the Euclidean distance and the cosine similarity.
[0087] In a specific embodiment of the present invention, in step S3, the Pearson correlation coefficient is used to measure the similarity between two users in terms of channel state and the linear correlation of channel gains. For each sub-channel, the Pearson correlation coefficient between the channel gains of two users is calculated, and the average value is taken as the overall channel similarity.
[0088] In a specific embodiment of the present invention, two weights are designed to calculate the comprehensive similarity of the spatial similarity and the movement direction similarity respectively, and the similarity determined according to the relative importance of the path and the channel.
[0089] Specifically, when calculating the similarity between users, three main metrics are considered: Euclidean distance, cosine similarity, and Pearson correlation coefficient. Each metric emphasizes different aspects of similarity and plays a key role in different application scenarios. These three metrics are introduced in detail below.
[0090] The Euclidean distance is the most intuitive and easiest-to-understand way to measure the distance between vectors, used to calculate the actual distance between two points in an n-dimensional space. The calculation formula is shown as follows:
[0091]
[0092] where A = (a1, a2,..., a n ), B = (b1, b2,..., b n ) represent two n-dimensional vectors respectively. At this time, the similarity between vectors can be defined by formula (13). Among them, sim(A, B) ∈ (0, 1). The larger the distance, the closer to 0, and the lower the similarity; on the contrary, the smaller the distance, the closer to 1, and the higher the similarity.
[0093]
[0094] The cosine similarity focuses on reflecting the difference in the direction of two vectors. Specifically, it measures the cosine value of the angle between two vectors, and judges whether the vectors are similar according to the cosine value. Formula (14) represents the calculation method of the cosine value between two vectors, and thus the calculation method of the cosine similarity between vector A and vector B can be deduced as shown in formula (15).
[0095] A·B = ||A||||B||cosθ (13)
[0096]
[0097] Among them, the measure of cosine similarity is usually used in the positive space, that is, it is defaulted that the directions of two vectors are the same. Therefore, cossim(A, B) ∈ [0, 1]. When the cosine similarity is 0, it is regarded as the least similarity between two vectors. When the cosine similarity is 1, it is regarded as the two vectors having exactly the same direction, that is, the highest similarity degree is regarded as such.
[0098] The Pearson correlation coefficient is an index to measure the linear relationship between two sets of data. Its value ranges between -1 and 1. Among them, 1 indicates a perfect positive correlation, -1 indicates a perfect negative correlation, and 0 indicates no correlation. Its calculation formula is as follows:
[0099]
[0100] Among them, r is the Pearson coefficient between vector A and vector B; A i and B i are individual observations of random variables A and B; and are the means of A and B respectively.
[0101] The ultra-dense network scenario considered in this chapter contains K users, and each user will move randomly on its moving path trajectory. At any time t, there exists a two-dimensional coordinate. Among them, the two-dimensional coordinate of the old user i is expressed as (x i,t , y i,t ), and the two-dimensional coordinate of the new user j is expressed as (x j,t , y j,t ). Therefore, when calculating the path similarity between the old user and the new user, not only the starting and ending positions should be considered, but also each position on the path of the user should be taken into account, so as to ensure that the user with the highest path similarity to the new user can be selected as the source user to ensure the reliability of transfer learning.
[0102] In the ultra-dense network, the moving path and channel state of users have a profound impact on the network performance. In the scenario, each user has its unique moving path and channel state. To formalize this information, the information of each user can be expressed as U i , more specifically
[0103] U i = {Path i , Ch i} (16)
[0104] Among them, Path iRepresents the movement path information of the i-th user; Ch i Represents the channel state information of the i-th user.
[0105] To train the model faster in the new environment, a similarity-based transfer learning strategy is proposed. In this part, we will introduce in detail how to calculate the comprehensive similarity between users and how to perform effective transfer learning based on this similarity.
[0106] The movement paths of users are a key part of their behavior patterns. In this chapter, it is assumed that the movement paths of all users are straight lines. To more accurately determine the similarity between users, the Euclidean distance and cosine similarity are combined for calculation. First, using the linear interpolation method, the movement path of each user from its starting position to the ending position is divided into m sampling points, which can provide a high-precision approximation of the user's movement pattern.
[0107]
[0108] For any two users A and B, calculate their Euclidean distance and cosine similarity at each sampling point. In this way, not only the similarity in space (through the Euclidean distance) is captured, but also the similarity in their movement directions (through the cosine similarity) can be known.
[0109]
[0110]
[0111] Finally, to combine these two similarities, a weight w1 is used to weight and combine them.
[0112]
[0113] In this way, it can not only ensure that the similarities in space and direction are considered, but also balance the influence of the two by adjusting w1, making the final similarity evaluation more accurate.
[0114] When considering the similarity between users, in addition to the spatial path, the channel state also plays a crucial role. In the system, each user can choose any one of the available sub-channels for communication. The gains of these sub-channels follow the Rayleigh distribution because in wireless communication, the signal fading caused by the multipath effect often shows the characteristics of the Rayleigh distribution.
[0115] To measure the similarity between two users in terms of channel state, the Pearson correlation coefficient is used to measure the linear correlation of their channel gains. For each sub-channel, we calculate the Pearson correlation coefficient between the channel gains of the two users and take the average as the overall channel similarity.
[0116]
[0117] Among them, and are the average Rayleigh fading of user a and user b on all sub-channels respectively.
[0118] Considering that both the mobile path and channel characteristics affect the communication performance of users, a comprehensive similarity score is calculated by combining these two similarities. By using the weighted average method, according to the relative importance of the path and the channel, a weight w2 is given, and the specific formula is as follows:
[0119]
[0120] Based on this similarity, it can be decided which user to migrate to a new user, so as to train the model faster in the new environment.
[0121] Step S4: Transfer the policy of the source scenario user with the highest similarity to the new user according to the transfer learning algorithm.
[0122] In a specific embodiment of the present invention, based on the above calculation and analysis of the similarity between the old users and the newly added users in the source scenario, the old user with the highest similarity to the newly added user is selected from the historical scenario as the migration object. Considering that the resource optimization for each old user in the source scenario adopts the deep reinforcement learning algorithm, the transfer learning algorithm proposed in this part adopts the transfer of the policy and experience of the most similar user in the source scenario.
[0123] Specifically, referring to Figure 3 shown below, the following are the steps and considerations of this transfer learning algorithm:
[0124] Similarity calculation: Calculate the comprehensive similarity between the new user and each old user in the source scenario according to Equation (23), and select the old user with the highest similarity to the new user as the migration object.
[0125] Policy transfer: Transfer the policy of this old user to the newly added user, which means that the new user will adopt the same or similar decision-making process as the old user with the highest similarity.
[0126] Experience reuse: Considering that deep reinforcement learning depends on the experience pool to train and update its policy, part of the content of the experience pool of the old user with the highest similarity is transferred to the experience pool of the new user. It provides a richer learning starting point for the new user, enabling it to learn and adjust its policy faster.
[0127] Reduced exploration rate training: Since the experience and policy have been transferred, new users can start with a lower exploration rate and focus on exploiting the transferred policy. Over time, adjust the exploration rate gradually during the training process based on the feedback from the new environment.
[0128] The present paper provides a simulation experiment Figure 4 showing the utility comparison of different learning rate settings in the deep reinforcement learning model, including four values: 0.0001, 0.00001, 0.00005, and 0.000005. The analysis shows that the learning rate configuration of 0.00005 achieves the optimal convergence speed and performance stability during the model training process. In contrast, a higher learning rate such as 0.0001 rapidly improves the performance at the beginning but then shows diminishing returns; while a lower learning rate such as 0.000005 shows slow progress throughout the training cycle. Based on this, this paper's research determines 0.00005 as the optimal learning rate to facilitate the model's fast learning and stable performance in efficiently handling the ultra-dense network resource allocation task.
[0129] Figure 5 The performance of deep reinforcement learning models with different neuron scales in the ultra-dense network resource allocation task is compared. This study explores four network configurations: 32-32-32, 64-64-64, 128-128-128, and 256-256-256. It is found that although the learning speeds of each configuration are similar at the beginning of training, as the training progresses, the 128-128-128 configuration shows significant advantages in performance, demonstrating the effectiveness of a moderately sized network structure in handling resource allocation problems with complex decision-making and dynamic environmental changes. Selecting too small a network may not fully capture the complexity of the environment, while too large a network may slow down the learning efficiency due to the computational burden. Therefore, a medium-sized network structure provides an ideal balance for the tasks in this chapter, taking into account both performance and computational efficiency.
[0130] Figure 6Shows the performance comparison of different deep reinforcement learning strategies in a six-user scenario in an ultra-dense network environment. This study compared three training strategies: one is the basic deep reinforcement learning with an initial exploration rate set to 0.1; the second is the enhanced exploration deep reinforcement learning with an initial exploration rate set to 0.5; the third is transfer learning, which applies the strategy of the most similar existing old user to the newly added user and sets the initial exploration rate to 0.5. The experimental results reveal that although the basic deep reinforcement learning strategy has a slower convergence speed, it reaches performance stability after 1200 iterations, while the enhanced exploration deep reinforcement learning and transfer learning strategies achieve fast convergence within 600 iterations. At the same time, the transfer learning strategy not only accelerates the convergence process but also achieves the best final performance, which reflects that the combination of an appropriate exploration rate and effective knowledge transfer can significantly improve the learning efficiency and resource allocation ability of the model. In contrast, the performance of the enhanced exploration deep reinforcement learning strategy is slightly inferior, perhaps because the higher exploration rate fails to find a sufficiently optimized strategy within the given number of iterations, highlighting the need for a more refined balance in the selection of the exploration rate.
[0131] Figure 7 Compares the performance of the basic deep reinforcement learning strategy and the transfer learning strategy in the ultra-dense network resource allocation task in a seven-user scenario. The basic deep reinforcement learning strategy is trained from scratch, while the transfer learning strategy is adjusted based on the training of the previous six-user scenario to adapt to the newly added user. The results show that the transfer learning strategy quickly surpasses the basic strategy in the initial stage of training, demonstrating the potential to accelerate the adaptation to new scenarios by leveraging existing learning results. During the iteration process, the transfer learning strategy quickly reaches a high performance level and maintains high stability, while the other two deep reinforcement learning strategies have a slower performance improvement. This result further confirms the effectiveness of transfer learning in quickly adjusting resource allocation strategies in dynamic environments.
[0132] Figure 8 Details the comparison of the average system throughput performance under different methods in the iteration range of 700 to 900. The results show that in the middle stage of training, the deep reinforcement learning with low exploration degree and transfer learning reach a higher system throughput faster than the traditional deep reinforcement learning. Figure 9 Shows the comparison of the system throughput in the iteration stage of 1300 to 1500. In this late training stage, due to the too low exploration degree of the deep reinforcement learning strategy with low exploration degree, it shows a relatively weak state in subsequent iterations, resulting in a lower system throughput. However, the transfer learning algorithm proposed in this paper can still maintain high performance in the late training stage, surpassing the traditional deep reinforcement learning and achieving a higher total system throughput.
[0133] Example Two
[0134] An embodiment of the present invention provides a fast spectrum allocation system for ultra-dense networks based on transfer learning. Refer to Figure 10 As shown in the figure, it includes:
[0135] Model construction module 10: used to construct a maximum throughput problem model under spectrum reuse and co-channel interference conditions based on the application scenario of dense base station deployment in an ultra-dense network;
[0136] Function determination module 20: used to set the state function, action function and reward function of each user equipment, and train the maximum throughput problem model;
[0137] Similarity calculation module 30: used to calculate the similarity between new users and source scenario users according to Euclidean distance, cosine similarity and Pearson correlation coefficient after the number of users changes dynamically;
[0138] Policy migration module 40: used to migrate the policy of the source scenario user with the highest similarity to the new user according to the transfer learning algorithm.
[0139] In a specific embodiment of the present invention, the total system throughput is maximized under the conditions of user association and sub-channel allocation restrictions:
[0140]
[0141] In the formula, is the transmission power of base station i on sub-channel n, is the signal-to-interference-plus-noise ratio of user k and base station i on signal n at time t. C1 and C2 are binary variables, representing the association relationship between the user and the base station and the channel respectively; C3 and C4 respectively represent that a user can only select one base station and one channel within time t; C5 restricts that the transmission power of the base station cannot exceed its maximum power.
[0142] In a specific embodiment of the present invention, the user is represented as a two-dimensional vector (x k (t), y k (t)) at position t. Each agent in the network needs to know the information of the entire network to determine the action selection. At the same time, each agent needs to determine the optimal policy according to its own position. Therefore, the state of user k consists of a channel selection matrix, a power selection matrix and a position state:
[0143] s k (t) = {C(t), B(t), d k (t)}
[0144] Among them, s k(t) is the state observed by user k at time slot t, C(t) is the set of channel selections of all users, and B(t) is the set of associated base station selections of all users.
[0145] In a specific embodiment of the present invention, based on the base station location and signal strength in the dynamic network environment, the optimal associated base station and channel of the user are determined, expressed as:
[0146] a k (t) = {b k (t), c k (t)}
[0147] where a k (t) is the action selection of user k at time t; b k (t) is the base station selection of user k at time t; c k (t) is the channel selection of user k at time t.
[0148] In a specific embodiment of the present invention, according to the system throughput, the agent is incentivized to take actions to improve the throughput, and the incentive function is expressed as:
[0149]
[0150] where B is the set of base station indices, U is the set of users, n is the selected channel, r i is the service radius of base station i, and d is the service range of the user connected to the associated base station.
[0151] In a specific embodiment of the present invention, the movement path and channel state of each user are determined according to the following formula:
[0152] U i = {Path i , Ch i}
[0153] where U i is the information of each user; Path i represents the movement path information of the i-th user; Ch i represents the channel state information of the i-th user.
[0154] In a specific embodiment of the present invention, using the linear interpolation method, the movement path of each user from its starting position to the ending position is divided into m sampling points, and the spatial similarity and movement direction similarity when the new user and the source scenario user move are calculated according to the Euclidean distance and cosine similarity respectively.
[0155] In a specific embodiment of the present invention, the similarity of channel states between two users and the linear correlation of channel gains are measured according to the Pearson correlation coefficient. For each sub-channel, the Pearson correlation coefficient between the channel gains of two users is calculated, and the average value is taken as the overall channel similarity.
[0156] In a specific embodiment of the present invention, two weights are designed to calculate the comprehensive similarity of the spatial similarity and the moving direction similarity, and the similarity determined according to the relative importance of the path and the channel.
[0157] The present invention provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the fast spectrum resource optimization method based on transfer learning according to the first aspect.
[0158] Optionally, the above electronic device may be a server.
[0159] In addition, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the fast spectrum resource optimization method based on transfer learning as described in the first aspect.
[0160] The rapid development of network technology has promoted the densification of network infrastructure. Aiming at the problem that the trained allocation strategy is difficult to quickly adapt to network changes due to the continuous change of the number of users in the ultra-dense network, a model considering the problem of maximizing the long-term throughput of the network after the dynamic increase of the number of users is constructed, and a fast resource allocation algorithm based on transfer learning is proposed. In the source scenario, each user has its fixed moving path, and deep reinforcement learning is used to optimize the strategy of associating base stations and selecting channels. After the number of users changes dynamically, a comprehensive similarity calculation method is designed to analyze the similarity between the newly added users and the old users in the source scenario, and the strategy of the old user with the highest similarity is migrated to the new user as the initial strategy and then further trained and optimized. The simulation results show that the resource allocation algorithm adopting the transfer learning strategy shows higher efficiency and effect in terms of the speed of learning the optimal strategy and the optimization of system throughput compared with the traditional deep reinforcement learning method.
[0161] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0162] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1Steps of the functions specified in one or more boxes. Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention. Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0163] The above has introduced the method and device provided by the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.
[0164] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one specific embodiment" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic description of the terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0165] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A fast spectrum allocation method for ultra-dense networks based on transfer learning, characterized in that: include: Step S1: Based on the application scenario of dense base station deployment in an ultra-dense network, a maximum throughput problem model under spectrum reuse and co-channel interference conditions is constructed; Step S2: setting the state function, action function and reward function of each user device, and training the maximum throughput problem model; Step S3: according to the maximum throughput problem model, after the number of users changes dynamically, the similarity between the new user and the source scenario user is calculated according to the Euclidean distance, cosine similarity and Pearson correlation coefficient; Step S4: Migrating the strategy of the source scenario user with the highest similarity to the new user according to the transfer learning algorithm; In step S3, a linear interpolation method is used to divide the moving path of each user from its starting position to its ending position into m sampling points, and the spatial similarity and moving direction similarity between the new user and the source scene user are calculated respectively according to the Euclidean distance and the cosine similarity; In step S3, the similarity of the channel states of the two users and the linear correlation of the channel gains are measured according to the Pearson correlation coefficient. For each subchannel, the Pearson correlation coefficient between the channel gains of the two users is calculated, and the average value is taken as the overall channel similarity; Two weights are designed to respectively calculate the comprehensive similarity of the spatial similarity and the moving direction similarity, and the final similarity is determined in combination with the channel similarity.
2. The ultra-dense network fast spectrum allocation method based on transfer learning according to claim 1, characterized in that: In step S1, the maximum system total throughput is calculated under the constraints of user association and subchannel allocation: In the formula, is the transmission power of base station i on subchannel n, is the signal-to-interference-to-noise ratio of user k and base station i on signal n at time t, C1 and C2 are binary variables, representing the association between the user, base station and channel respectively; C3 and C4 respectively represent that a user can only select one base station and one channel at time t; C5 is the limit on the base station's transmission power not to exceed its maximum power, is the association vector of user k, Assign vector to the kth user binary channel, is the association vector between user k and base station i at time t, is the allocation vector of user k on signal n at time t.
3. The ultra-dense network fast spectrum allocation method based on transfer learning according to claim 1, characterized in that: In step S2, the user at time t is represented as a two-dimensional vector (x k (t),y k (t)), each agent in the network needs to know the information of the entire network to decide the action selection, and each agent needs to determine the optimal strategy according to its own position. Therefore, the user k state is composed of the channel selection matrix, the power selection matrix and the position state: s k (t)={C(t),B(t),d k (t)} Among them, s k (t) is the state observed by user k at time t, C(t) is the set of channel selection matrices of all users, B(t) is the set of power selection matrices of all user-associated base stations, d k (t) The location status of user k at time t.
4. The method for fast spectrum allocation in ultra-dense networks based on transfer learning according to claim 3, characterized in that: According to the base station location and signal strength in the dynamic network environment, the best associated base station and channel for the user are determined, which is expressed as: a k (t)={b k (t),c k (t)} Among them, a k (t) is the action selected by user k at time t; b k (t) is the base station selection of user k at time t; c k (t) is the channel selection of user k at time t.
5. The ultra-dense network fast spectrum allocation method based on transfer learning according to claim 4 is characterized in that: According to the system throughput, the agent is encouraged to take actions to improve the throughput. The reward function is expressed as: Among them, B is the base station index set, U is the user set, n is the selected channel, r i is the service radius of base station i, d is the service range of the user-associated base station, is the signal to interference and noise ratio of user k and base station i on signal n at time t.
6. The method for fast spectrum allocation in ultra-dense networks based on transfer learning according to claim 1, characterized in that: In step S3, the moving path and channel state of each user are determined according to the following formula: U i ={Path i ,Ch i } Among them, U i Information for each user; Path i represents the mobile path information of the i-th user; Ch i represents the channel state information of the i-th user.
7. A system for fast spectrum allocation in ultra-dense networks based on transfer learning, the system being used to implement the method according to any one of claims 1 to 6, characterized in that: include: Model building module: used to build a maximum throughput problem model under spectrum reuse and co-channel interference conditions based on the application scenario of dense base station deployment in ultra-dense networks; Function determination module: used to set the state function, action function and reward function of each user device and train the maximum throughput problem model; Similarity calculation module: used to calculate the similarity between new users and source scenario users based on Euclidean distance, cosine similarity and Pearson correlation coefficient after the number of users changes dynamically; Strategy migration module: used to migrate the strategy of the source scenario user with the highest similarity to the new user according to the migration learning algorithm.
Citation Information
Patent Citations
Unmanned aerial vehicle base station dynamic deployment method based on user trajectory prediction
CN114710786A
Multi-resource optimization method for isomerized semantic and bit communication network
CN117858124A