A reinforcement learning-based method for enhancing user association in dynamic air-ground heterogeneous networks
By optimizing the association between users and base stations in an integrated air-ground heterogeneous network using reinforcement learning, this approach solves the complex problems caused by non-uniform user distribution and dual mobility, maximizing network capacity and achieving rapid convergence, thus improving network performance.
Patent Information
- Application Number
- CN202310493523.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-04
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2043-05-04
AI Technical Summary
In the context of integrated air-ground heterogeneous networks, the relationship between users and base stations is complex, especially affected by the non-uniform distribution of users, the dual mobility of air base stations and ground users, and the heterogeneity of network structure. Existing research is unable to effectively optimize user access selection and resource allocation, leading to a decline in network performance.
A reinforcement learning-based dynamic air-ground heterogeneous network user association enhancement method is adopted. By constructing an air-ground heterogeneous network model and initializing parameters, the optimization problem is solved based on the Q-learning reinforcement algorithm. Considering the joint processing of air base stations and ground base stations, the association between users and base stations is optimized. The Q-value table is used to train iteratively to obtain the user association solution that maximizes the channel capacity.
It effectively improved the overall network capacity, shortened the algorithm convergence time, and enhanced network coverage and quality, adapting to user association needs in dynamic scenarios.
Smart Images

Figure CN116634450B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication network technology, and more specifically, to a method for enhancing user association in dynamic air-ground heterogeneous networks based on reinforcement learning. Background Technology
[0002] With the development of communication technology, network architecture is gradually changing, and the concept of integrated air-space network has been proposed, which will adopt a multi-layered and heterogeneous network architecture. The integrated air-ground heterogeneous network is a new type of network system that combines terrestrial and air networks. In the field of wireless communication, the new air-ground heterogeneous network structure can effectively improve network reliability, increase network bandwidth and channel capacity, and solve the explosive demand for user services.
[0003] However, user association and resource allocation issues in heterogeneous air-ground network scenarios remain a major concern. The new network structure has various types of base stations and terminals, which will generate complex user association problems. The advantages and characteristics of the new network structure require new user association strategies to realize their benefits. User association strategies can optimize user access selection, maximize system benefits, avoid resource waste, and improve the overall network coverage and quality.
[0004] However, changes in network structure have brought new challenges to user association strategies. In heterogeneous network scenarios integrating air and ground, the non-uniformity of user distribution leads to the non-uniformity of ground base station deployment, which makes the spatial correlation between users and base stations very complex. At the same time, the dual mobility of air base stations and users, the heterogeneity of network structure, and the interweaving of multiple frequency bands all bring certain difficulties to user association strategies.
[0005] Therefore, how to rationally allocate network system resources in a new network environment and how to help users choose appropriate access points in a complex network structure are issues that cannot be ignored in improving overall network performance and ensuring user service quality. Summary of the Invention
[0006] To address the issue of user-base station association in dynamic scenarios of heterogeneous air-ground integrated networks, this application provides a reinforcement learning-based method for enhancing user association in dynamic heterogeneous air-ground networks.
[0007] The embodiments of this application are implemented as follows:
[0008] This application provides a method for enhancing user association in dynamic air-ground heterogeneous networks based on reinforcement learning, including:
[0009] Construct an air-ground heterogeneous network and initialize the network parameters. The air-ground heterogeneous network includes ground macro base stations, ground small base stations, airborne base stations, and ground users.
[0010] The system channels and channel capacity are modeled based on the aforementioned air-ground heterogeneous network;
[0011] Determine the optimization problem that maximizes network capacity and its constraints;
[0012] The optimization problem is solved using the Q-learning reinforcement algorithm to obtain the optimal user association solution.
[0013] In one possible implementation, the construction of the air-to-ground heterogeneous network and the initialization of network parameters, wherein the air-to-ground heterogeneous network includes ground macro base stations, ground small base stations, airborne base stations, and ground users, including:
[0014] Given the set of the airborne base stations as The set of ground base stations is The total set of base stations is The set of ground users is ;
[0015] in, The number of the aforementioned airborne base stations. The number of the ground base stations. The number of ground users is defined as follows: the ground base station includes ground macro base stations and ground small base stations; the overall base station includes the ground base stations and the airborne base stations.
[0016] In one possible implementation, the construction of the air-to-ground heterogeneous network and the initialization of network parameters, wherein the air-to-ground heterogeneous network includes terrestrial macro base stations, terrestrial small base stations, airborne base stations, and terrestrial users, includes:
[0017] Given the beamwidth of the airborne base station as The radius of the ground coverage area of the airborne base station is ;
[0018] in, This refers to the flight altitude of the drone;
[0019] The trajectory of the airborne base station is a circle centered on the ground macro base station with a radius of [missing information]. Circular motion, for The location coordinates of the aforementioned airborne base station at that time. For the air base station in The location at a given time, and the speed of the airborne base station. ;
[0020] in, The coverage radius of the ground macro base station. The moving speed of a single airborne base station.
[0021] In one possible implementation, the construction of the air-to-ground heterogeneous network and the initialization of network parameters, wherein the air-to-ground heterogeneous network includes terrestrial macro base stations, terrestrial small base stations, airborne base stations, and terrestrial users, and further includes:
[0022] Given the movement speed of each of the ground users as The angle of movement is The user's movement speed and movement angle are updated in real time, and the parameters are set. For the update time slot of the aforementioned air-to-ground heterogeneous network, each The speed and angle of the ground macro base station, the ground small base station, the air base station, and the ground user are updated every second, and the network topology is regenerated.
[0023] In one possible implementation, the modeling of system channels and channel capacity based on the air-to-ground heterogeneous network includes:
[0024] Given the ground user in The position of the moment is and the ground base station in The position of the moment is The distance between the ground user and the ground base station is ;
[0025] in, The altitude of the ground base station;
[0026] Due to the cross-domain and heterogeneous nature of networks, the channel gain of terrestrial communication needs to consider both large-scale and small-scale fading. At any given time, the independent channel gain between the ground user and the ground base station is ,in, This is a small-scale fading component;
[0027] According to the Jakes model, the small-scale fading components Represented as a first-order complex Gaussian Markov process: ;
[0028] in, , It is a zeroth-order Bessel function of the first kind. It is an independent and identically distributed circularly symmetric complex Gaussian random variable with unit variance;
[0029] The large-scale fading component is , The large-scale fading component is inversely correlated with the distance between the ground base station and the ground user;
[0030] in, Channel power gain per unit distance for reference. Represents the speed of light. Indicates the carrier frequency. This represents the path fading coefficient of the channel link between the ground user and the ground base station.
[0031] In one possible implementation, the modeling of system channels and channel capacity based on the air-to-ground heterogeneous network further includes:
[0032] To maintain consistency, the service channel of the airborne base station also considers both large-scale and small-scale fading. Given that the airborne base station... The position of the moment is The distance between the ground user and the airborne base station is ;
[0033] in, The altitude of the aerial base station;
[0034] Analogous to the channel model definition of the ground base station, in At any given time, the channel gain between the ground user and the airborne base station is ;
[0035] in, Let be the path fading coefficient of the channel link between the ground user and the air base station, satisfying the relationship .
[0036] In one possible implementation, the modeling of system channels and channel capacity based on the air-to-ground heterogeneous network further includes:
[0037] The signal-to-noise ratio between the ground user and the overall base station is ;
[0038] in, For the overall base station in time slot The transmission power, Indicates in Time slot, the total interference power received by the base station at the ground user;
[0039] and ;
[0040] in, This represents the received power from the overall base station received by the ground user. The power of additive white Gaussian noise is represented by [insert power here], and the number of ground users connected to the total base station in the air-to-ground heterogeneous network is represented by [insert power here]. .
[0041] In one possible implementation, the modeling of system channels and channel capacity based on the air-to-ground heterogeneous network includes:
[0042] The transmission rate between the overall base station and the ground user is ;
[0043] in, This represents the spectrum resources allocated to the total number of base stations;
[0044] All users under the same base station share bandwidth resources equally, and the total channel capacity of the system is... .
[0045] In one possible implementation, determining the optimization problem and constraints for maximizing network capacity includes:
[0046] The solution to the problem of associating the total number of base stations with the ground users is matrix A:
[0047] ;
[0048] in, Binary parameters representing the association between the ground user and the overall base station, when the ground user is associated with the overall base station, ,otherwise ;
[0049] The optimization problem for maximizing network capacity is defined as follows:
[0050] ;
[0051] Wherein, C1 is the limitation on the movement trajectory of the airborne base station, and C2 is the limitation that the total number of users served by the base station cannot exceed the maximum number of connections of the current base station. The limitation C3 is a restriction on the association of users, ensuring that each user has one and only one base station for downlink data transmission in the same time slot.
[0052] In one possible implementation, solving the optimization problem using the Q-learning reinforcement algorithm to obtain the optimal user association solution includes: decomposing the user association problem into tuples. The tuple This is a series of parts in the Markov decision-making process, among which... Represents an intelligent agent. This indicates the user association status collected by the air-to-ground base station. This indicates the action taken by the intelligent agent. Indicates an immediate reward function;
[0053] Design the Q-value function table as follows A table in the format of [format], where the number of rows is [number]. , representing the total number of state spaces. The number of columns in the table represents a certain state. The total number of actions to choose from. ,and This indicates the state. Next, action The selected Q value;
[0054] Based on tuples Using the Q-value function table, the optimal user association solution between the ground user and the overall base station is obtained.
[0055] The technical solution provided in this application can achieve at least the following beneficial effects:
[0056] This application provides a reinforcement learning-based method for enhancing user association in dynamic air-to-ground heterogeneous networks. This method considers the joint processing of airborne and ground base stations in the air-to-ground heterogeneous network scenario, particularly taking into account the dual mobility of airborne base stations and ground users. It solves the user association problem in air-to-ground heterogeneous networks using reinforcement learning, obtaining a user association solution that maximizes channel capacity through iterative training of the Q-value table. Furthermore, it ensures optimal overall network capacity even when the coverage areas of air-to-ground heterogeneous network base stations overlap. The method also improves the Q-value function table, action selection, and reward function in Q-learning, accelerating the convergence of the user association algorithm, reducing its time consumption, and effectively speeding up the convergence and improving the overall network capacity. Attached Figure Description
[0057] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This is a flowchart illustrating an exemplary embodiment of the present application of a method for enhancing user association in dynamic air-ground heterogeneous networks based on reinforcement learning;
[0059] Figure 2 This is a schematic diagram illustrating the process of constructing an air-ground heterogeneous network and initializing the network parameters, as shown in an exemplary embodiment of this application.
[0060] Figure 3This is a schematic diagram illustrating a process for modeling system channels and channel capacity based on an air-ground heterogeneous network, as shown in an exemplary embodiment of this application.
[0061] Figure 4 This is a flowchart illustrating an exemplary embodiment of this application, showing the optimization problem and constraints for maximizing network capacity.
[0062] Figure 5 This is a schematic diagram illustrating an exemplary embodiment of the present application of solving the optimization problem using the Q-learning reinforcement algorithm to obtain the optimal user association solution;
[0063] Figure 6 This is a schematic diagram illustrating the process of obtaining the optimal user association solution through the Q-learning reinforcement algorithm in an exemplary embodiment of this application;
[0064] Figure 7 This is a schematic diagram of a system model illustrating an exemplary embodiment of this application of a dynamic air-ground heterogeneous network user association enhancement method based on reinforcement learning;
[0065] Figure 8 This is a flowchart illustrating an exemplary embodiment of the Q-learning reinforcement algorithm of this application;
[0066] Figure 9 This is a schematic diagram comparing the Q-learning reinforcement algorithm and the Q-learning universal algorithm according to an exemplary embodiment of this application;
[0067] Figure 10 This is a schematic diagram comparing the system channel capacity of an exemplary embodiment of this application with that of the prior art. Detailed Implementation
[0068] To make the objectives, implementation methods and advantages of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0069] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0070] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0071] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0072] Before explaining the reinforcement learning-based method for enhancing user association in dynamic air-ground heterogeneous networks provided in this application, the application scenarios and implementation environment of this application embodiment will be introduced first.
[0073] With the development of communication technology, network architecture is gradually changing, and the concept of an integrated air-space network has been proposed, which will adopt a multi-layered and heterogeneous network architecture.
[0074] Air-ground integrated heterogeneous networks are a new type of network architecture that combines terrestrial and airborne networks. Designed for the field of wireless communication, they offer several advantages over traditional cellular networks:
[0075] First, integrated air-ground heterogeneous networks can achieve wider coverage through the deployment of airborne networks. At the same time, the flexibility and high reliability of airborne networks can ensure network connectivity in various complex environments.
[0076] Secondly, the air-ground integrated heterogeneous network combines terrestrial and airborne networks, allowing for more efficient use of spectrum resources and achieving higher data rates and capacities. Especially in high-density user areas and temporary scenarios, the airborne network can provide greater capacity and higher speeds.
[0077] Finally, the air-ground integrated heterogeneous network can select the optimal communication method according to different application scenarios, thereby achieving lower latency and energy consumption. For example, in high-speed vehicle scenarios, the air network can provide lower latency and higher bandwidth.
[0078] In conclusion, the novel air-ground heterogeneous network structure can effectively improve network reliability, increase network bandwidth and channel capacity, and address the explosive growth in user service demands.
[0079] However, user association and resource allocation issues in heterogeneous air-ground network scenarios remain a major concern. The new network structure contains various types of base stations and terminals, which can lead to complex user association problems.
[0080] The advantages and characteristics of the new network architecture require new user association strategies to realize their benefits. User association strategies can optimize user access choices, maximize system benefits, avoid resource waste, and improve the overall network coverage and quality.
[0081] However, changes in network structure have brought new challenges to user association strategies. In heterogeneous network scenarios integrating air and ground, the non-uniformity of user distribution leads to the non-uniformity of ground base station deployment, which makes the spatial correlation between users and base stations very complex.
[0082] Meanwhile, the dual mobility of airborne base stations and users, the heterogeneity of network structures, and the interweaving of multiple frequency bands all pose certain challenges to user association strategies.
[0083] Therefore, how to rationally allocate network system resources in a new network environment and how to help users choose appropriate access points in a complex network structure are issues that cannot be ignored in improving overall network performance and ensuring user service quality.
[0084] Existing research on user association strategies in heterogeneous air-ground networks mainly focuses on the coverage and energy consumption of airborne base stations, and most studies consider static scenarios, with little consideration for the dual mobility of airborne base stations and ground users.
[0085] In the face of increasingly complex network structures, considering a new scenario is of great significance for today's wireless communication.
[0086] The air-ground integrated heterogeneous wireless communication network differs significantly from terrestrial networks, as it has multiple types of base stations, making the problem of overlapping base station coverage more complex.
[0087] Air-to-ground networks require consideration of different channels and the construction of separate channel information models. Furthermore, the dynamic characteristics of these new heterogeneous air-to-ground networks are more pronounced, with rapid movement of both airborne base stations and ground users necessitating fast strategy development. These new network scenarios require user association strategies adapted to their dynamic nature, thereby improving overall network channel capacity.
[0088] Based on this, this application provides a reinforcement learning-based method for enhancing user association in dynamic air-ground heterogeneous networks. Specifically, it considers user association methods under complex network structures in integrated air-ground heterogeneous network scenarios. That is, it performs user association in scenarios where airborne base stations and terrestrial cellular heterogeneous networks coexist. It considers the different channel characteristics of air-ground networks and the dual mobility of ground users and airborne base stations, formulates an optimization problem to maximize network capacity, and proposes a Q-learning-based user association algorithm for dynamic air-ground heterogeneous networks. Through Q-learning training, the optimal user association solution under the scenario model is obtained.
[0089] The scenarios considered mainly include aerial drone base stations, ground macro base stations, ground small base stations, and mobile users. The overlap and asymmetry of base station coverage, as well as the dual mobility of aerial and ground base stations, are the main challenges to improving channel capacity in this scenario.
[0090] Next, the technical solutions of this application and how they solve the aforementioned technical problems will be described in detail through embodiments and in conjunction with the accompanying drawings. The embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application.
[0091] Figure 1 This is a flowchart illustrating an exemplary embodiment of the present application of a method for enhancing user association in dynamic air-ground heterogeneous networks based on reinforcement learning.
[0092] In one exemplary embodiment, such as Figure 1 As shown, a method for enhancing user association in dynamic air-ground heterogeneous networks based on reinforcement learning is provided. In this embodiment, the method may include the following steps:
[0093] Step 100: Construct an air-ground heterogeneous network and initialize the network parameters. The air-ground heterogeneous network includes ground macro base stations, ground small base stations, airborne base stations, and ground users.
[0094] Step 200: Model the system channels and channel capacity based on the aforementioned air-ground heterogeneous network;
[0095] Step 300: Determine the optimization problem and constraints for maximizing network capacity;
[0096] Step 400: Solve the optimization problem using the Q-learning reinforcement algorithm to obtain the optimal user association solution.
[0097] As can be seen, this embodiment solves the user association problem in air-to-ground heterogeneous networks based on reinforcement learning. Through iterative training of the Q-value table, a user association solution maximizing channel capacity is obtained. This method improves the Q-value function table, action selection, and reward function in Q-learning, effectively accelerating convergence and increasing the overall network capacity.
[0098] Figure 2 This is a schematic diagram illustrating the process of constructing an air-to-ground heterogeneous network and initializing network parameters, as shown in an exemplary embodiment of this application. Figure 7 This is a schematic diagram of a system model illustrating an exemplary embodiment of this application of a dynamic air-ground heterogeneous network user association enhancement method based on reinforcement learning.
[0099] In one possible implementation, such as Figure 2 As shown, the construction of the air-to-ground heterogeneous network and the initialization of network parameters, the air-to-ground heterogeneous network includes ground macro base stations, ground small base stations, airborne base stations, and ground users, including:
[0100] Step 110: Given the set of the airborne base stations as follows The set of ground base stations is The total set of base stations is The set of ground users is .
[0101] When initializing the various parameters of the network, it is necessary to set up a multi-layered heterogeneous network scenario that integrates air and ground, such as... Figure 7 As shown, the air-ground heterogeneous network in this embodiment includes ground macro base stations, ground small base stations, airborne base stations, and ground users. The airborne base stations are provided using drones. The number of the aforementioned airborne base stations. The number of the ground base stations. The number of ground users is defined as follows: the ground base station includes ground macro base stations and ground small base stations; the overall base station includes ground base stations and airborne base stations. In the network of this embodiment, the airborne base station mainly plays the role of expanding the communication coverage area and providing services to the edge areas of the scene.
[0102] In one possible implementation, such as Figure 2 As shown, the construction of the air-to-ground heterogeneous network and the initialization of network parameters, the air-to-ground heterogeneous network includes ground macro base stations, ground small base stations, airborne base stations, and ground users, including:
[0103] Step 120: Given the beamwidth of the airborne base station as... The radius of the ground coverage area of the airborne base station is ;
[0104] Step 130: The trajectory of the airborne base station is a circle centered on the ground macro base station with a radius of... Circular motion, for The location coordinates of the aforementioned airborne base station at that time. For the air base station in The location at a given time, and the velocity of the airborne base station, are determined by the set of velocities. To determine this.
[0105] in, The flight altitude of the drone. The coverage radius of the ground macro base station. The moving speed of a single airborne base station.
[0106] In one possible implementation, such as Figure 2 As shown, the construction of the air-to-ground heterogeneous network and the initialization of network parameters, the air-to-ground heterogeneous network includes ground macro base stations, ground small base stations, airborne base stations, and ground users, and also includes:
[0107] Step 140: Given the movement speed of each of the ground users as... The angle of movement is The user's movement speed and movement angle are updated in real time, and the parameters are set. For the update time slot of the aforementioned air-to-ground heterogeneous network, each The speed and angle of the ground macro base station, the ground small base station, the air base station, and the ground user are updated every second, and the network topology is regenerated.
[0108] The user's movement speed and angle are updated in real time, and it is assumed that the user will move in the opposite direction after encountering an edge.
[0109] As can be seen, by assuming that the altitude of the region is uniform and there are no basins or mountains, and assuming that the flight altitude of the drone base station is fixed, it is possible to simulate the dual mobility of the airborne base station and the ground user.
[0110] Figure 3 This is a schematic diagram illustrating a process for modeling system channels and channel capacity based on an air-ground heterogeneous network, as shown in an exemplary embodiment of this application.
[0111] In one possible implementation, such as Figure 3 As shown, the modeling of system channels and channel capacity based on the air-to-ground heterogeneous network includes:
[0112] Step 210: Given the ground user in The position of the moment is and the ground base station in The position of the moment is The distance between the ground user and the ground base station is ;
[0113] Step 220: Based on the cross-domain and heterogeneous nature of networks, the channel gain of terrestrial communication needs to consider both large-scale and small-scale fading. At any given time, the independent channel gain between the ground user and the ground base station is ;
[0114] Step 230: Based on the Jakes model, separate the small-scale fading components. Represented as a first-order complex Gaussian Markov process: ;
[0115] Step 240: The large-scale fading component is , The large-scale fading component is inversely correlated with the distance between the ground base station and the ground user.
[0116] in, The altitude of the ground base station. For small-scale fading components, It is an absolute value operation. , It is a zeroth-order Bessel function of the first kind. It is an independent and identically distributed circularly symmetric complex Gaussian random variable with unit variance. Channel power gain per unit distance for reference. Represents the speed of light. Indicates the carrier frequency. This represents the path fading coefficient of the channel link between the ground user and the ground base station.
[0117] In one possible implementation, such as Figure 3 As shown, the modeling of system channels and channel capacity based on the air-to-ground heterogeneous network further includes:
[0118] Step 250: To maintain consistency, the service channel of the air base station also considers both large-scale and small-scale fading. Given the air base station in... The position of the moment is The distance between the ground user and the airborne base station is ;
[0119] Step 260: Analogous to the channel model definition of the ground base station, in At any given time, the channel gain between the ground user and the airborne base station is .
[0120] in, The altitude of the aerial base station. Let be the path fading coefficient of the channel link between the ground user and the air base station, satisfying the relationship .
[0121] In one possible implementation, such as Figure 3 As shown, the modeling of system channels and channel capacity based on the air-to-ground heterogeneous network further includes:
[0122] Step 270: The signal-to-noise ratio between the ground user and the overall base station is ;
[0123] and .
[0124] in, For the overall base station in time slot The transmission power, Indicates in Time slot, the total interference power received by the base station at the ground user, This represents the received power from the overall base station received by the ground user. The power of additive white Gaussian noise is represented by [insert power here], and the number of ground users connected to the total base station in the air-to-ground heterogeneous network is represented by [insert power here]. .
[0125] In one possible implementation, such as Figure 3 As shown, the modeling of system channels and channel capacity based on the air-to-ground heterogeneous network further includes:
[0126] Step 280: The transmission rate between the overall base station and the ground user is:
[0127] ;
[0128] Step 290: All users under the same base station share bandwidth resources equally, and the total channel capacity of the system is:
[0129] .
[0130] in, This represents the spectrum resources allocated to the overall base stations.
[0131] Figure 4 This is a flowchart illustrating an exemplary embodiment of this application, showing the optimization problem and constraints for maximizing network capacity.
[0132] In one possible implementation, such as Figure 4 As shown, the optimization problem and constraints for determining the maximum network capacity include:
[0133] Step 310: Given the solution to the overall base station and ground user association problem as matrix A:
[0134] ;
[0135] Step 320: Determine the optimization problem and constraints for maximizing network capacity as follows:
[0136] ;
[0137] in, Binary parameters representing the association between the ground user and the overall base station, when the ground user is associated with the overall base station, ,otherwise ;
[0138] C1 defines the movement trajectory of the aerial base station and the UAV, assuming the UAV's trajectory is known prior and it performs rapid circular motion around the scene edge. Furthermore, this invention considers dynamic scenes and segments time slots, with scene information... Updates every second;
[0139] C2 means that the total number of users served by the base station cannot exceed the maximum number of connections of the current base station. The limitation;
[0140] C3 is a restriction on user association, ensuring that each user has one and only one base station for downlink data transmission in the same time slot.
[0141] Figure 5 This is a schematic diagram illustrating an exemplary embodiment of the present application, showing how the Q-learning reinforcement algorithm is used to solve the optimization problem and obtain the optimal user association solution.
[0142] In one possible implementation, such as Figure 5 As shown, the step of solving the optimization problem using the Q-learning reinforcement algorithm to obtain the optimal user association solution includes:
[0143] Step 410: Decompose the user association problem into tuples The tuple This refers to multiple parts of the Markov decision-making process;
[0144] Step 420: Design the Q-value function table as follows A formatted table;
[0145] Step 430: Based on tuples Using the Q-value function table, the optimal user association solution between the ground user and the overall base station is obtained.
[0146] in, Representing the intelligent agent, this strategy assigns the decision-making task to the core network. The user association problem is viewed as a many-to-one interconnection problem between each user and all base stations. User association state information for the entire network is collected by airborne and ground base stations. The decision-maker's task is to update the Q-value function table through the algorithm flow in each iteration.
[0147] This represents the user association status collected by the air-to-ground base stations. This status is determined by the overall network association matrix. This is to ensure the continuity of the associated state. Among them... This represents the upper limit of the number of states, which can be calculated. The elements of the state space form each row of the Q-value function table, and the number of state spaces is the total number of rows in the Q-value function table.
[0148] This represents the action taken by the intelligent agent; each element represents a specific user. Select a base station Access is then established. Among other things, This represents the number of motion spaces. The total number of motion spaces can be calculated as follows: The elements of the action space constitute each column of the Q-value function table, and the number of elements in the action space is the total number of columns in the Q-value function table;
[0149] This represents the immediate reward function, which is positively correlated with the difference in the overall network capacity between two consecutive iterations, and gradually decreases with time and the number of iterations.
[0150] Figure 6 This is a schematic diagram illustrating the process of obtaining the optimal user association solution through the Q-learning reinforcement algorithm in an exemplary embodiment of this application. Figure 8 This is a flowchart illustrating an exemplary embodiment of the Q-learning reinforcement algorithm of this application.
[0151] like Figure 6 and Figure 8 As shown, the overall process of obtaining the optimal user association solution through the Q-learning reinforcement algorithm includes:
[0152] Step 431: Initialize the network scenario and related parameters for Q-Learning, including the number and location of users and base stations, and the number of training iterations. Number of iterations Learning rate Discount Factor Action selection probability wait;
[0153] Step 432: Determine the iteration number. If it is the first iteration, generate the current state using the nearest association algorithm, and set... Use this as the initial state. Otherwise, use the state from the previous iteration as the current state.
[0154] Step 433: Collect data from airborne and ground base stations, and record the current status. Input to the intelligent agent In the middle, intelligent agents Based on the Q-value function table, the enhanced action selection strategy proposed in this invention is executed, and the currently changed user action selection is output. And the user's changed base station selection That is, the optimal user association solution.
[0155] In the above embodiments, the enhanced action selection strategy consists of two parts:
[0156] Randomly select the user to be changed ;
[0157] Select User The base station that will be replaced.
[0158] Base station selection strategies include random exploration, greedy algorithm for maximizing signal-to-noise ratio, and Q-value-maximizing strategy. The probability distributions of these three strategies are as follows: , , ,and The probabilities of the three strategies are determined by the following formula:
[0159]
[0160] in, Choose a probability for the action. The parameters are adjustable and depend on the scene size. This indicates the current iteration number, i.e. . This represents the upper limit of the number of iterations, i.e. .
[0161] In one possible implementation, the greedy base station selection strategy that maximizes the signal-to-noise ratio includes each user... Action The probability of selection is:
[0162]
[0163] The QQ learning-enhanced algorithm is used to evaluate the selected actions. If the change in action improves the overall channel capacity, the selected action is returned. Otherwise, the old action is returned, ensuring that the overall channel capacity continues to improve.
[0164] The difference in the overall channel capacity of the system before and after is calculated using the following formula, and is used as the reward value after performing the selected action in the current state.
[0165]
[0166] in, This represents a function for calculating the overall channel capacity of the network system. This represents the correlation matrix before the action is executed, while This represents the correlation matrix after the action strategy is executed.
[0167] If the reward is positive, round the reward value and use it as the reward, then update the Q-value table using the following formula. When the reward value is negative, let... Similarly, the Q-value function table is updated using the following formula.
[0168]
[0169] A saturation check is used to determine if the current environment is saturated. If so, the loop exits immediately. Similarly, when the maximum number of iterations is reached, the loop exits, and the algorithm terminates. Otherwise, the current state is saved. Returning to the state selection phase of the algorithm, and using the saved state... Continue the loop iteration until the end and exit to obtain the final Q-value function table.
[0170] The rows of the Q-value function table represent the state information at a certain moment, which is the correlation matrix. The columns of the Q-value function table represent the users whose current state will change. The selectable action base stations, the Q value is the user's The expected value of the selected base station action, constrained by the reward function, corresponds to the optimal base station action selection based on the maximum Q value in the Q-table. The correlation matrix is then calculated according to the maximum Q value. By iterating through and updating the Q table according to its maximum value, we can obtain the optimal user association strategy, i.e., the optimal user association solution.
[0171] Some embodiments of this application are applied to dynamic air-ground heterogeneous network scenarios. Based on the Q-learning user association method, an air-ground integrated heterogeneous network downlink transmission link model is established. The network mainly includes air base stations, ground small base stations, ground macro base stations, and mobile users. It focuses on the dual mobility of air base stations and users, with the goal of maximizing system network capacity, optimizing the association results between users and base stations, and improving the overall network performance.
[0172] Furthermore, in the context of heterogeneous air-ground networks, a user association enhancement method based on Q-learning is proposed. This invention addresses the unique characteristics of heterogeneous air-ground network scenarios and the dual mobility of ground users and airborne base stations by enhancing the Q-learning algorithm to adapt to dynamic scenarios, accelerate network capacity increases, and expedite the convergence process.
[0173] To demonstrate the effectiveness of this application, some embodiments of this application are compared with the user association algorithm proposed in the literature "Prioritized User Association for Sum-Rate Maximization in UAV-Assisted Emergency Communication: A Reinforcement Learning Approach" under the same air-to-ground heterogeneous scenario. The main difference lies in the overall channel performance. Simulation parameters are shown in Table 1. Due to its universality, this algorithm is temporarily referred to as the Q-learning universal algorithm.
[0174]
[0175] Table 1 Simulation Parameters
[0176] Figure 9 This is a schematic diagram comparing the Q-learning reinforcement algorithm and the Q-learning universal algorithm according to an exemplary embodiment of this application, as shown below. Figure 9 The figure shows a convergence comparison between the Q-learning-based user association enhancement algorithm and the Q-learning universal algorithm with 100 users. The results demonstrate the stability and adaptability of the algorithm of this invention in dynamic scenarios.
[0177] Some embodiments of this application have been optimized in three aspects to adapt to the set scenario: first, the Q-table associated with the initial user; second, the action selection strategy; and third, the reward function. The simulation experiment scenario of this invention is in the simulation area... In this network scenario, one macro base station and three small base stations are configured, along with three airborne base stations. User distribution considers both randomness and clustering characteristics. 60% of users follow a Poisson distribution, randomly generated across various network areas to simulate randomness in user distribution. The remaining 40% of users are densely distributed around the small base stations, simulating the presence of cells.
[0178] Figure 10 This is a schematic diagram comparing the system channel capacity of an exemplary embodiment of this application with that of the prior art, as shown below. Figure 10As shown, a comparison of network capacity performance for four association methods is presented. With the increase in the number of users, the Q-learning reinforcement algorithm and the Q-learning universal algorithm exhibit relatively higher network capacity, significantly exceeding the traditional nearest-neighbor association and maximum signal-to-noise ratio (SNR) association algorithms, reaching near saturation when the number of users reaches 100. The proposed Q-learning reinforcement algorithm achieves a channel capacity of 1361.1 Mbps, the Q-learning universal algorithm reaches 1345 Mbps, while the maximum SNR connection algorithm only achieves 892 Mbps. The Q-learning reinforcement algorithm improves channel capacity by 68% compared to the maximum SNR association algorithm, by 72% compared to the nearest-neighbor strategy, and by 1.5% compared to the Q-learning universal algorithm.
[0179] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially as indicated, these steps are not necessarily executed in the indicated order. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0180] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0181] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A dynamic air-ground heterogeneous network user association enhancement method based on reinforcement learning, characterized in that, The method comprises the following steps: constructing an air-ground heterogeneous network and initializing parameters of the network, the air-ground heterogeneous network comprising ground macro base stations, ground small base stations, air base stations and ground users, comprising: Given the set of aerial base stations , the set of ground base stations , the set of overall base stations , the set of ground users ; wherein, is a number of the aerial base stations, is a number of the ground base stations, is a number of the ground users, the ground base stations including the ground macro base stations and the ground small base stations, the total base stations including the ground base stations and the aerial base stations; modeling system channels and channel capacities based on the air-ground heterogeneous network; determining an optimization problem of maximizing network capacity and a constraint condition, comprising: given a solution of the total base station and ground user association problem as matrix A: ; wherein, denotes a binary parameter associated with the ground user and the overall base station, when the ground user is associated with the overall base station, , otherwise ; determining the optimization problem of maximizing network capacity as: ; wherein is the position coordinates of the aerial base station at the moment, is the coverage radius of the ground macro base station, C1 is a limitation on the movement trajectory of the aerial base station, C2 is a limitation that the number of users served by the overall base station cannot exceed the maximum number of connections of the current base station , C3 is a limitation on the association of users, ensuring that each user has and only has one base station for downlink data transmission in the same time slot. solving the optimization problem by a Q-learning reinforcement algorithm to obtain an optimal user association solution, comprising: Decomposing the user association problem into tuples , the tuples are a plurality of parts in a Markov decision process, wherein, denotes an agent, denotes a user association state collected by a base station, denotes an action taken by the agent, denotes an immediate reward function; The Q-value function table is designed as a table in the format, where the number of table rows is , representing the number of state spaces of the overall system, is the number of table columns, representing a certain state of the overall system, the number of action choices, , and represents the Q-value of the action being selected under the state . Based on the tuple and a Q-value function table, an optimal user association solution for the ground users and the overall base stations is obtained.
2. The method of claim 1, wherein the method further comprises: in the step of constructing an air-ground heterogeneous network and initializing parameters of the network, the air-ground heterogeneous network comprising ground macro base stations, ground small base stations, air base stations and ground users, comprising: Given that the beam width of the airborne base station is , the coverage radius of the airborne base station to the ground is ; wherein, is the flight altitude of the drone; The trajectory of the airborne base station is a circle centered on the ground macro base station with a radius of [missing information]. Circular motion, for The location coordinates of the aforementioned airborne base station at that time. For the air base station in The location at a given time, and the speed of the airborne base station. ; wherein, is a coverage radius of a ground macro base station, is a moving speed of a single said airborne base station. 3.The method of claim 1, wherein, in the step of constructing an air-ground heterogeneous network and initializing parameters of the network, the air-ground heterogeneous network comprising ground macro base stations, ground small base stations, air base stations and ground users, further comprising: The moving speed of each ground user is given as , and the moving angle is , the moving speed and the moving angle of the user are updated in real time, and the parameter is set as the update time slot of the air-ground heterogeneous network, and the speed and angle of the ground macro base station, the ground small base station, the air base station and the ground user are updated every second, and the network topology is regenerated.
4. The method of claim 1, wherein, in the step of modeling system channels and channel capacities based on the air-ground heterogeneous network, comprising: Given the position of the ground user at time is and the position of the ground base station at time is , the distance between the ground user and the ground base station is ; wherein is the height of the ground base station; Due to the cross-domain and heterogeneous nature of networks, the channel gain of terrestrial communication needs to consider both large-scale and small-scale fading. At any given time, the independent channel gain between the ground user and the ground base station is ,in, This is a small-scale fading component; According to the Jakes model, the small-scale fading components are represented as first-order complex Gaussian Markov processes: ; wherein , is a first kind zero order Bessel function, are independent and identically distributed circularly symmetric complex Gaussian random variables with unit variance, is an update time slot for a heterogeneous network of airspaces; The large-scale fading component is inversely related to a distance between the ground base station and the ground user. , , the large-scale fading component is inversely related to a distance between the ground base station and the ground user. wherein is the channel power gain per unit distance of reference, denotes the speed of light, denotes the carrier frequency, denotes the path loss coefficient of the terrestrial user to terrestrial base station channel link.
5. The method of claim 4, wherein, in the step of modeling system channels and channel capacities based on the air-ground heterogeneous network, further comprising: To maintain consistency, the service channel of the airborne base station also considers both large-scale and small-scale fading. Given that the airborne base station... The position of the moment is The distance between the ground user and the airborne base station is ; wherein, H is the height of the aerial base station; The channel model definition is analogous to that of the ground base station, in At time t, the channel gain between the ground user and the air base station is ; wherein is the path fading coefficient for the terrestrial user to air base station channel link, satisfying the relationship .
6. The method of claim 5, wherein, in the step of modeling system channels and channel capacities based on the air-ground heterogeneous network, further comprising: The signal-to-noise ratio between the ground user and the overall base station is ; wherein is the transmit power of the overall base station in the time slot , denotes the interference power received by the overall base station at the ground user in the time slot , and ; wherein, denotes the received power from the overall base stations, denotes the additive white Gaussian noise power, the number of ground users accessing the overall base stations in the air-ground heterogeneous network is denoted as .
7. The method of claim 6, wherein, in the step of modeling system channels and channel capacities based on the air-ground heterogeneous network, comprising: The transmission rate between the overall base station and the ground user is ; wherein, represents the spectrum resources allocated to the base stations in the group. All users under the same base station share bandwidth resources, and the total channel capacity of the system is .
Citation Information
Patent Citations
Unmanned aerial vehicle heterogeneous network energy efficiency optimization method based on deep reinforcement learning
CN114189891A
A Resource Allocation Method for Cellular Heterogeneous Networks Based on Deep Reinforcement Learning
CN114938543A