DQN parameter optimization system and method based on Voronoi boundary
Through the DQN parameter optimization system based on Voronoi boundaries, the switching process is simplified and the switching boundaries are optimized, and the problem of frequent switching in super-dense networks is solved, thereby improving user data rate and system efficiency.
Patent Information
- Application Number
- CN202510633676.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-16
AI Technical Summary
In ultra-dense network environments, traditional switching mechanisms based on signal strength lead to frequent switching, reducing user service quality, and traditional methods require complex signaling interactions, increasing switching costs.
The DQN parameter optimization system based on Voronoi boundary is adopted. Through the path loss module, user mobile module, user switching module and reinforcement learning module, combined with the self-feedback mechanism, the switching boundary and switching process are optimized, the user switching rate is reduced, and the data rate is improved.
Simplify the switching process, reduce unnecessary information interaction, adaptively optimize the switching boundaries, reduce user switching rates, improve user data rates, and accelerate network convergence.
Smart Images

Figure CN120498574A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wireless communication switching applications based on reinforcement learning, and in particular to a DQN parameter optimization system and method based on Voronoi boundaries. Background Art
[0002] With the widespread adoption of millimeter-wave networks, ultra-dense networks are experiencing rapid development. In ultra-dense network environments, due to the high density of base station deployments, mobile users frequently switch base stations while traveling, making handover management a key challenge to be solved. Regarding the currently widely used signal strength-based handover mechanism, practical systems typically utilize the A3 handover trigger event defined in 3GPP standardization to execute the handover process. In ultra-dense network environments, this approach results in frequent handovers, which can degrade user service quality, such as reduced user data rates.
[0003] To this end, numerous studies have leveraged machine learning techniques based on this handover mechanism to optimize handover parameters, reducing handover frequency and improving network performance. However, traditional signal strength-based handover methods require frequent and complex signaling interactions between the base station and the user, involving reporting critical information such as network channel status, the user's received power in the serving cell and neighboring cells, and the level of interference caused by surrounding devices. The traditional signal strength-based handover process involves complex signaling interactions, which not only reduces system efficiency but also increases handover costs. In recent years, the introduction of neural networks has reduced the handover rate to a certain extent, but existing research has not yet improved traditional handover schemes.
[0004] In view of this, the present invention proposes a DQN parameter optimization system and method based on Voronoi boundaries. Summary of the Invention
[0005] The purpose of the present invention is to provide a DQN parameter optimization system and method based on Voronoi boundaries to solve the problems mentioned in the background technology.
[0006] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:
[0007] A DQN parameter optimization system based on Voronoi boundaries, including the following modules:
[0008] Path loss module: used to complete channel modeling and describe the power attenuation law of the signal during propagation;
[0009] User Mobility Module: This module uses a random walk model, assuming that the user's current and future locations are random, and that the user's speed and direction of movement are also random and unrelated to past and future moments.
[0010] User handover module: Based on the Voronoi model, the optimal handover boundary problem is transformed into a handover region problem. The handover region is defined as the area where users experience high handover frequency during the handover process. The handover process is divided into the handover preparation phase, the candidate base station selection phase, and the handover execution phase.
[0011] Reinforcement Learning Module: Combines the reinforcement learning algorithm with the Voronoi model to reduce user switching rates and increase user data rates. It also introduces a self-feedback mechanism to correct the action output of the DQN network and accelerate network convergence.
[0012] A DQN parameter optimization method based on Voronoi boundaries includes the following steps:
[0013] S1. Build a path loss model, complete channel modeling, and describe the power attenuation law of the signal during propagation;
[0014] S2. Select the random walk model as the user movement model, assuming that the user's current and future locations are random, and the user's speed and movement direction are random and have nothing to do with the past and future moments;
[0015] S3. Construct a user handoff model based on Voronoi boundaries, transform the problem of solving the optimal handoff boundary into solving the handoff region. Define the area where users experience high-frequency handoffs during the handoff process as the handoff region, and divide the handoff process into the handoff preparation phase, the candidate base station selection phase, and the handoff execution phase.
[0016] S4. Combine the reinforcement learning algorithm with the Voronoi model to build a reinforcement learning network DQN switching framework to reduce the user switching rate and increase the user data rate; and introduce a self-feedback mechanism to correct the action output of the DQN network and accelerate network convergence.
[0017] Preferably, the division of the handover process in S3 specifically includes the following contents:
[0018] S3.1, Handover preparation stage:
[0019] During the user's random walk, the system continuously obtains the user's location information. When the user approaches the handover boundary, the system enters the handover preparation phase to prepare for the subsequent handover.
[0020] The optimal switching boundary is represented by the exponentially decaying path loss model, which is expressed as follows:
[0021]
[0022] Among them, P T represents the base station transmission power; u represents the initial position of the user; x sIndicates the location of the serving base station; x i represents the location of other base stations; η represents the path loss factor; Φ represents the uniform random point process;
[0023] Simplifying the above formula, the optimal switching boundary is:
[0024] Γ(x s ,x i )={(x,y)|(xx s ) 2 +(yy s ) 2 =(xx i ) 2 +(yy i ) 2}
[0025] When the user moves the position and Γ(x s ,x i ) When the distance d of the trajectory is less than or equal to a certain critical value D, it enters the switching preparation stage, where D is the switching distance threshold;
[0026] S3.2, candidate base station selection stage:
[0027] The time constraint TTT in the A3 handover event is used to enhance the reliability of the handover process. Based on the A3 handover trigger event, the concept of handover region is proposed. The range enclosed by the handover threshold D when the user enters the handover preparation stage and the handover boundary is called the handover region. The formula is expressed as:
[0028]
[0029] Among them, (a) changes to the Cartesian coordinate system representation; (b) changes to the polar coordinate system representation; r represents the user as the pole and the boundary Γ(x s ,x i ); the polar axis is a line parallel to the switching boundary; θ represents the angle between the polar axis and the target motion direction, ranging from [0,π]; r s and θ s Represents the distance and angle between the user and the serving base station; r i and θ i Indicates the distance and angle between the user and the target base station;
[0030] During user mobility, the system calculates the handover area between the user and multiple base stations, and screens potential target base stations based on the TTT condition. It uses the distance vector sum algorithm to calculate the vector distance between the user and each candidate base station within the specified TTT time, and accumulates and sums these distances. The formula is as follows:
[0031]
[0032] Among them, d i (t) represents the distance d between the user and base station i at time t, i represents the base station number; sign (α) v' represents the vertical component of the user's movement toward the target base station, v' = v·sinθ, L represents the cumulative sum of the vector distances to the base station within the time window. If the cumulative sum of the vector distances L between the user and a base station is always positive, the base station is considered a candidate base station for handover. If the cumulative sum of the vector distances L between the user and a base station is negative at a certain moment, the base station is disqualified as a candidate base station.
[0033] S3.3, switching execution phase:
[0034] The system selects the base station with the maximum cumulative sum of vector distances from the candidate base stations as the best target base station. The formula is:
[0035]
[0036] Wherein, I represents the best target base station.
[0037] Preferably, the reinforcement learning network DQN switching framework includes:
[0038] Communication environment module: responsible for recording and processing key information of users during movement, including: user's instantaneous movement speed, instantaneous movement direction, currently connected serving base station, distance between user and serving base station and distances to other base stations, transmission power of each base station, channel propagation loss, and real-time data rate received by the user;
[0039] The switching module is responsible for receiving user and base station information sent by the communication environment module, processing the information, and selecting the best target base station; it is also responsible for calculating the state information learned by the agent module;
[0040] Agent module: responsible for receiving status information and immediate rewards, feeding them into the decision module for decision-making; also responsible for receiving the latest action output returned by the decision module and outputting the action to the communication environment module;
[0041] Decision module: Responsible for integrating and calculating the state information and switching decision information in the communication environment module, switching module and agent module, storing the current state information and its corresponding action value, and making decisions based on the immediate reward function.
[0042] Preferably, a self-feedback mechanism is introduced into the decision module to modify the action D and the time limit condition TTT, specifically including the following contents:
[0043] A simplified A3 handover event mechanism is proposed. A hysteresis delay factor (HOM) is introduced into the standardized A3 handover trigger event mechanism. When the reference signal received power (RSRP) of the target base station is greater than or equal to the RSRP of the serving base station and a certain HOM margin is met, the system considers the handover successful. The formula is expressed as follows:
[0044] RSRP T RSRP S +HOM
[0045] The calculation simplifies the A3 handover event and outputs the lowest HOM parameter value that meets the conditions. If the output value is positive, it means that the current handover meets the conditions and there is no need to adjust the output action. If the output value is negative, it indicates that the RSRP of the target base station is lower than the RSRP of the serving base station, indicating that the handover is not ideal. The system immediately adjusts the D and TTT values based on the correction function and feeds back the corresponding actions to the intelligent agent module to optimize the handover decision and improve the handover success rate. The formula of the correction function is expressed as:
[0046]
[0047] Among them, α represents the correction factor, which is a constant used to control the weighting degree; W represents the weight factor output by the correction function; the self-feedback mechanism corrects the action D and TTT according to W, that is, way, where W∈[0,1].
[0048] The present invention further protects a computer device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the instruction, program, code set, or instruction set is loaded and executed by the processor to implement the above-mentioned DQN parameter optimization method based on Voronoi boundaries.
[0049] The present invention further protects a computer-readable storage medium, characterized in that the computer-readable storage medium stores at least one instruction, at least one program, code set or instruction set, and the instruction, program, code set or instruction set is loaded and executed by a processor to implement the above-mentioned DQN parameter optimization method based on Voronoi boundaries.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] The present invention continuously obtains user location information to determine whether to execute a handoff, thereby reducing unnecessary information exchange and simplifying the handoff process. Furthermore, to further alleviate the problem of frequent user handoffs, the present invention combines the Voronoi model with deep reinforcement learning (DQN) to construct a new reinforcement learning-driven handoff framework. This framework can adaptively optimize the Voronoi handoff boundary to better adapt to complex communication environments, thereby effectively reducing user handoff rates and increasing user data rates. Furthermore, the present invention introduces a self-feedback mechanism into the reinforcement learning-driven handoff framework, which can better correct the action output of the DQN network and accelerate network convergence. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, a brief introduction to the drawings involved in the embodiments is now provided. It is obvious that the drawings described below are only schematic illustrations of some embodiments of the present invention. Those skilled in the art can construct other forms of drawings based on these drawings without inventive effort.
[0053] Figure 1 This is a schematic diagram of user movement mentioned in Example 1 of the present invention;
[0054] Figure 2 This is a schematic diagram of the switching process mentioned in Example 1 of the present invention;
[0055] Figure 3 This is a schematic diagram of the DQN framework structure mentioned in Example 1 of the present invention;
[0056] Figure 4-6 This is a diagram showing the network performance characterization results mentioned in Example 2 of the present invention. DETAILED DESCRIPTION
[0057] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments.
[0058] Example 1:
[0059] The present invention proposes a DQN parameter optimization system and method based on Voronoi boundaries, which specifically include:
[0060] 1. Path loss model
[0061] This example assumes a single-layer ultra-dense network where the distribution of base stations follows a uniform random point process Φ with strength u. The received power P of the mobile device is R As shown below:
[0062] PR (d) = P T ·|d| -η ·l(|d|)
[0063] Among them, the symbol P T represents the base station transmission power, d represents the distance between the signal transmitter and the receiver. The symbol η represents the path loss factor, and η is set to ≥ 4. l represents the shadow fading model, and lnl(|d|) obeys the mean The variance is σ 2 The shadow of fading.
[0064] 2. User Mobility Model
[0065] This example sets the user mobility model to a random walk model, which assumes that the user's current and future locations are random, that is, the user's speed and movement direction are random and have nothing to do with past and future moments.
[0066] Assuming the user's initial position is u(x,y), at each time t, the user's position is updated by a combination of speed and direction:
[0067] u(t+1)=u(t)+v·Δt
[0068] Among them, u(t) represents the current user location, u(t+1) represents the user location at the next moment, v represents the user's velocity, and Δt is the time step.
[0069] 3. User switching model based on Voronoi boundary
[0070] The Voronoi model is widely used as an analytical method for spatial division. It divides the space into several regions, and the distance between the points in each region and the corresponding base point of the region is closer than the distance to other base points. The region generated by each base point is called the Voronoi cell of the base point, and two adjacent Voronoi cells are connected by sharing a common boundary. The irregular Voronoi cell is simulated as the signal coverage area of the base station, and the boundary of the signal coverage area of the adjacent base station is regarded as the user's switching boundary. The base station in the user's area is called the serving base station B. S The base station to which the user switches is called the target base station B. T .
[0071] Based on the Voronoi model, the handover boundary is represented by the area where the power is equal. That is, on this boundary, the signal power received by the user from the two base stations is equal:
[0072]
[0073] Among them, the location of the serving base station is recorded as x s (xs ,y s ), the other base station locations are denoted as x i (x i ,y i ), where x i ∈Φ / x s .
[0074] The service base station to which the user initially connects satisfies the formula:
[0075]
[0076] Because actual handover boundaries fluctuate over time, accurately determining a user's optimal handover boundary becomes extremely complex. To address this challenge, this paper proposes transforming the problem into determining the handover region. Specifically, the handover region is defined as an area where a user may experience a high frequency of handovers during a handover, i.e., an area with high uncertainty in user signal coverage. This paper divides the handover process into a handover preparation phase, a candidate base station selection phase, and a handover execution phase, specifically encompassing the following aspects.
[0077] (1) Switching preparation phase
[0078] During the user's random walk, the system continuously obtains the user's location information. When the user approaches the handoff boundary, the system enters the handoff preparation phase to prepare for the subsequent handoff.
[0079] First, the optimal switching boundary is a path loss model that decays with an exponential function, which is expressed as:
[0080]
[0081] Simplifying the above formula, the optimal switching boundary is:
[0082] Γ(x s ,x i )={(x,y)|(xx s ) 2 +(yy s ) 2 =(xx i ) 2 +(yy i ) 2}
[0083] When the user moves the position and Γ(x s ,x i ) When the distance d of the trajectory is less than or equal to a certain critical value D, it enters the switching preparation stage, where D is called the switching distance threshold.
[0084] (2) Candidate base station selection stage
[0085] Due to the high degree of randomness in a user's random walk speed and direction, this uncertainty makes relying solely on the handover preparation phase insufficient to ensure a smooth handover. Specifically, factors such as signal fluctuations and interference may prevent users from accurately predicting handover timing, leading to handover delays or failures. Therefore, simply relying on users approaching the handover boundary and entering the handover preparation state is insufficient to guarantee handover stability and success rates.
[0086] In order to improve the stability and reliability of the handover process, the present invention adopts the time limit condition TTT in the A3 handover event. The A3 handover event refers to the time threshold condition set when handover occurs between base stations, which limits the signal quality change of the user within a certain period of time and prevents frequent handovers caused by short-term signal fluctuations. Specifically, the TTT condition stipulates that when the received signal quality of the user is continuously higher than the set threshold M S +Hys and the duration exceeds the preset TTT value before the handover is executed. This condition helps avoid invalid handovers caused by instantaneous signal fading or interference, thereby enhancing the reliability of handovers.
[0087] Based on the A3 handover trigger event, the present invention defines a candidate base station selection phase to further optimize the handover process. To better explain the candidate base station selection phase, the concept of a handover region is first proposed. The range enclosed by the handover threshold D and the handover boundary when the user enters the handover preparation phase is called the handover region, namely:
[0088]
[0089] Here (a) is changed to Cartesian coordinate system; (b) is changed to polar coordinate system; r represents the user as the pole and the boundary Γ(x s ,x i ); the polar axis is a line parallel to the switching boundary; θ represents the angle between the polar axis and the target motion direction, ranging from [0,π]; r s and θ s Represents the distance and angle between the user and the serving base station; r i and θ i Represents the distance and angle between the user and the target base station.
[0090] See also Figure 1 , Figure 1This is a schematic diagram of user movement. The blue curve represents the handoff boundary, the black triangles represent base stations, the red dots represent poles, the gray curve represents the polar axis, the black curve represents the x-axis, and the red arrow curve represents the user's movement direction. The solid line represents the system's calculation of the current candidate base station as 10, and the dashed line represents the system's calculation of the current candidate base station as 3. For example, the solid line shows that the user's current polar axis is parallel to the handoff boundary, the x-axis is perpendicular to the handoff boundary, the angle between the polar axis and the movement direction is θ, and the angle between the x-axis and the movement direction is α. r represents the distance from the user to the handoff boundary.
[0091] During this phase, the system calculates the handover zones between the user and multiple base stations and screens potential target base stations based on the time-to-distance (TTT) criteria. Specifically, the present invention uses a distance vector sum algorithm to determine candidate base stations. During this process, the system calculates the vector distances between the user and each candidate base station within a specified time-to-distance (TTT) and accumulates and sums these distances. The formula is as follows:
[0092]
[0093] Among them, d i (t) represents the distance d between the user and base station i at time t, i represents the base station number. sign(α)v' represents the vertical component of the user's movement toward the target base station, where v' = v·sinθ, L is the cumulative sum of the vector distances to the base station within the time window. If the cumulative sum of the vector distances between the user and a base station is always positive, the base station is considered a candidate for handover.
[0094] Specifically, during the TTT period, due to the user's random movement, candidate base stations are constantly updated over time, meaning the number and number of candidate base stations may change at any moment. If, at any given moment, the cumulative sum of the vector distances L between the user and a base station is negative, that base station will be disqualified as a candidate. Conversely, if the cumulative sum of the vector distances between the user and a base station is consistently positive, that base station will continue to be a candidate. This dynamic update mechanism allows the system to flexibly select the most appropriate target base station based on the user's real-time movement and signal conditions, ensuring the stability and accuracy of the handover process.
[0095] The goal of this phase is to continuously monitor and evaluate the signal quality of multiple base stations to ensure that the selected target base station is sufficiently stable and can provide reliable coverage during the actual handover, thereby reducing the probability of handover failure and improving system performance. By introducing the TTT condition and the candidate base station selection phase, the accuracy and stability of the handover process are effectively improved, unnecessary handover operations are reduced, and overall system performance is optimized.
[0096] (3) Switch execution phase
[0097] The system is responsible for selecting a unique optimal base station from the candidate base stations as the target base station to ensure that the user's data rate reaches the optimal level during the handover process. That is, the candidate base station with the maximum vector distance and the maximum value is selected as the optimal target base station, which is expressed as:
[0098]
[0099] Through this method, the system can select the most suitable target base station based on the vector distance information between the user and each candidate base station, thereby ensuring a smooth switching process, optimizing the user experience, and maximizing the stability and reliability of the data rate.
[0100] See also Figure 2 , which shows the handover process in detail. In the figure, the yellow dot represents the user's starting position, the green dot represents the user's ending position, and the red curve represents the user's movement trajectory. The blue solid line represents the handover boundary, the black dotted line indicates that the user enters the handover preparation phase, the red dotted line marks the start of the candidate base station selection phase, and the green dotted line represents the start of the handover execution phase. The symbol D represents the handover distance threshold, which is defined as the distance from the black dotted line to the blue solid line. TTT represents the handover time threshold, which is defined as the time period from the red dotted line to the green dotted line. S (D,TTT) Represents the switching area, d 10 (0) represents the distance between the user and the switching boundary at the initial moment, d 10 (TTT) represents the distance between the user and the handover boundary at the TTT time.
[0101] The user starts from the starting position and moves along the red track to the black dot, which means that the user enters the handover preparation phase. This area is represented by the yellow area in the figure. The yellow area changes with the user's step size per unit time. After the system receives the signal that the user has entered the handover preparation phase, the timer starts and continuously monitors the TTT period, recording the vector distance d between the user and the handover boundary. i (t), that is, if the angle θ is positive, the distance is positive; if the angle θ is negative, the distance is negative. This process is shown in the figure by starting from the initial time d i (0) to the end time d i The red and green areas within the timer (TTT) are represented. When the user reaches the green dot, the TTT timer has expired, and the system enters the handover execution phase. During this phase, the system selects the base station with the largest sum of the distance vectors between the user and each candidate base station as the target base station and completes the handover with the user via signaling messages. At this point, the user disconnects from the serving base station and establishes a new connection with the target base station, completing the handover process.
[0102] In summary, this invention addresses two key handover parameters in the A3 event and reconstructs the handover process based on the Voronoi model. This foundation also proposes a novel handover model based on user location. Unlike the traditional A3 event, which relies on signal reception strength, this model determines the handover entry conditions using the handover control parameter D and delineates the handover area based on the TTT time delay. Furthermore, this solution further improves the candidate base station screening mechanism and optimizes the optimal base station selection strategy, thereby enhancing the accuracy of handover decisions and network performance.
[0103] 4. Switching framework model based on reinforcement learning
[0104] RL (reinforcement learning) is a machine learning method based on the interaction between an agent and its environment. By obtaining feedback (rewards or penalties) during trial and error, the agent gradually optimizes its decision-making strategy, thereby maximizing long-term rewards. RL is particularly suitable for complex optimization problems, especially the switching parameter optimization problem based on the Voronoi diagram proposed in this paper. In a UDN environment, users frequently switch to maintain the stability of the data rate, and the switching process is affected by multiple factors, including the user's speed, movement direction, received power, and data rate. These factors increase the complexity of the switching decision to a certain extent. Therefore, how to optimize the switching strategy in a dynamic environment has become a challenge that needs to be solved urgently.
[0105] In order to effectively address this challenge, the present invention proposes to combine RL with the Voronoi model, and use the adaptive ability of RL to optimize key control parameters in the switching process, such as the switching time threshold TTT and the switching distance threshold D. The Voronoi model divides the environment into multiple non-overlapping areas, which not only simplifies the decision space but also optimizes the efficiency of switching decision execution. The combination of RL and the Voronoi model can improve the optimization effect of the switching process, effectively reduce the complexity of the system, and provide an effective strategy for dealing with dynamic changes in the UDN environment. By optimizing the switching parameters in real time, the model proposed in the present invention can significantly improve the accuracy of switching decisions and the overall performance of the system. The specific contents are as follows.
[0106] Reinforcement Learning Network DQN Switching Framework
[0107] This paper proposes a deep Q network (DQN) switching framework that combines RL with switching. Through the adaptive capability of RL, the framework continuously explores and learns the optimal strategy to deal with complex switching optimization problems.
[0108] The DQN switching framework is divided into four main parts according to its functions, such as Figure 3Through this structure, the agent can make accurate decisions based on the environment state information during the learning process, thereby achieving real-time optimization of the switching process, including the following:
[0109] (1) Communication environment module
[0110] In the communication network environment, the communication environment module plays a vital role. It is a platform that integrates a large amount of complex network information and is responsible for recording and processing key information about users during their movement. The main functions of this module include real-time tracking of the user's instantaneous movement speed, instantaneous movement direction, currently connected service base station, the distance between the user and the service base station, and the distance to other base stations. At the same time, the communication environment module is also responsible for monitoring and recording the transmission power of each base station, channel propagation loss, and the real-time data rate received by the user. Specifically, the communication environment module has the following four core functions.
[0111] 1) Provide user and base station information to the switching module, namely user speed, user direction, user location and all base station locations, i.e. [v, Dir, (x, y) user ,(x i ,y i ) bs , i∈Φ]
[0112] 2) The communication environment module is responsible for receiving the target base station information sent by the switching module, and sending the key information of the target base station to the decision module for switching decision. The switching decision information is the target base station RSRP, the serving base station RSRP, that is, [RSRP S , RSRP T ].
[0113] 3) Receive actions sent by the intelligent agent module, update the status information according to the latest action, and transmit the new user and base station information to the switching module.
[0114] 4) Feedback the immediate reward generated by the latest action to the agent module. The definition of the reward function R is given below. The reward function is a cumulative reward mechanism with a switching rate λ HO and data rates For performance indicators:
[0115]
[0116] Among them, λ HO represents the average user switching rate, in times / second, λ HO =HO count / time,HO countRepresents the number of user switches, in times; time represents the time the user has experienced, in seconds.
[0117] Handover control parameters directly impact the user's handover rate and average data rate. Larger D and TTT values lead to frequent handovers, impacting the user's service experience. Smaller D and TTT values can lead to handover failures or even disconnections, which also impacts the user's handover success rate. The reward function incorporates both the user's average data rate and handover rate. This is primarily due to the need to increase the user's average data rate while ensuring a lower average handover rate. Furthermore, a balance is established between the user's average data rate and handover rate to prevent network non-convergence and a handover rate approaching zero.
[0118] (2) Switching module
[0119] The handover module selects the best handover target base station based on the above handover modeling. The handover module has the following functions:
[0120] 1) The handover module is responsible for receiving user and base station information sent by the communication environment, processing the information, and selecting the best handover base station.
[0121] 2) The switching module is responsible for calculating the state information learned by the agent. The state information selected in this paper is given below: Where v represents the user's speed, reflecting the instantaneous speed of the user's movement in the network, which is an important factor in the handover decision. Dir represents the user's movement direction, which affects the user's future movement trajectory, thereby affecting the handover decision and the selection of the target base station. L represents the number of candidate base stations that the user may access during the handover process, which provides the relative position information of the candidate base stations for the handover decision. I and BS T Indicates the vector distance and target base station number of the target base station, which is used to indicate the target base station to which the user switches. BS S Indicates the currently connected service base station, which is used to represent the user's current service status. The switching module sends the above status information F to the agent.
[0122] (3) Agent Module
[0123] The agent module is responsible for receiving state information and immediate rewards, and outputting the latest actions to the environment to achieve a process of continuous exploration and learning. The functions of the agent module are as follows:
[0124] 1) Receive the status information sent by the switching module and send the current status to the decision module for decision making.
[0125] 2) Receive the latest action output returned by the decision module and apply the action output to the communication environment.
[0126] 3) Receive the immediate reward generated by the communication environment and send the immediate reward feedback to the decision module.
[0127] (4) Decision-making module
[0128] The decision module plays a crucial role in the entire system. It integrates state information and switching decision information from the switching module, the agent, and the environment information module, and performs calculations. It is responsible for storing the current state information and its corresponding action value, and making decisions based on the immediate reward function. The core functions of this module include the following aspects:
[0129] 1) To further improve the accuracy and robustness of handover decisions, this paper enhances the decision module and introduces a self-feedback function. This function corrects the action D and the time-to-touch (TTT) function. Specifically, a simplified A3 event mechanism from the standardization is introduced, which adds a hysteresis delay factor (HOM). The decision module considers the impact of user received power on handover. When the reference signal received power (RSRP) of the target base station is greater than or equal to the RSRP of the serving base station and meets a certain HOM margin, the system considers the handover successful, namely:
[0130] RSRP T RSRP S +HOM
[0131] Its main purpose is to ensure the user's handover success rate by providing dual protection from the perspective of received power. The system first calculates and simplifies the A3 handover event and outputs the lowest HOM parameter value that meets the conditions. If the output value is positive, it means that the current handover meets the conditions and there is no need to adjust the output action; if the output value is negative, it means that the RSRP of the target base station is lower than the RSRP of the serving base station, which means that the handover is not ideal. At this time, the system will immediately adjust the D and TTT values and feedback the corresponding actions to the intelligent agent to optimize the handover decision and improve the handover success rate. The correction function is given below:
[0132]
[0133] Among them, α is the correction factor used to control the weighting degree, which is a constant. Self-feedback outputs the weight factor W according to the correction function to correct the action D and TTT, that is, method, where W∈[0,1]. By modifying D and TTT actions, the handover decision process can be adjusted, making the system more stable in complex network environments. Secondly, the application of the correction mechanism effectively avoids handover failures caused by network instability or signal fluctuations, and accelerates the convergence of the network loss function.
[0134] 2) The decision module outputs an action based on the received state information using a greedy strategy, namely the ε-greedy strategy: it randomly selects an action (exploration) with probability ε and selects the current optimal action with probability 1-ε. The output action is the combination of the handover distance threshold D and the handover time threshold TTT (D, TTT).
[0135] In the DQN handover framework, the agent continuously learns state information F, exploring and adjusting its decision-making strategy to determine the optimal handover distance threshold D and handover time threshold TTT that best suit the current communication environment. Furthermore, the framework incorporates a self-feedback mechanism to effectively ensure the success rate of user handovers. This mechanism not only plays a key role in the handover decision-making process but also improves the stability and accuracy of the handover process by dynamically adjusting the strategy. This optimization process reduces handover failures caused by network fluctuations or signal strength variations, improving overall system performance and user experience, accelerating network convergence, and optimizing the reward function, further enhancing system performance.
[0136] Example 2:
[0137] Based on Example 1, but different in that the contents mentioned in Example 1 are simulated using the Python platform, and the specific simulation pseudo code is as follows:
[0138]
[0139]
[0140]
[0141] A comparative experiment was designed to compare the DQN-based switching scheme proposed in this invention with the A3 event-based switching scheme and the fixed switching control parameter scheme. The performance characteristics are as follows: Figure 4-6 As shown in the figure, based on the characterization data, it can be seen that the present invention uses the DQN network to reduce the switching rate and increase the user data rate. In addition, the self-feedback mechanism of the DQN network improves the network loss convergence speed and optimizes the intelligent agent action output.
[0142] It should be noted that, in the present invention patent, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0143] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A DQN parameter optimization system based on Voronoi boundaries, characterized in that: Includes the following modules: Path loss module: used to complete channel modeling and describe the power attenuation law of the signal during propagation; User Mobility Module: This module uses a random walk model, assuming that the user's current and future locations are random, and that the user's speed and direction of movement are also random and unrelated to past and future moments. User handover module: Based on the Voronoi model, the optimal handover boundary problem is transformed into a handover region problem. The handover region is defined as the area where users experience high handover frequency during the handover process. The handover process is divided into the handover preparation phase, the candidate base station selection phase, and the handover execution phase. Reinforcement Learning Module: Combines the reinforcement learning algorithm with the Voronoi model to reduce user switching rates and increase user data rates. It also introduces a self-feedback mechanism to correct the action output of the DQN network and accelerate network convergence.
2. A DQN parameter optimization method based on Voronoi boundaries implemented by the system of claim 1, characterized in that: The following steps are involved: S1. Build a path loss model, complete channel modeling, and describe the power attenuation law of the signal during propagation; S2. Select the random walk model as the user movement model, assuming that the user's current and future locations are random, and the user's speed and movement direction are random and have nothing to do with the past and future moments; S3. Construct a user handoff model based on Voronoi boundaries, transform the problem of solving the optimal handoff boundary into solving the handoff region. Define the area where users experience high-frequency handoffs during the handoff process as the handoff region, and divide the handoff process into the handoff preparation phase, the candidate base station selection phase, and the handoff execution phase. S4. Combine the reinforcement learning algorithm with the Voronoi model to build a reinforcement learning network DQN switching framework to reduce the user switching rate and increase the user data rate; A self-feedback mechanism is introduced to correct the action output of the DQN network and accelerate network convergence.
3. The method for optimizing DQN parameters based on Voronoi boundaries according to claim 2, characterized in that: The handover process described in S3 is divided into the following specific contents: S3.1, Handover preparation stage: During the user's random walk, the system continuously obtains the user's location information. When the user approaches the handover boundary, the system enters the handover preparation phase to prepare for the subsequent handover. The optimal switching boundary is represented by the exponentially decaying path loss model, which is expressed as follows: Among them, P T represents the base station transmission power; u represents the initial position of the user; x s Indicates the location of the serving base station; x i represents the location of other base stations; η represents the path loss factor; Φ represents the uniform random point process; Simplifying the above formula, the optimal switching boundary is: Γ(x s ,x i )={(x,y)|(xx s ) 2 +(yy s ) 2 =(xx i ) 2 +(yy i ) 2 } When the user moves the position and Γ(x s ,x i ) When the distance d of the trajectory is less than or equal to a certain critical value D, it enters the switching preparation stage, where D is the switching distance threshold; S3.2, candidate base station selection stage: The time constraint TTT in the A3 handover event is used to enhance the reliability of the handover process. Based on the A3 handover trigger event, the concept of handover region is proposed. The range enclosed by the handover threshold D when the user enters the handover preparation stage and the handover boundary is called the handover region. The formula is expressed as: Among them, (a) changes to the Cartesian coordinate system representation; (b) changes to the polar coordinate system representation; r represents the user as the pole and the boundary Γ(x s ,x i ); the polar axis is a line parallel to the switching boundary; θ represents the angle between the polar axis and the target motion direction, ranging from [0,π]; r s and θ s Represents the distance and angle between the user and the serving base station; r i and θ i Indicates the distance and angle between the user and the target base station; During user mobility, the system calculates the handover area between the user and multiple base stations, and screens potential target base stations based on the TTT condition. It uses the distance vector sum algorithm to calculate the vector distance between the user and each candidate base station within the specified TTT time, and accumulates and sums these distances. The formula is as follows: Among them, d i (t) represents the distance d between the user and base station i at time t, i represents the base station number; sign (α) v' represents the vertical component of the user's movement toward the target base station, v' = v·sinθ, L represents the cumulative sum of the vector distances to the base station within the time window. If the cumulative sum of the vector distances L between the user and a base station is always positive, the base station is considered a candidate base station for handover. If the cumulative sum of the vector distances L between the user and a base station is negative at a certain moment, the base station is disqualified as a candidate base station. S3.3, switching execution phase: The system selects the base station with the maximum cumulative sum of vector distances from the candidate base stations as the best target base station. The formula is: Wherein, I represents the best target base station.
4. The method for optimizing DQN parameters based on Voronoi boundaries according to claim 3, characterized in that: The reinforcement learning network DQN switching framework includes: Communication environment module: responsible for recording and processing key information of users during movement, including: user's instantaneous movement speed, instantaneous movement direction, currently connected serving base station, distance between user and serving base station and distances to other base stations, transmission power of each base station, channel propagation loss, and real-time data rate received by the user; The switching module is responsible for receiving user and base station information sent by the communication environment module, processing the information, and selecting the best target base station; it is also responsible for calculating the state information learned by the agent module; Agent module: responsible for receiving status information and immediate rewards, feeding them into the decision module for decision-making; also responsible for receiving the latest action output returned by the decision module and outputting the action to the communication environment module; Decision module: Responsible for integrating and calculating the state information and switching decision information in the communication environment module, switching module and agent module, storing the current state information and its corresponding action value, and making decisions based on the immediate reward function.
5. The method for optimizing DQN parameters based on Voronoi boundaries according to claim 4, characterized in that: The decision module introduces a self-feedback mechanism for correcting the action D and the time constraint TTT, which specifically includes the following: A simplified A3 handover event mechanism is proposed. A hysteresis delay factor (HOM) is introduced into the standardized A3 handover trigger event mechanism. When the reference signal received power (RSRP) of the target base station is greater than or equal to the RSRP of the serving base station and a certain HOM margin is met, the system considers the handover successful. The formula is expressed as follows: RSRP T >RSRP S +HOM Calculate and simplify the A3 switching event and output the lowest HOM parameter value that meets the conditions. If the output value is positive, it means that the current switching condition is met and there is no need to adjust the output action. If the output value is negative, it indicates that the RSRP of the target base station is lower than the RSRP of the serving base station, indicating that the handover is not ideal. The system immediately adjusts the D and TTT values based on the correction function and feeds the corresponding action back to the intelligent agent module to optimize the handover decision and improve the handover success rate. The formula of the correction function is expressed as: Among them, α represents the correction factor, which is a constant used to control the weighting degree; W represents the weight factor output by the correction function; the self-feedback mechanism corrects the action D and TTT according to W, that is, way, where W∈[0,1].
6. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the instruction, program, code set, or instruction set is loaded and executed by the processor to implement the DQN parameter optimization method based on Voronoi boundaries according to any one of claims 2 to 5.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, and the instruction, program, code set, or instruction set is loaded and executed by the processor to implement the DQN parameter optimization method based on Voronoi boundaries as described in any one of claims 2 to 5.
Citation Information
Patent Citations
Dynamic spectrum access method based on layered deep reinforcement learning in NOMA system
CN113207127A
DQN-assisted low earth orbit satellite switching decision generation method
CN116916409A
Internet of vehicles channel congestion control method based on multi-agent deep reinforcement learning
CN118283700A
Low-orbit giant constellation satellite switching method and device based on deep reinforcement learning
CN118611728A
Processing of communications signals using machine learning
US10396919B1