A comprehensive optimization method for base station sleep control strategy and beam width selection strategy

By constructing an optimization model and deep reinforcement learning algorithm, dynamically adjusting the base station sleep and beam width, the balance problem between base station energy consumption and user initial access delay in high-frequency mobile communications is solved, and the base station energy consumption is reduced without significantly affecting the user delay performance, thereby improving the calculation accuracy and reducing the calculation complexity.

CN119854828BActive Publication Date: 2025-10-03HUAZHONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411983548.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-10-03
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing technologies fail to effectively balance base station energy consumption and user initial access delay. Especially in high-frequency mobile communication networks, existing research has not fully considered the impact of base station sleep state on beam width selection.

Method used

By building an optimization model with the goal of minimizing the total energy consumption of the base station and the average beam scanning delay, combined with a deep reinforcement learning algorithm, the DQN model is used for solution, the base station sleep and beam width are dynamically adjusted, and a sector antenna model is used to approximate the actual antenna radiation pattern to improve the calculation accuracy.

Benefits of technology

Without significantly affecting user delay performance, the base station energy consumption is reduced, the base station sleep control and beam width selection are optimized, which solves the balance problem between base station energy consumption and user initial access delay, reduces calculation complexity and improves calculation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119854828B_ABST
    Figure CN119854828B_ABST
Patent Text Reader

Abstract

The present invention discloses a comprehensive optimization method for a base station sleep control strategy and a beam width selection strategy, belonging to the field of wireless communications. The method takes minimizing the total energy consumption of the base station and the average beam scanning delay as the optimization goal, establishes constraint conditions for the base station sleep control and the beam width selection, and builds an optimization model. The base station sleep decision and the beam width selection decision of the target cell are obtained by solving the method. The method fully considers the influence of the base station sleep on the beam width selection and the initial access delay, and can solve the problem of the difficulty in balancing the base station energy consumption and the user initial access delay in the traditional base station sleep control and beam width selection methods. By comprehensively optimizing the base station sleep control and beam width selection, the base station energy consumption is reduced without significantly affecting the user delay performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communications, and more specifically, relates to a comprehensive optimization method for a base station sleep control strategy and a beam width selection strategy. Background Art

[0002] Future mobile communication networks will increasingly utilize ultra-high-bandwidth high-frequency bands (such as millimeter waves, terahertz, and visible light bands) to meet users' rapidly growing demand for transmission speeds. In high-band mobile communication networks, users and base stations utilize directional, narrow beams to improve directional gain and ensure successful signal reception. When a base station enters a dormant state, users previously connected to that base station must switch to a more distant base station to maintain their connection. However, the channel fading between these users and the newly connected base station is greater, forcing users and base stations to use narrower beams to ensure communication quality. Narrower beams require users and base stations to scan more directions during initial access (the process by which users and base stations discover each other and establish a connection through beam alignment) to find the optimal transmit and receive direction, thereby increasing initial access latency. On the other hand, when fewer base stations are dormant, users can connect to closer base stations, although base station energy consumption increases. Therefore, both base stations and users can utilize wider beams for transmission, reducing initial access latency. However, the impact of base station sleep on beamwidth selection and initial access delay has not been considered in existing research. Furthermore, minimizing initial access delay through beamwidth selection under different base station sleep states remains to be investigated. Therefore, studying a comprehensive optimization scheme for base station sleep control and beamwidth selection in high-band mobile communication networks, to achieve the goal of balancing base station energy consumption and user initial access delay, has important theoretical research significance and application value. Summary of the Invention

[0003] In response to the above defects or improvement needs of the existing technology, the present invention provides a comprehensive optimization method for base station sleep control strategy and beam width selection strategy, thereby solving the problem of difficulty in balancing base station energy consumption and user initial access delay in traditional methods.

[0004] To achieve the above objectives, according to a first aspect of the present invention, a comprehensive optimization method for a base station sleep control strategy and a beam width selection strategy is provided, comprising:

[0005] An optimization model is constructed with the goal of minimizing the total energy consumption and average beam scanning delay of the base station of the target cell, and the optimization model is solved to obtain the base station sleep decision and beam width selection decision of the target cell;

[0006] Wherein, the optimization model is:

[0007]

[0008] a k,j ≤b j ,

[0009] a k,j ∈{0,1},

[0010] b j ∈{0,1},

[0011] α and β are weight factors, J and K are the number of base stations and users in the target cell, respectively; are the number of beams required for coarse beam scanning and fine beam scanning for the jth base station, The number of beams required for coarse beam scanning and fine beam scanning for the kth user, b j Indicates whether the jth base station is in sleep state, b j =0 means the jth base station is in sleep state, b j =1 indicates that the jth base station is in working state; a k,j Indicates the association relationship between the base station and the user, a k,j =1 means the kth user is associated with the jth base station, a k,j =0 means that the kth user is not associated with the jth base station, a is the association relationship vector between all base stations and users, b is the sleep decision vector of all base stations, Selecting decision vectors for beam widths when performing coarse beam scanning and fine beam scanning between all base stations and users; represents the constant power consumption when the jth base station is turned on; represents the power consumption generated when the jth base station turns on a radio frequency chain; P Ref and D Ref are the total power consumption constant and the average delay constant respectively; Ω BS Ω is an optional set of the number of antennas of the base station BS ;Ω UE An optional set of the number of antennas for the user; M j is the upper limit of the number of users that can be served by the j-th base station; is the mathematical expectation of the average beam scanning delay.

[0012] According to a second aspect of the present invention, there is provided an electronic device comprising: a computer-readable storage medium and a processor;

[0013] The computer-readable storage medium is used to store executable instructions;

[0014] The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the method according to the first aspect.

[0015] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the method according to the first aspect.

[0016] According to a fourth aspect of the present invention, there is provided a computer program product comprising a computer program or instructions, which implement the method according to the first aspect when executed by a processor.

[0017] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0018] The method provided by the present invention takes minimizing the total energy consumption of the base station and the average beam scanning delay as the optimization goal, establishes the constraint conditions for the base station sleep control and the beam width selection, and builds an optimization model. The base station sleep decision and the beam width selection decision of the target cell are obtained by solving the optimization model. The method fully considers the impact of the base station sleep on the beam width selection and the initial access delay, and can solve the problem of the difficulty in balancing the base station energy consumption and the user initial access delay in the traditional base station sleep control and beam width selection methods. By comprehensively optimizing the base station sleep control and beam width selection, the base station energy consumption is reduced without significantly affecting the user delay performance.

[0019] As a further preference, the method provided by the present invention adopts a deep reinforcement learning algorithm to solve the above-mentioned decision-making problem, and uses DQN to learn and make decisions in complex environments to achieve dynamic adjustment of base station sleep and beam width. Compared with traditional solution methods for NP-hard problems, it can reduce computational complexity and accuracy.

[0020] As a further preference, the method provided by the present invention uses a sector antenna model to approximate the actual antenna radiation pattern to obtain a calculation formula for the beam width, thereby being as close as possible to the antenna radiation characteristics in the actual operating environment of the base station and improving the calculation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A flow chart of a comprehensive optimization method for a base station sleep control strategy and a beam width selection strategy provided by an embodiment of the present invention;

[0022] Figure 2 A schematic diagram of the DQN structure used in the comprehensive optimization method for the base station sleep control strategy and beam width selection strategy provided in an embodiment of the present invention;

[0023] Figure 3 Schematic diagram of the DQN training process based on experience replay adopted by the comprehensive optimization method of the base station sleep control strategy and beamwidth selection strategy provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0025] An embodiment of the present invention provides a comprehensive optimization method for a base station sleep control strategy and a beam width selection strategy, including:

[0026] An optimization model is constructed with the goal of minimizing the total energy consumption and average beam scanning delay of the base station of the target cell, and the optimization model is solved to obtain the base station sleep decision and beam width selection decision of the target cell.

[0027] First, the communication system model of the target cell is introduced.

[0028] The target cell includes a millimeter wave cellular communication system consisting of J base stations and K users. The location distribution of base stations and users obeys the density ρ BS and ρ UE Poisson point process, let r be the distance between any base station and user, the probability density function (PDF) of its distribution is Let r be the distance between a user and the base station serving the user. Without loss of generality, in the base station and user distribution model, it can be assumed that the user is served by the base station closest to it. Then the probability density function (PDF) of r can be obtained by a two-dimensional Poisson process, expressed as

[0029] The base station and user equipment are equipped with M BS and M UE For a uniform linear array of active antenna elements, the half-power beamwidth (HPBW) can be given by and Estimation, d is the antenna spacing;

[0030] The channel between the base station and the user follows a statistical channel model. The channel consists of a set of C different clusters, each corresponding to a different scattering path. The arrival angle and departure angle of these (cluster) paths follow a uniform distribution (0 to 2π). Let the number of distribution clusters used by the base station sector and the user sector be K and L, respectively, satisfying L≤C and K≤L. If the AoD / AoA of a cluster is not within the range of the base station / user sector, then the cluster cannot be used to send / receive synchronization signals. The distribution clusters that the base station and the user may use are indexed by u=1,2,...,U and v=1,2,...,V, respectively. Then the distribution of the uth cluster on the base station side can be expressed as satisfy Similarly, the distribution of the vth cluster on the user side can be expressed as satisfy N BS is the number of beams of the base station, N UE is the number of beams for the user.

[0031] To fully consider the impact of base station sleep on beam width selection and initial access delay, and to reduce base station energy consumption without significantly affecting user latency performance, an optimization model for base station sleep and beam width selection is constructed with the goal of minimizing total base station energy consumption and average beam scanning delay:

[0032]

[0033] C5:a k,j ≤b j ,

[0034] C6:a k,j ∈{0,1},

[0035] C7:b j ∈{0,1},

[0036] Among them, α∈(0,1) and β∈(0,1) are weight factors, indicating the proportion of total power consumption and delay in the optimization target, satisfying α+β=1; are the number of beams required for coarse beam scanning and fine beam scanning for the jth base station, are the number of beams required for coarse beam scanning and fine beam scanning for the kth user, respectively; J and K represent the number of base stations s and users s in the millimeter wave cellular network, respectively; b j Indicates whether the jth base station is in sleep state. If b j =0, it means the base station is in sleep state. j =1, it means the base station is in working state; a k,j Indicates the association relationship between the base station and the user, a k,j =1 means the kth user is associated with the jth base station, a k,j =0 means that the kth user is not associated with the jth base station; represents the constant power consumption when the jth base station is turned on; represents the power consumption generated when the jth base station turns on a radio frequency chain; P Ref and D Ref They are respectively a sufficiently large total power consumption constant and an average delay constant, which are used to normalize the total power consumption and average delay in the optimization target; the constraint condition C1 represents The value needs to be selected from the optional set Ω of the number of antennas of the base station BS Select; constraint C2 represents An optional set of the number of antennas required from the user Ω UE Constraint C3 indicates that each user can be served by at most one base station; Constraint C4 indicates the upper limit of the number of users that can be served by the j-th base station, that is, the upper limit of the number of RF chains that can be activated is M j ; Constraint C5 means that only base stations in non-dormant state can provide services to users; Constraints C6 and C7 mean that a k,j and b j Can only take values ​​from the set {0, 1}.

[0037] By solving the above objective function, the beam width selection decision vector for all base stations and users when performing coarse beam scanning and fine beam scanning can be obtained: And the optimal sleep decision vector b of all base stations, according to b, the sleep strategy of each base station in the target cell can be obtained. The beam width of each base station and user when performing coarse beam scanning and fine beam scanning can be calculated.

[0038] It is understood that the beam width can be calculated based on the number of beams, and the calculation formula for the beam width is determined by the antenna model used to approximate the actual antenna pattern, such as the uniform linear array model, the Taylor distribution array model, etc. To facilitate analysis and to closely follow the antenna radiation characteristics under the actual operating environment of the base station, the embodiment of the present invention preferably uses a sector antenna model to approximate the actual antenna pattern to obtain the calculation formula for the beam width:

[0039] Assume that η and η′ are the angles of deviation from the broadside of the user and the base station, respectively, then the antenna gain G BS (η′) and G UE The calculation formula of (η) is set up and are the beam widths when the base station performs coarse beam scanning and fine beam scanning, respectively. If the beam covers the entire two-dimensional plane, the number of beams required is:

[0040] Mathematical expectation of average beam scanning delay The calculation method is:

[0041]

[0042] Among them, f r (r) is the PDF function of the distribution of base stations and users given in S11; is the expected delay of beam scanning between the base station and the user at a distance r.

[0043] The expected delay of beam scanning between the base station and the user in the mathematical expectation of the average beam scanning delay The calculation method is:

[0044]

[0045] in, is the average false detection probability during the coarse beam scanning phase; The average false detection probability during the beamlet scanning phase; T SS N is the period of the synchronization signal block (SSB) burst set in 5G NR; tot is the total number of beams required in the entire beam scanning process, given by Calculated; N SS is the number of SSBs in one SSB burst set period; T slot is the duration of a time slot, and n and m are the number of scans required for coarse beam and fine beam scanning.

[0046] Specifically, the average false detection probability during the beam scanning phase is The calculation method is:

[0047]

[0048] Where C is the number of channel distribution clusters, C ~ max{Poisson(κ),1}; is the average false detection probability under the given condition C=C′, Calculated; For sending SSB through a fixed sector of a base station, the false detection probability of all users failing to receive is given by Calculated; is the false detection probability of user reception failure in a fixed sector under given L and K values, which is given by Calculated, where P BS is the transmission power of the base station, PL is the path loss between the base station and the user, γ L,K is the power scaling factor, G BS and G UE are the antenna gains of the base station and the user respectively, P th It is the minimum power threshold for users to successfully receive signals.

[0049] The above objective function is an NP-hard problem and can be solved using existing methods, such as the relaxation constraint method. Taking into account the computational complexity and accuracy, preferably, the embodiment of the present invention uses the DQN reinforcement learning model to solve the above objective function.

[0050] Specifically, the base station sleep control and beamwidth selection problem is modeled as an MDP, with the network controller that determines the base station sleep and beamwidth control in the target area as the intelligent agent. A DQN reinforcement learning model is constructed, including:

[0051] (1) Based on the optimization objective of the resource allocation optimization problem model, the entire base station sleep control and beamwidth selection process is described as a Markov decision process (MDP). The DQN reinforcement learning algorithm is used as a training model to find the sleep strategy and beamwidth control strategy of each base station for the target cell.

[0052] (2) Set the state space of the agent to:

[0053]

[0054] Among them, a t-1 is the association status information of all users and base stations in the previous time slot, b t-1 is the sleep state information of all base stations in the previous time slot, It is the channel state information of all base stations and users in the current time slot.

[0055] (3) Set the action space of the agent to:

[0056] a t ={b t ,d t}

[0057] Among them, b t represents the sleep decision of all base stations in time slot t, d t represents the beamwidth selection decision of all base stations in time slot t,

[0058] (4) Set the agent's reward function to the negative value of the cost function:

[0059]

[0060] Based on the state space and action space defined above, a DQN network is constructed. After parameter initialization and setting the experience replay pool, training is performed using the ε-greedy strategy, updating the parameters using the loss function, and continuing training until the model converges.

[0061] The DQN structure is as follows Figure 2 As shown in Figure 1, a trained neural network can map the input system state to the Q-function values ​​of different behaviors in that state. However, due to the correlation between the sample data collected by the agent during training, directly using neural networks for Q-learning may lead to problems such as unstable training process and failure to converge.

[0062] In order to solve the problem of correlation between sample data, a deep reinforcement learning scheme based on experience replay training is used to train the DQN network. The DQN training process based on experience replay is as follows: Figure 3 As shown, the agent will explore multiple experience segments generated in the process of environment t =(s t ,a t ,r t ,s t+1 ) is stored and randomly sampled from it after a period of time to obtain an experience fragment. The input (state and behavior) of this sampled experience fragment is sent to the DQN, and the output (the transferred state and the reward obtained) is sent to the target neural network. Based on the sampled experience fragment, the DQN weights are updated by minimizing the following loss function: w i and are the weights of the DQN and target neural networks at the time of update in round i, respectively. The aforementioned loss function is the mean squared error between the DQN and target neural networks' estimates of the Q function value, and the solution that minimizes it can be found using stochastic gradient descent. To reduce the correlation between the DQN and target neural networks, the target neural network's weights are updated less frequently than the DQN. After DQN training is complete, the agent determines its strategy based on the Q function value output by the DQN.

[0063] Finally, the observation of the system state is input into the trained DQN reinforcement learning model to obtain the corresponding base station sleep control and beam width selection strategy.

[0064] An embodiment of the present invention provides an electronic device, comprising: a computer-readable storage medium and a processor;

[0065] The computer-readable storage medium is used to store executable instructions;

[0066] The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the method described in any one of the above embodiments.

[0067] An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the method described in any of the above embodiments.

[0068] An embodiment of the present invention provides a computer program product, including a computer program or instructions, which implements the method described in any of the above embodiments when executed by a processor.

[0069] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A comprehensive optimization method for base station sleep control strategy and beam width selection strategy, characterized in that: include: An optimization model is constructed with the goal of minimizing the total energy consumption and average beam scanning delay of the base station of the target cell, and the optimization model is solved to obtain the base station sleep decision and beam width selection decision of the target cell; Wherein, the optimization model is: a k,j ≤b j , a k,j ∈{0,1}, b j ∈{0,1}, α and β are weight factors, J and K are the number of base stations and users in the target cell, respectively; are the number of beams required for coarse beam scanning and fine beam scanning for the jth base station, The number of beams required for coarse beam scanning and fine beam scanning for the kth user, b j Indicates whether the jth base station is in sleep state, b j =0 means the jth base station is in sleep state, b j =1 indicates that the jth base station is in working state; a k,j Indicates the association relationship between the base station and the user, a k,j =1 means the kth user is associated with the jth base station, a k,j =0 means that the kth user is not associated with the jth base station, a is the association relationship vector between all base stations and users, b is the sleep decision vector of all base stations, Selecting decision vectors for beam widths when performing coarse beam scanning and fine beam scanning between all base stations and users; represents the constant power consumption when the jth base station is turned on; represents the power consumption generated when the jth base station turns on a radio frequency chain; P Ref and D Ref are the total power consumption constant and the average delay constant respectively; Ω BS Ω is an optional set of the number of antennas of the base station BS ;Ω UE An optional set of the number of antennas for the user; M j is the upper limit of the number of users that can be served by the j-th base station; is the mathematical expectation of the average beam scanning delay.

2. The method according to claim 1, wherein Using a pre-trained DQN reinforcement learning model to solve the optimization model; Among them, the state space of the agent a t-1 is the association status information of all users and base stations in the previous time slot, b t-1 is the sleep state information of all base stations in the previous time slot, The channel state information of all base stations and users in the current time slot; The action space a of the agent t ={b t ,d t }; b t represents the sleep decision vector of all base stations in time slot t, d t represents the decision vector for selecting the beam width of all base stations in time slot t, 3. The method according to claim 2, wherein The agent's reward function 4. The method according to claim 1, wherein and The beam widths for coarse beam scanning and fine beam scanning for the base station are respectively, and The beam widths used when the user performs coarse beam scanning and fine beam scanning, respectively.

5. An electronic device, characterized in that: include: Computer-readable storage media and processor; The computer-readable storage medium is used to store executable instructions; The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the method according to any one of claims 1 to 4.

6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the method according to any one of claims 1 to 4.

7. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Method by which terminal determines beam in wireless communication system and terminal therefor

    CN110089048A

  • 5G base station collaborative energy saving method based on near-end strategy optimization

    CN117979397A