Base Station Intelligent Energy Saving Method Based on Fuzzy Logic System and Deep Reinforcement Learning

By applying a combination method of fuzzy logic system and deep reinforcement learning in the base station, the energy consumption and user access strategy of the base station are evaluated and optimized, and the problems of insufficient energy-saving decision-making basis and insufficient response capabilities in the existing technology are solved, and more efficient energy saving and service quality balance are achieved.

CN115568006BActive Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211149492.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-05-27
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

The existing base station energy-saving technology lacks a more appropriate decision-making basis, it is difficult to balance energy conservation and user service quality, and it lacks the ability to respond to emergencies in complex and changeable mobile networks.

Method used

The intelligent energy-saving method of base stations based on fuzzy logic systems and deep reinforcement learning is adopted, and the energy consumption evaluation factor between users and micro base stations is evaluated through the fuzzy logic system, and the global energy-saving strategy and user access strategy are formulated in combination with deep reinforcement learning algorithms to balance energy saving and service quality, and improve the ability to respond to emergencies.

Benefits of technology

It provides a more appropriate energy-saving decision-making basis, balances the energy consumption of base stations and the quality of user services, and improves the ability to prevent and respond to emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115568006B_ABST
    Figure CN115568006B_ABST
Patent Text Reader

Abstract

The present invention proposes a smart energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning, belonging to the field of base station energy saving. Specifically, in a communication scenario with macro base stations and micro base stations, each micro base station reports its user and base station information to the macro base station in real time. The macro base station counts the total number of all users in the scenario and determines whether it changes. If so, the macro base station uses an energy-saving access algorithm based on fuzzy logic and deep reinforcement learning to obtain the energy-saving control strategy of the micro base stations and the access base station selection strategy of the users, and issues them to the micro base stations for execution, taking into account both base station energy saving and user experience. Otherwise, each micro base station uses an energy-saving handover algorithm based on fuzzy logic to adjust the user access distribution in real time, thereby further preventing emergencies. The present invention balances the contradiction between energy saving and user service quality; utilizes the prediction ability of the energy consumption evaluation factor for future energy consumption to perform predictive user handovers and make a more proactive response to emergencies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of base station energy saving, and particularly relates to a base station intelligent energy saving method based on a fuzzy logic system and deep reinforcement learning. Background Art

[0002] With the development of mobile networks, users' demand for high-rate data transmission is increasing continuously. To solve the problems of insufficient signal coverage and transmission resources of macro base stations, the construction and deployment scale of micro base stations are increasing continuously, and a heterogeneous network with coexistence of macro / micro base stations has become a typical mobile network scenario. However, the energy consumption of base stations accounts for a relatively large proportion in the total energy consumption of mobile communication networks. While deploying more base stations to meet user needs, how to take measures to reduce the energy consumption of base stations is a key issue to be solved for the green development of future networks.

[0003] It is difficult to guarantee the applicability and accuracy of algorithms based on manual analysis. The intelligent energy saving technology based on artificial intelligence can reduce manual participation and improve the automation and intelligence of energy saving. Existing intelligent energy saving technologies focus on statistical parameters on the base station side, learn and predict service rules, and adopt different base station shutdown strategies for different time periods and different service rules.

[0004] However, the above technologies have defects:

[0005] 1) The existing technology uses the predicted traffic volume as the basis for energy saving decisions. In fact, the transmission energy consumption is also related to the signal transmission quality. Therefore, the existing technology lacks a more appropriate basis for energy saving decisions;

[0006] 2) Simply focusing on taking energy saving measures on the base station side, but the access selection of users also has an important impact on the energy saving effect and user service quality;

[0007] 3) In an increasingly complex and changeable mobile network, the ability of the service scenario rules learned from the statistical parameters on the base station side to cope with emergencies is insufficient. Summary of the Invention

[0008] In view of the above problems, the present invention provides a base station intelligent energy saving method based on a fuzzy logic system and deep reinforcement learning, aiming to focus on user-side parameters, estimate the energy consumption required by users, and based on this, give an energy saving strategy for the global base station and an access strategy for global users, balance the contradiction between energy saving and user service quality, and at the same time make more active prevention and response to emergencies.

[0009] The base station intelligent energy saving method based on a fuzzy logic system and deep reinforcement learning includes:

[0010] Step 1: In a communication scenario with macro base stations and micro base stations, each micro base station reports its user and base station information to the macro base station in real time;

[0011] The described user information includes: the number of users, the user's historical signal reception situation, and the user's service type; the described base station information includes: the power level of the micro base station.

[0012] Step 2: The macro base station counts the total number of all users in the scenario and determines whether it changes. If it does, the macro base station uses an energy-saving access algorithm based on fuzzy logic and deep reinforcement learning to obtain the energy-saving control strategy of the micro base station and the access base station selection strategy of the users, and sends them to the micro base station for execution, taking into account both base station energy saving and user experience; otherwise, proceed to Step 3.

[0013] Set the initial number of users for each base station according to the actual user service capabilities of each base station.

[0014] The specific steps are as follows:

[0015] Step 201: For a certain target micro base station B, input the information of the user u to be connected to this micro base station and the base station information of the target micro base station B as variables into the fuzzy logic system.

[0016] Step 202: Set the membership functions corresponding to each fuzzy concept in the fuzzy logic system.

[0017] The membership degree represents the degree to which the input / output exact value of the fuzzy logic system conforms to this fuzzy concept; the same input / output exact value conforms to multiple fuzzy concepts simultaneously.

[0018] The linear membership functions are:

[0019] Triangular membership function

[0020] where a 1 、b 1 、c 1 respectively represent the abscissas of the three vertices of the triangle; represents the upper and lower limit ranges of the input exact value.

[0021] Trapezoidal membership function

[0022] where a 2 、b 2 、c 2 、d 2 respectively represent the abscissas of the four vertices of the trapezoid; represents the upper and lower limit ranges of the input exact value.

[0023] The non-linear membership functions are:

[0024] Gaussian membership function

[0025] a 3 represents the mean of the Gaussian distribution, affecting the abscissa of the center of the curve; b3 represents the variance of the Gaussian distribution, which affects the width of the curve;

[0026] Bell-shaped membership function

[0027] a 4 represents the length that affects the bell shape, b 4 affects the width of the bell shape; c 4 represents the abscissa of the center of the bell shape;

[0028] Step 203: For a single input variable S, substitute the exact value of the input variable into each membership function, and output the membership degree of the input variable for each fuzzy concept;

[0029] Similarly, obtain the membership degrees of each input variable for each fuzzy concept;

[0030] Step 204: Map the membership degrees of each input variable for all fuzzy concepts through fuzzy rules to obtain the membership degrees of the corresponding output variable for each fuzzy concept, and further defuzzify them to convert them into the exact value of the final output variable, which is the energy consumption evaluation factor between user u and target micro base station B;

[0031] Step 205: At time t, statistically collect the set of energy consumption evaluation factors between each user and each micro base station as the state space of the reinforcement learning model; use the transmission power of each micro base station and the access base station numbers of each user as the action decisions; use the weighted sum of the energy consumption and user service performance indicators in the heterogeneous network with macro / micro base station coexistence as the reward function; and deploy the trained reinforcement learning model on the macro base station;

[0032] The state space set is:

[0033]

[0034] where M and N are the total number of all users and the total number of all micro base stations in the heterogeneous network, respectively, is the energy consumption evaluation factor between the i-th micro base station and the j-th user;

[0035] The action decisions are:

[0036] a t ={[p 1 , p 2 , …, p i , …, p N , [β 1 , β 2 , …, β j , …, β M};

[0037] p i is the transmission power of the i-th micro base station, βj The access base station number of the j-th user;

[0038] The reward function is:

[0039] r t = -α 1 W 1 -α 2 W 2 -α 3 W 3 +α 4 Q

[0040] where α 1 , α 2 , α 3 , α 4 are normalized weighting coefficients; W 1 is the total network transmission energy consumption at time t, W 2 is the constant energy consumption of the base station at time t, W 3 is the sum of the energy consumption evaluation factors between all users and all base stations in the network at time t; Q is the service quality reward for users;

[0041] Step 206: The macro base station receives the newly reported energy consumption evaluation factor information from the micro base stations, uses the trained reinforcement learning model to output the power control decisions for all micro base stations and the access decisions for all users, and sends the decisions to each micro base station;

[0042] Step 207: The micro base stations receive the sent decisions, adjust their own transmission powers according to the power control decisions, and judge whether the users they serve need to switch to other base stations according to the access decisions. If so, send the corresponding access decisions to the users, otherwise do nothing; Each user receives the access decision and applies for access to its new target base station; Each new target base station feeds back the access request.

[0043] Step 3: If the number of users remains unchanged, each micro base station uses the energy-saving handover algorithm based on fuzzy logic to adjust the user access distribution in real time, so as to further prevent emergencies.

[0044] The specific steps are as follows:

[0045] Step 301: Each micro base station records the average energy consumption [θ 1 , θ 2 , …, θ N in the past T time, and updates its own energy consumption threshold:

[0046] [Θ 1 , Θ 2 , …, Θ i , …, Θ N = ω × [θ 1 , θ 2,…,θ N

[0047] Θ i is the energy consumption threshold of the i-th small base station, and ω is the artificially set energy consumption threshold coefficient;

[0048] Step 302: Each small base station uses a fuzzy logic system to calculate and monitor in real time the energy consumption evaluation factor and the sum of energy consumption evaluation factors between each user accessing the base station and the base station:

[0049] [F 1 ,F 2 ,…,F i ,…,F N ;

[0050] F i is the sum of the energy consumption evaluation factors of all users of the i-th small base station;

[0051] Step 303: For each small base station, if the sum of the energy consumption evaluation factors of the small base station is greater than the energy consumption threshold, it means that the current transmit power configuration of the small base station is low; set the small base station as the small base station to be adjusted and report it to the macro base station, and enter Step 304; otherwise, do nothing;

[0052] Step 304: For the current small base station to be adjusted n1, the macro base station traverses the users served by the small base station to be adjusted n1, and finds the user m1 with the largest energy consumption evaluation factor from the current small base station to be adjusted n1 as the user to be switched;

[0053] Step 305: The macro base station finds the base station k1 with the smallest energy consumption evaluation factor from other base stations except the small base station to be adjusted n1 for the user to be switched m1 as the candidate target base station;

[0054] Step 306: Use the sum of the energy consumption evaluation factors F k1 of the base station k1, and the energy consumption evaluation factor f m1,k1 between the user to be switched m1 and the base station k1, and judge whether the sum of the two is less than the energy consumption threshold, that is, F k1 +f m1,k1 <Θ k1 , if so, switch the user m1 to the base station k1, and the base station k1 cannot be used as a candidate target base station; otherwise, if no qualified small base station is found, connect the user to be switched to the macro base station;

[0055] In particular, since the user m1 was not connected to the base station k1 before, the change in the received signal strength within the recent T seconds in the input item is set to 0 when calculating f m1,k1 ;

[0056] ​Step 307: The macro base station repeatedly searches for the next micro base station to be adjusted n2, the user to be switched m2, and the candidate target base station k2, and switches the user until there is no micro base station to be adjusted.

[0057] The advantages of the present invention are as follows:

[0058] 1) The intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning of the present invention represents the environmental state based on the energy consumption evaluation factor of fuzzy logic, comprehensively considers base station and user-side information such as channels and services, and provides a more appropriate decision-making basis;

[0059] 2) The intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning of the present invention, the designed energy-saving access algorithm based on deep reinforcement learning gives the access decision of the user while giving the power control decision of the micro base station, balancing the contradiction between energy saving and user service quality;

[0060] 3) The intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning of the present invention, the designed energy-saving handover algorithm based on fuzzy logic utilizes the prediction ability of the energy consumption evaluation factor for future energy consumption, performs anticipatory user handover, and makes a more active response to emergencies. Description of the Drawings

[0061] Figure 1 is the flowchart of the intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning of the present invention;

[0062] Figure 2 is the flowchart of the energy-saving access algorithm based on fuzzy logic and deep reinforcement learning of the present invention;

[0063] Figure 3 is the flowchart of the energy-saving handover algorithm based on fuzzy logic of the present invention. Detailed Embodiments

[0064] The following further explains the present invention in detail in conjunction with embodiments and drawings.

[0065] The intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning of the present invention, as Figure 1 shown, includes:

[0066] Step 1: In a communication scenario with a macro base station and micro base stations, each micro base station reports its user and base station information to the macro base station in real time;

[0067] The user information includes: the number of users, the historical signal reception situation of the users, and the user service type; the base station information includes: the power level of the micro base station.

[0068] Step 2: The macro base station counts the total number of all users in the scenario and determines whether it changes. If it does, the macro base station uses an energy-saving access algorithm based on fuzzy logic and deep reinforcement learning to obtain the energy-saving control strategy of the micro base station and the access base station selection strategy of the users, and sends them to the micro base station for execution, taking into account both base station energy saving and user experience; otherwise, go to Step 3;

[0069] Set the initial number of users for each macro and micro base station according to their actual user service capabilities.

[0070] The energy-saving access algorithm based on fuzzy logic and deep reinforcement learning includes: an energy consumption evaluation factor based on fuzzy logic and an energy-saving access algorithm based on deep reinforcement learning; as Figure 2 shown, the specific steps are as follows:

[0071] Step 201: For a target micro base station B, input the information of user u to be connected to this micro base station and the base station information of target micro base station B into the fuzzy logic system as variables;

[0072] Specifically include: the current signal reception situation of this user, the change situation of the received signal strength in the recent T seconds, the transmission requirement of this user, and the transmission power of this base station. The first two items can be calculated from the reported historical signal reception situation of the user, and the last two items are mapped from the reported user service type and the reported transmission power level respectively; in the embodiment, T takes the value of 300;

[0073] Step 202: Set the membership functions corresponding to each fuzzy concept in the fuzzy logic system;

[0074] Fuzzy concepts are used to evaluate input / output variables. Each fuzzy concept corresponds to a membership function curve, and the value of the membership degree is between 0 and 1, which is obtained by substituting the exact value of the input variable into the membership function; it represents the degree to which the exact value of the input / output variable of the fuzzy logic system conforms to this fuzzy concept;

[0075] The abscissa range where the membership function value is non-zero reflects the range of the exact value of the input / output that conforms to this fuzzy concept. The membership function curves overlap, and the same exact value of the input / output conforms to multiple fuzzy concepts at the same time;

[0076] Linear membership functions include:

[0077] Triangular membership function

[0078] where a 1 、b 1 、c 1 represent the abscissas of the three vertices of the triangle respectively; they represent the upper and lower limit ranges of the input exact value.

[0079] Trapezoidal membership function

[0080] Among them, a 2 , b 2 , c 2 , d 2 respectively represent the abscissas of the four vertices of the trapezoid; represent the upper and lower limit ranges of the input exact values.

[0081] The non-linear membership functions are as follows:

[0082] Gaussian membership function

[0083] a 3 represents the mean of the Gaussian distribution and affects the abscissa of the center of the curve; b 3 represents the variance of the Gaussian distribution and affects the width of the curve;

[0084] Bell-shaped membership function

[0085] a 4 represents the length affecting the bell shape, b 4 affects the width of the bell shape; c 4 represents the abscissa of the center of the bell shape;

[0086] The parameters of all membership functions are obtained through neural network training.

[0087] Step 203, for a single input variable S, substitute the exact value of the input variable into each membership function for fuzzy inference, and output the membership degree of the input variable to each fuzzy concept;

[0088] Similarly, obtain the membership degrees of each input variable to each fuzzy concept;

[0089] Step 204, map the membership degrees of each input variable to all fuzzy concepts to obtain the membership degrees of the corresponding output variable to each fuzzy concept, and further use the centroid method for defuzzification to convert it into the exact value of the final output variable, which is the energy consumption evaluation factor of user u and target micro base station B;

[0090] The fuzzy rule takes the membership degree of each input variable to a single input fuzzy concept and the fuzzy concept as the premise, maps to obtain the membership degree of the output variable to the output fuzzy concept, constructs a T-S type fuzzy logic system with an adaptive neuro-fuzzy system, obtains the fuzzy rule by training the neural network, and uses the centroid method to convert the membership degree of the output variable mapped by the fuzzy rule into the exact value of the output variable; use historical data to train the adaptive neuro-fuzzy system, and the trained fuzzy logic system outputs the energy consumption evaluation factor. When user u accesses the target micro base station B, the energy consumption situation within the next T seconds, the larger the value of the energy consumption evaluation factor, the greater the energy consumption, and vice versa, the smaller the energy consumption;

[0091] For example:

[0092] First, the exact values of the two input variables are P and Q respectively. There are three corresponding fuzzy concepts for each input variable, namely "high", "medium", "low" and "fast", "general", "slow". It is assumed that there are three output fuzzy concepts: "increase", "remain unchanged", and "decrease".

[0093] Let the membership degree of the exact value P of the input variable conforming to "high" be x, the membership degree conforming to "medium" be y, and the membership degree conforming to "low" be z; let the membership degree of the exact value Q of the input variable conforming to "fast" be u, the membership degree of "general" be v, and the membership degree of "low" be t.

[0094] The above membership degrees x, y, z, u, v, t are obtained by substituting the input variables P and Q into each membership function. The process of substituting and calculating the membership degrees is "fuzzification".

[0095] Then, the fuzzy inference rule is mapped as: If the input variable P belongs to "high" and the input variable Q belongs to "fast", then the fuzzy concept that the output variable belongs to is "increase", and the membership degree belonging to "increase" is obtained by x and u according to neural network training.

[0096] Finally, according to the membership degrees of the output variable belonging to the three fuzzy concepts of "increase", "remain unchanged", and "decrease", and there are also membership functions (obtained by neural network training) of the three fuzzy concepts of "increase", "remain unchanged", and "decrease", using the "centroid method" can combine the membership degrees and membership functions to give the exact value of the output variable. This process is "defuzzification"; the exact value of the final output variable is the energy consumption evaluation factor.

[0097] The present invention constructs a T-S type fuzzy logic system with an adaptive neural fuzzy system, uses simulation and historical data to train the neural network to obtain the above input / output membership functions and fuzzy inference rules, and deploys the trained adaptive neural fuzzy system in each macro and micro base station.

[0098] Step 205: At time t, count the set of energy consumption evaluation factors between each user and each micro base station as the state space of the reinforcement learning model; use the transmission power of each micro base station and the access base station number of each user as the action decision; use the weighted sum of the energy consumption and user service performance indicators in the heterogeneous network where the macro / micro base stations coexist as the reward function; and train the reinforcement learning model and then deploy it in the macro base station.

[0099] The state space set is:

[0100]

[0101] Where M and N are the total number of all users and the total number of all micro base stations in the heterogeneous network, respectively, is the energy consumption evaluation factor between the i-th micro base station and the j-th user;

[0102] The action decision is:

[0103] a t ={[p 1 , p 2 ,…, p i ,…, p N ,[β 1 , β 2 ,…, β j ,…, β M};

[0104] p i is the transmission power of the i-th micro base station, and β j is the access base station number of the j-th user;

[0105] The reward function is:

[0106] r t =-α 1 W 1 -α 2 W 2 -α 3 W 3 +α 4 Q

[0107] Where α 1 , α 2 , α 3 , α 4 are normalized weighting coefficients, whose function is to normalize W 1 , W 2 , W 3 and Q, and at the same time attach different weights to adjust the influence degree of each item on the reward value result; in the embodiment, let the total transmission energy consumption, total constant energy consumption, and sum of energy consumption evaluation factors when all micro base stations transmit at full power be V 1 , V 2 , V 3 , respectively, then we can take W 1 is the total network transmission energy consumption at time t, W 2 is the constant energy consumption of the base station at time t, and W 3 is the sum of the energy consumption evaluation factors between all users and all base stations in the network at time t; Q is the service quality reward of the user;

[0108] The reward function consists of 4 parts, and the specific calculation method includes:

[0109]

[0110] Among them, P 0 is the maximum transmission power of the base station, which is set to 10W in the embodiment, and p i ∈[0, 1] is the base station power level given by the energy-saving access algorithm based on deep reinforcement learning, and τ i is the longest time for the base station to send data to serve users, and W 1 The smaller it is, the smaller the total transmission energy consumption of the base station;

[0111]

[0112] Among them p i being 0 indicates that the micro base station is in a dormant state, and p off is the constant power when the base station is in a dormant state, which is set to 5W in the embodiment, and p on is the constant power when the base station is not in a dormant state, which is set to 40W in the embodiment, and W 2 The smaller it is, the smaller the total constant energy consumption of the base station;

[0113]

[0114] W 3 is the sum of the energy consumption evaluation factors between all users and all base stations. The smaller W 3 is, the smaller the total transmission energy consumption of the base station in a future period of time, and the decision made can prevent emergencies;

[0115]

[0116] Among them R i is the actual transmission rate of the user, and R 0 is the minimum acceptable transmission rate of the user, which is set to 100Mbps in the embodiment. The larger Q is, the better the user service quality is, and the decision made can take into account the user service quality.

[0117] Step 206: The macro base station receives the newly reported energy consumption evaluation factor information of the micro base stations, uses the trained reinforcement learning model to output the power control decisions of all micro base stations and the access decisions of all users, and sends the decisions to each micro base station;

[0118] In reinforcement learning, it is necessary to model the Markov decision process, that is, define the input state S and action A; when the action a is performed at a certain time t and the environment state is s, the state s' at the next time and the reward r for performing action a can be obtained. When training the neural network in reinforcement learning, the network continuously tries actions in different states and updates the parameters in the network. The goal is to make the reward r after performing the action larger, so that after the training is completed, the network can give the optimal action decision for different states.

[0119] Therefore, s, a, and r are all used when training the reinforcement learning network. In particular, s and a are also the input and output of the actual deployment after the reinforcement learning network training is completed. However, r is only useful during training to tell the neural network in which direction to update the parameters.

[0120] The two decisions refer to the fact that the definition of a contains two arrays, p and β, which respectively determine the transmission power of the micro base station and the base station to which the user accesses; the decision a can be obtained by inputting the s set obtained from the network and deploying it in the reinforcement learning intelligent model of the macro base station.

[0121] Step 207, the micro base station receives the decision sent down, adjusts its own transmission power according to the power control decision, and determines whether the user it serves needs to switch access to other base stations according to the access decision. If so, the corresponding access decision is sent down to the user, otherwise no processing is performed; each user receives the access decision and applies for access to their respective new target base stations; each new target base station feeds back the access request.

[0122] Step 3: If the number of users remains unchanged, each micro base station uses an energy-saving switching algorithm based on fuzzy logic to adjust the user access distribution in real time, thereby further preventing emergencies.

[0123] like Figure 3 As shown, the specific steps are:

[0124] Step 301: Each micro base station records the average energy consumption [θ 1 ,θ 2 ,…,θ N ], update their respective energy consumption thresholds:

[0125] [Θ 1 ,Θ 2 ,…,Θ i ,…,Θ N ]=ω×[θ 1 ,θ 2 ,…,θ N ]

[0126] Θ i is the energy consumption threshold of the i-th micro base station, ω is the artificially set energy consumption threshold coefficient;

[0127] Step 302: Each micro base station uses a fuzzy logic system to calculate and monitor in real time the energy consumption evaluation factors between each user accessing the base station and the base station, as well as the sum of the energy consumption evaluation factors:

[0128] [F 1 ,F 2 ,…,F i ,…,F N ;

[0129] F i is the sum of the energy consumption evaluation factors of all users of the i-th micro base station;

[0130] Step 303: For each micro base station, if the sum of the energy consumption evaluation factors of the micro base station is greater than the energy consumption threshold of the base station, it indicates that the current transmit power configuration of the micro base station is too low; set the micro base station as the micro base station to be adjusted, report it to the macro base station, and enter Step 304; otherwise, do not process;

[0131] When the transmit power configuration of the micro base station is too low, if there are future dynamic changes of users in the coverage area of the base station or other emergencies, it will have a greater impact on the energy consumption and user service quality of the cell.

[0132] Step 304: For the current micro base station to be adjusted n1, the macro base station traverses the users served by the micro base station to be adjusted n1, and finds the user m1 with the largest energy consumption evaluation factor from the current micro base station to be adjusted n1 as the user to be switched;

[0133] Step 305: The macro base station finds the base station k1 with the smallest energy consumption evaluation factor from other base stations except the micro base station to be adjusted n1 for the user to be switched m1 as the candidate target base station;

[0134] Step 306: Use the sum of the energy consumption evaluation factors F k1 of the base station k1, and the energy consumption evaluation factor f m1,k1 between the user to be switched m1 and the base station k1, and judge whether the sum of the two is less than the energy consumption threshold, that is, F k1 +f m1,k1 <Θ k1 , if so, switch the user m1 to the base station k1, and the base station k1 cannot be used as the candidate target base station; otherwise, if no qualified micro base station is found, connect the user to be switched to the macro base station;

[0135] In particular, since the user m1 was not connected to the base station k1 before, when calculating f m1,k1 , the change in the received signal strength within the last T seconds in the input item is set to 0;

[0136] Step 307: The macro base station repeatedly searches for the next micro base station to be adjusted n2, the user to be handed over m2, and the candidate target base station k2, and hands over the user until there is no micro base station to be adjusted.

[0137] Embodiment:

[0138] In order to comprehensively consider the impact of traffic volume and signal transmission quality on energy consumption, this invention evaluates and predicts energy consumption through user-related parameters. The input variables of the fuzzy logic system are user information and base station information, so as to evaluate the impact of each user on the energy consumption of the base station, and the macro base station processes the reported information. Since the relationship between input and output variables is not a simple linear relationship, in the embodiment, the membership function of the fuzzy logic system is selected as a non-linear bell-shaped membership function. The current signal reception situation of the user, the change in received signal strength within the last T seconds, and the transmission requirement of the user in the input variables respectively correspond to the three fuzzy concepts of "low", "medium", and "high", and the transmit power of the base station corresponds to the five fuzzy concepts of "very low", "low", "medium", "high", and "very high".

[0139] An adaptive neural fuzzy system is used to construct the fuzzy logic system, and the fuzzy rule mapping rule is obtained by training the neural network, avoiding the workload and limitations of manually adjusting the parameters of the fuzzy system.

[0140] In order to obtain a more accurate output and better express the calculation result of the output membership function, in the embodiment, the defuzzification method of the fuzzy logic system is selected as the centroid method.

[0141] The adaptive neural fuzzy system is trained using historical data. After training, the fuzzy logic system can output an energy consumption evaluation factor to evaluate the energy consumption situation of the user within the next T seconds if the user accesses the base station. The rules learned from the user-side parameters can better handle the emergencies brought by the randomness of user services and the movement of user locations.

[0142] Combined with the Shannon formula, the energy consumption of the base station for transmitting data to the user:

[0143]

[0144] where D is the amount of data that the user needs to transmit, N 0 is the noise power, P is the transmit power of the base station. The larger P is, the higher the transmission rate and the better the service quality, but it also consumes more energy. L is the link loss between the base station and the user, which is related to the selection of the access base station and also affects the energy consumption and service quality.

[0145] Therefore, in order to find a balance between energy saving and service quality, the energy-saving access algorithm based on deep reinforcement learning includes:

[0146] Construct a deep reinforcement learning network with Proximal Policy Optimization (PPO). PPO is designed based on Actor-Critic and can solve control problems in continuous or large action spaces such as multi-level power control and access selection. In PPO, it is assumed that the policy follows a specific distribution, and the neural network outputs the parameters of this distribution. Then, action decisions are obtained through random sampling based on the obtained parameters. Moreover, PPO uses the same batch of data for multiple gradient updates, tries new policies in each iteration, which can minimize the loss function and ensure a relatively small deviation from the previous iteration's policy, balancing the difficulty of implementation, the complexity of sampling, and the workload required for debugging;

[0147] When the number of users is large, the difficulty of training the deep reinforcement learning network increases significantly. To reduce the training difficulty and comprehensively consider the impact of traffic volume and signal transmission quality on energy consumption, the energy-saving access algorithm based on deep reinforcement learning uses the energy consumption evaluation factor between each user and each small cell at time t as the environmental state, and preprocesses various user and base station parameters, that is

[0148] uses the transmit power level p of each small cell i , and the access base station number β of each user j as the action decision; The base station adopts multi-level power control to save energy, which is more flexible. While taking energy-saving measures on the base station side, it also gives user access strategies to balance the energy consumption distribution in the system and ensure user service quality;

[0149] Uses the weighted sum of the energy consumption and user service performance indicators in the heterogeneous network with macro / small cell coexistence as the reward function, specifically including: the total network transmission energy consumption W at time t 1 , the constant energy consumption W of the base station at time t 2 , the sum W of the energy consumption evaluation factors between all users and all base stations in the network at time t 3 , the service quality reward Q of the user, that is r t =-α 1 W 1 -α 2 W 2 -α 3 W 3 +α 4 Q, where α 1 、α 2 、α 3 、α 4 are normalized weighting coefficients.

[0150] To make a more proactive response to emergencies and avoid the additional impact on cells without emergencies caused by the above energy-saving access algorithm based on deep reinforcement learning, the energy-saving handover algorithm based on fuzzy logic includes:

[0151] Each micro base station records the average power consumption [θ 1 , θ 2 , …, θ N within the past T time. In the embodiment, the energy consumption threshold [Θ 1 , Θ 2 , …, Θ N = 1.2 × [θ 1 , θ 2 , …, θ N ;

[0152] Each micro base station uses a fuzzy logic system to calculate and monitor in real time the energy consumption evaluation factors and the sum of energy consumption evaluation factors [F 1 , F 2 , …, F N between each user accessing the base station and the base station;

[0153] If the sum of the energy consumption evaluation factors of a base station is greater than the energy consumption threshold of the base station, these base stations are defined as micro base stations to be adjusted and reported to the macro base station; the macro base station traverses the energy consumption evaluation factors between the users served by the micro base stations to be adjusted and each base station, and finds the user m with the largest energy consumption evaluation factor among the users served by the current micro base station n to be the user to be handed over;

[0154] The macro base station finds the base station k with the smallest energy consumption evaluation factor between the user to be handed over and other base stations except the micro base station to be adjusted as the candidate target base station;

[0155] If F k + f m,k < Θ k , then the user m is handed over to the base station k, otherwise the base station k can no longer be used as the candidate target base station. In particular, since the user m was not connected to the base station k before, the change in the received signal strength within the past T seconds is set to 0 in the input item when calculating f m,k ;

[0156] The macro base station repeats finding the micro base station n to be adjusted, the user m to be handed over, and the candidate target base station k and performs user handover until there are no micro base stations to be adjusted.

Claims

1. Base station intelligent energy-saving method based on fuzzy logic system and deep reinforcement learning, characterized in that, the specific steps are as follows: First, in a communication scenario with macro base stations and micro base stations, each micro base station reports its user and base station information to the macro base station in real time; Then, the macro base station counts the total number of all users in the scenario and determines whether it changes. If so, the macro base station uses an energy-saving access algorithm based on fuzzy logic and deep reinforcement learning to obtain the energy-saving control strategy of the micro base stations and the access base station selection strategy of the users, and issues them to the micro base stations for execution, taking into account both base station energy saving and user experience; otherwise, each micro base station uses an energy-saving handover algorithm based on fuzzy logic to adjust the user access distribution in real time, so as to further prevent emergencies; The energy-saving access algorithm based on fuzzy logic and deep reinforcement learning, the specific steps are as follows: Step 201, for a target micro base station B, take the information of user u to be connected to this micro base station and the base station information of target micro base station B as variables and input them into the fuzzy logic system; Step 202, set the membership functions corresponding to each fuzzy concept in the fuzzy logic system; The membership degree represents the degree to which the input / output exact value of the fuzzy logic system conforms to this fuzzy concept; the same input / output exact value conforms to multiple fuzzy concepts at the same time; Step 203, for a single input variable S, substitute the exact value of the input variable into each membership function, and output the membership degree of this input variable to each fuzzy concept; Similarly, obtain the membership degrees of each input variable to each fuzzy concept; Step 204, map the membership degrees of each input variable to all fuzzy concepts through fuzzy rules, obtain the membership degrees of the corresponding output variables to each fuzzy concept, and further defuzzify them to convert them into the exact value of the final output variable, which is the energy consumption evaluation factor between user u and target micro base station B; Step 205, at time t, count the set of energy consumption evaluation factors between each user and each micro base station as the state space of the reinforcement learning model; take the transmission power of each micro base station and the access base station number of each user as the action decision; take the weighted sum of the energy consumption and user service performance indicators in the heterogeneous network with coexistence of macro / micro base stations as the reward function; and deploy the trained reinforcement learning model on the macro base station; The state space set is: where M and N are the total number of all users and the total number of all micro base stations in the heterogeneous network, respectively, is the energy consumption evaluation factor between the i-th micro base station and the j-th user; The action decision is: a t = {[p 1 , p 2 , …, p i , …, p N , [β 1 , β 2 , …, β j , …, β M}; p i is the transmission power of the i-th small base station, β j is the access base station number of the j-th user; The reward function is: r t = -α 1 W 1 -α 2 W 2 -α 3 W 3 +α 4 Q where α 1 , α 2 , α 3 , α 4 are normalized weighting coefficients; W 1 is the total network transmission energy consumption at time t, W 2 is the constant energy consumption of the base station at time t, W 3 is the sum of the energy consumption evaluation factors between all users and all base stations in the network at time t; Q is the service quality reward for users; Step 206, the macro base station receives the newly reported energy consumption evaluation factor information of the micro base stations, uses the trained reinforcement learning model to output the power control decisions of all micro base stations and the access decisions of all users, and issues the decisions to each micro base station; Step 207, the micro base station receives the issued decisions, adjusts its own transmission power according to the power control decision, and determines whether the users it serves need to switch to other base stations according to the access decision. If so, issue the corresponding access decision to this user, otherwise do nothing; each user receives the access decision and applies for access to its new target base station; each new target base station feeds back the access request; The specific steps of the energy-saving handover algorithm based on fuzzy logic are: Step 301: Each micro base station records the average power consumption [θ 1 , θ 2 , …, θ N within the past T time, and updates its respective energy consumption threshold: [Θ 1 , Θ 2 , …, Θ i , …, Θ N = ω × [θ 1 , θ 2 , …, θ N ​ Θ i is the energy consumption threshold of the i-th micro base station, and ω is the artificially set energy consumption threshold coefficient; Step 302: Each micro base station uses a fuzzy logic system to calculate and monitor in real time the energy consumption evaluation factors between each user accessing the base station and the base station, as well as the sum of the energy consumption evaluation factors. [F 1 ,F 2 ,…,F i ,…,F N ; F i is the sum of the energy consumption evaluation factors of all users of the i-th micro base station; Step 303: For each micro base station, if the sum of the energy consumption evaluation factors of the micro base station is greater than the energy consumption threshold, it indicates that the current transmit power configuration of the micro base station is too low; set the micro base station as a micro base station to be adjusted, report it to the macro base station, and proceed to Step 304. Otherwise, no action is taken. Step 304: For the current micro base station to be adjusted n1, the macro base station traverses the users served by the micro base station to be adjusted n1, and finds the user m1 with the largest energy consumption evaluation factor among the users served by the current micro base station to be adjusted n1 as the user to be switched. Step 305: The macro base station finds the base station k1 with the smallest energy consumption evaluation factor among the other base stations except the micro base station to be adjusted n1 for the user to be switched m1 as the candidate target base station. Step 306: Use the sum of the energy consumption evaluation factors F of base station k1 k1 , and the energy consumption evaluation factor f of the user m1 to be handed over and base station k1 m1,k1 , and determine whether the sum of the two is less than the energy consumption threshold, that is, F k1 +f m1,k1 <Θ k1 , if so, switch user m1 to base station k1, and base station k1 cannot be used as a candidate target base station; Otherwise, if no eligible micro base station is found, connect the user to be switched to the macro base station. Step 307: The macro base station repeats the process of finding the next micro base station to be adjusted n2, user to be switched m2, candidate target base station k2 and switching the user until there are no micro base stations to be adjusted.

2. The intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning according to claim 1, characterized in that the user information includes: the number of users, the historical signal reception situation of the users, and the user service type; the base station information includes: the power level of the micro base station.

3. The intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning according to claim 1, characterized in that the initial value of the total number of all users in the scenario is: set the initial number of users for each base station according to the actual user service capabilities of each base station.

4. The intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning according to claim 1, characterized in that the membership functions include: linear membership functions are: Triangular membership function Among them, a 1 , b 1 , c 1 respectively represent the abscissas of the three vertices of the triangle; represent the upper and lower limit ranges of the input exact values; Trapezoidal membership function Among them, a 2 , b 2 , c 2 , d 2 respectively represent the abscissas of the four vertices of the trapezoid; represent the upper and lower limit ranges of the input exact values; non-linear membership functions are: Gaussian membership function a 3 represents the mean of the Gaussian distribution and affects the abscissa of the center of the curve; b 3 represents the variance of the Gaussian distribution and affects the width of the curve; Bell-shaped membership function a 4 represents the length affecting the bell shape, b 4 affects the width of the bell shape; c 4 represents the abscissa of the center of the bell shape.

5. The intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning according to claim 1, characterized in that in Step 205, the reward function is composed of 4 parts, and the specific calculation method includes: Where P 0 is the maximum transmission power of the base station, and p i ∈ [0, 1] is the base station power level given by the energy-saving access algorithm based on deep reinforcement learning, and τ i is the longest time for the base station to send data to serve users; Among them p i being 0 indicates that the micro base station is in a dormant state, and p off is the constant power when the base station is in a dormant state, and p on is the constant power when the base station is not in a dormant state; Among them R i is the actual transmission rate of this user, and R 0 is the lowest acceptable transmission rate for the user.

6. The intelligent energy-saving method for base stations based on a fuzzy logic system and deep reinforcement learning according to claim 1, characterized in that In step 306, since user m1 was not connected to base station k1 before, the change in received signal strength within the recent T seconds in the input item is set to 0 when calculating f m1,k1 .

Citation Information

Patent Citations

  • Algorithm for solving power distribution in cognitive radio based on reinforcement learning

    CN112367132A

  • Method of 5g / WLAN vertical handover based on fuzzy logic control

    US20170181052A1