Control device and control method
The control device addresses the inefficiencies in base station sleep control by using relationship information to optimize power consumption and communication quality through deep reinforcement learning, improving learning efficiency and accuracy.
Patent Information
- Application Number
- PCT/JP2024/020418
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-12-11
AI Technical Summary
Existing base station sleep control technologies in wireless networks lack direct input of relationship information between base stations, leading to reduced learning accuracy and prolonged training times, which affects communication quality and power consumption optimization.
A control device that calculates and inputs relationship information, such as graph-based reconnection probabilities between base stations, to optimize power consumption and communication quality using deep reinforcement learning.
Improves learning efficiency and control accuracy by directly incorporating relationship information between base stations, reducing the time required for training and enhancing the effectiveness of power consumption and communication quality management.
Smart Images

Figure JP2024020418_11122025_PF_FP_ABST
Abstract
Description
Control device and control method
[0001] The present invention relates to sleep control for communication devices such as base stations in wireless networks to optimize communication quality and power consumption.
[0002] In recent years, the rapid increase in traffic demand in cellular networks has led to constant demand for improved performance, while at the same time, the impact of environmental loads has led to the need to reduce the power consumption of communication equipment. More than half of the power consumed by communication equipment is consumed by base stations, and in order to reduce power consumption, studies are being conducted to temporarily put less used base station equipment into sleep mode.
[0003] Generally, base stations are designed based on the traffic usage volume during peak hours, and the traffic usage rate is low during non-peak hours. Non-Patent Document 1 discloses a technology that utilizes this property of base stations to reduce power consumption by putting the base station to sleep during non-peak hours.
[0004] On the other hand, excessive sleep time may degrade the communication quality of users using communication services in that area, so it is necessary to set appropriate rules for sleep control.Furthermore, because base stations have various locations and usage situations, it is important to set comprehensive sleep control rules for these diverse base stations.
[0005] In relation to the above, Non-Patent Documents 2, 3, and 4 disclose control that aims to optimize power reduction and communication quality by using a deep reinforcement learning approach that can autonomously learn rules (policies).
[0006] J. Wu, Y. Zhang, M. Zukerman, and E. K. -N. Yung, "Energy-Efficient Base-Stations Sleep-Mode Techniques in Green Cellular Networks: A Survey," in IEEE Communications Surveys & Tutorials, vol. 17, no. 2, pp. 803-826, Secondquarter 2015, doi: 10.1109 / COMST.2015.2403395.Li, Rongpeng, et al. "TACT: A transfer actor-critic learning framework for energy saving in cellular radio access networks." IEEE transactions on wireless communications 13.4 (2014): 2000-2011.Ye, Junhong, and Ying-Jun Angela Zhang. "DRAG: Deep reinforcement learning based base station activation in heterogeneous networks." IEEE Transactions on Mobile Computing 19.9 (2019): 2076-2087.Wu, Qiong, et al. "Deep reinforcement learning with spatio-temporal traffic forecasting for data-driven base station sleep control." IEEE / ACM Transactions on Networking 29.2 (2021): 935-948.
[0007] The inputs in the conventional technologies disclosed in Non-Patent Documents 2 to 4 are the traffic load and utilization rate of the base station to be controlled, and the current state of the base station (active / sleep or on / off), and relationship information representing the relationship between base stations is not used as input. Therefore, the conventional technologies have had the problem of taking time for learning and reducing the accuracy of control. Note that a base station is an example of a communication device to be controlled.
[0008] The present invention has been made in consideration of the above points, and aims to provide a technology that enables control actions to be output using relationship information between communication devices when controlling multiple communication devices that provide network services.
[0009] According to the disclosed technology, there is provided a control device for controlling a plurality of communication devices that provide network services, the control device comprising: a calculation unit that calculates, for each pair of two communication devices where reconnection may occur, a relationship between a certain communication device and another communication device that is a candidate for reconnection of a terminal that was connected to the communication device in response to a predetermined trigger, and the probability of reconnection to the candidate; and a control unit that outputs control behavior based on relationship information output from the calculation unit and observation values obtained from the environment in which the plurality of communication devices are installed.
[0010] The disclosed technology provides a technology that enables output of control actions using relationship information between communication devices in controlling a plurality of communication devices that provide network services.
[0011] It is a block diagram of the control device 100. It is a flowchart showing the operation of the control device 100. It is a diagram for explaining graph information. It is a diagram for explaining graph information. It is a diagram for explaining an example of the hardware configuration of the device.
[0012] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment.
[0013] Hereinafter, an embodiment will be described relating to sleep control of communication devices such as base stations or antennas in a wireless network for reducing power consumption while maintaining communication quality. However, the embodiment described below is merely an example, and the network is not limited to a wireless network, and the controlled object is not limited to a base station or an antenna. For example, the controlled object may be an individual server (an example of a communication device) in a data center that has many servers.
[0014] In the following, first, problems related to the technology according to the present embodiment will be described more specifically, and then the technology according to the present embodiment will be described in detail.
[0015] (Regarding the Issues) As mentioned above, Non-Patent Documents 2, 3, and 4 disclose control that aims to optimize power reduction and communication quality by using a deep reinforcement learning approach that can autonomously learn rules (policies). However, the inputs in the conventional technologies disclosed in Non-Patent Documents 2 to 4 are the traffic load and utilization rate of the base station to be controlled, and the current state of the base station (active / sleep or on / off), and relationship information that represents the relationship between base stations is not used as an input.
[0016] In base station sleep control, when a base station goes into sleep mode, users connected to that base station are reconnected to one of the other active base stations, and the communication quality of the users is affected by the usage status of the reconnection destination base station. Therefore, it is important to understand the usage status of the reconnection destination base station in order to perform control taking communication quality into consideration.
[0017] However, while existing methods input the utilization rate and current state (active / sleep or on / off) of each base station as feature quantities, they do not input information on the relationship between base stations, such as which base stations are candidates for reconnection when a base station is put to sleep. Information on the relationship between base stations refers to, for example, the candidate base stations to be reconnected to when a base station is put to sleep and the respective probability of reconnection.
[0018] In existing methods, the relationships between these base stations are thought to be indirectly reflected in the policy during the learning process, but by not inputting them directly, accuracy is reduced compared to when the relationship information between base stations is input directly. Furthermore, since more training data is required to obtain a model with the same accuracy, learning takes time.
[0019] (Outline of the embodiment) In this embodiment, the control device 100 directly inputs relationship information (described later) relating to the relationship between the asleep base station and the candidate base station to which it will reconnect to the optimization control unit 120, thereby reducing the time required for learning and improving the control accuracy.
[0020] That is, in the base station sleep control executed by the control device 100, in addition to the utilization rate and current state (active / sleep or on / off) of the base station as characteristics of the base station, relationship information between the asleep base station and the candidate base stations to reconnect to is input to the optimization control unit 120.
[0021] More specifically, control is performed by inputting candidate base stations to which a base station can be reconnected when it goes to sleep and the probability of reconnection as graph information in which the base stations are connected as nodes and the candidate reconnection destinations are connected as edges, and the probability of reconnection is represented as the weight of the edges.
[0022] (Device Configuration Example) Fig. 1 shows a configuration example of a control device 100 in this embodiment. As shown in Fig. 1, the control device 100 includes a relationship information calculation unit 110 and an optimization control unit 120. Fig. 1 also shows an environment 200 that is the target of control and observation by the control device 100. Note that the relationship information calculation unit 110 may be called the calculation unit, and the optimization control unit 120 may be called the control unit.
[0023] In this embodiment, the environment 200 is assumed to be a cellular network equipped with one or more base stations.
[0024] The control device 100 in this embodiment performs sleep control for one or more control objects (e.g., base stations or antennas provided in the base stations) to switch between a state in which power consumption can be reduced and an active state. The state in which power consumption can be reduced is, for example, a sleep mode or a state in which the power is turned off. In other words, the control device 100 according to this embodiment performs control to select "active or sleep" or "on or off" for multiple base stations or multiple antennas in a cellular network.
[0025] In this embodiment, the term "base station" may refer to a wireless device consisting of a single housing equipped with an antenna, or may refer to a "configuration equipped with a central station and one or more base stations," or may refer to a base station in a configuration equipped with a central station and one or more base stations, or may refer to an antenna that is a main part of the base station. In other words, "base station" and "antenna" may be synonymous. The controlled object may be a base station, an antenna, or a communication device that is different from either a base station or an antenna. Both a base station and an antenna are examples of communication devices.
[0026] An overview of the operation of the control device 100 will be described with reference to the flowchart shown in Fig. 2. When control starts, in S101 (step 101), the relationship information calculation unit 110 calculates relationship information between devices to be controlled based on device information acquired from the environment 200, and outputs the relationship information. The relationship information is input to the optimization control unit 120.
[0027] In S102, the optimization control unit 120 observes the feature quantities of the control target device from the environment 200 and acquires the observed values as observed values. The observed values include, for example, resource usage rates at base stations / antennas, user (UE) throughput, etc.
[0028] S101 and S102 may be executed in parallel, or may be executed serially, such as S101->S102 or S102->S101.
[0029] When the optimization control unit 120 acquires both the relationship information and the observed values, it calculates and outputs a control action in S103. S101 to S103 are performed, for example, at each time step.
[0030] Although the method for realizing the relationship information calculation unit 110 and the optimization control unit 120 is not limited to a specific method, in this embodiment, it is assumed that the relationship information calculation unit 110 and the optimization control unit 120 are each a neural network model. Also, the optimization control unit 120 may be a neural network model, and the relationship information calculation unit 110 may not be a neural network model.
[0031] The optimization control unit 120 retains the state after the action (e.g., activation / deactivation of each antenna) each time an action is output, and performs policy learning to optimize the policy (measure) based on the state. The model parameters correspond to the policy. More specifically, the policy is learned so that, for example, power consumption is reduced and the communication quality at the UE (terminal) obtained as an observed value satisfies constraints.
[0032] The control device 100 may learn the policy while operating the control over the environment 200, or may operate the control over the environment 200 using the learned policy.
[0033] (Detailed Operation Example) First, the operation of the control device 100 will be described in more detail, assuming that the controlled device is a "base station." As mentioned above, the "base station" may be replaced with an "antenna." The set of base stations is set to B = {1, 2, ..., B}. Also, assume that any two base stations have overlapping coverage areas. Note that there may be two base stations that do not have overlapping coverage areas. A pair of two base stations with overlapping coverage areas may be called a "pair of two base stations where reconnection may occur."
[0034] In this example, the environment 200 (cellular network) shown in Figure 1 includes multiple base stations and multiple users who use the base stations, where the users are referred to as UEs (User Equipment).
[0035] The relationship information calculation unit 110 acquires, as device information, information for configuring relationship information between base stations to be controlled from the environment 200. The relationship information calculation unit 110 generates relationship information between base stations (specifically, graph information) based on the information acquired from the environment 200, and passes the relationship information to the optimization control unit 120. The optimization control unit 120 observes the utilization rate and current state (active / sleep or on / off) of the base station to be controlled as device features, and performs control on the environment 200 to select active / sleep or on / off of multiple base stations to be optimized in terms of power consumption and communication quality based on the observed values obtained from the observation and the relationship information between base stations input from the relationship information calculation unit 110.
[0036] <Details of Relationship Information Calculation Unit 110> An example of the configuration of graph information (assumed to be a directed graph G) calculated by the relationship information calculation unit 110 will be described below.
[0037] The graph information is represented as a directed graph G, with a set of nodes as V, a set of edges as E, and a set of edge weights as W. Each of the multiple base stations in the environment 200 is a node.
[0038] The relationship information calculation unit 110 calculates a set of base stations to which a UE connected to a base station i corresponding to a node i (∈B) may reconnect when the base station i goes to sleep (or is turned off). i and the base station j (∈B i ) to the node corresponding to edge e i→j Set.
[0039] Furthermore, the relationship information calculation unit 110 calculates the reconnection probability p i→j , edge e i→j Here, the probability of reconnection may be set based on the ratio of the overlapping area size of the cover area of base station j to the cover area of base station i, or may be set based on parameters related to handover or load balancing set by the base station, or may be set based on information on reconnection or handover actually measured by the base station.
[0040] B iThe nodes included in are node 2 and node 3, and an image of the edges and weights is shown in FIG.
[0041] For example, if the cover areas of node i, node 2, and node 3 are as shown in FIG. 4, the overlapping area size between the cover area of node i and the cover area of node 2 is larger than the overlapping area size between the cover area of node i and the cover area of node 3. i→2 >p i→3 (For example, p i→2 = 0.8, p i→3 = 0.2).
[0042] The relationship information calculation unit 110 may treat the edges as undirected and the graph information as an undirected graph. i→j and e j→i p set as the weight of i→j and p j→i The weight is set by taking the average value of ij (=p ji ) edge e ij (= e ji )
[0043] For example, in the example of FIG. 4, for the edge between node i and node 2, for example, p i→2 = 0.8, p 2→i = 0.6, then the edge e between node i and node 2 i2 (= e 2i ) is the weight p i2 (=p 2i ) = 0.7 (= (0.8 + 0.6) ÷ 2).
[0044] <Details of Optimization Control Unit 120> Next, a description will be given of the optimization control unit 120. The optimization control unit 120 receives the graph information calculated by the relationship information calculation unit 110 as an input.
[0045] The optimization control unit 120 in this embodiment is basically a neural network model that receives the feature quantities of the base station as input and outputs control actions, similar to existing methods.
[0046] In this embodiment, a graph neural network layer is added to the input layer of the model, so that the optimization control unit 120 (neural network model) receives relationship information between base stations (features having a graph structure) and the features of the base stations as inputs, and outputs control actions.
[0047] Note that using graph neural network layers is just one example. A model that does not use graph neural network layers may also be used. Since relationship information between base stations (features with a graph structure) is used in the input, even if graph neural network layers are not used, it is possible to select actions that reflect the relationships between base stations.
[0048] <Example of Optimization Operation Related to Power Reduction and Communication Quality> An example of optimization operation related to power reduction and communication quality executed by the optimization control unit 110 will be described. Here, the optimization control unit 110 selects whether to activate or shut down a plurality of antennas that provide different frequency bands in a cellular network corresponding to the environment 200 as control targets. One or more UEs (User Equipments) that communicate with the antennas exist under the control of the antennas. The UEs may also be called terminals. "Activation / shutdown" may also be called "on / off" or "active / sleep."
[0049] The antenna set is A all ={1,2,...,A all}, and from the viewpoint of maintaining connectivity, the set of antennas that cannot be stopped is designated as A coverage ={1,2,...,A coverage}, and the set of antennas that can be stopped is A capacity ={1,2,...,A capacity}. This is the antenna set A capacity The coverage area that can provide service is coverage This means that the area is covered by the coverage area of at least one antenna.
[0050] The set of antennas that are not shut down is called A. on (∈A all) and in sleep control, the antenna set A on Determine the following. capacity When the antenna of is stopped, the UE connected to the stopped antenna will on The mobile station will reconnect (handover) to one of the antennas that has a coverage area overlapping with that of the stopped antenna.
[0051] The number of resource blocks (RB) each antenna has is B total Let RB usage rate of antenna i at time t be ρ i Let (t)∈[0,1].
[0052] Antenna set A on The power consumption P of all antennas at time t c is expressed by the following formula:
[0053] Here, P i represents the maximum power consumption of antenna i, and q i ∈(0,1) represents the proportion of the maximum power consumption that is consumed in a fixed manner. However, a different power consumption model may be used.
[0054] Antenna set A all Let N = {1, 2, ..., N} be the set of UEs served by the UE. If all antennas are not stopped (if no antennas are stopped) at the time of traffic demand at time t, the power consumption P a (t) is expressed by the following formula:
[0055] where ^ρ i (t) is the RB usage rate of antenna i at time t when all antennas are not stopped. The power reduction amount P(t) at time t is P(t) = 1 - P c (t) / P a It is expressed as (t).
[0056] Let Q(t) be the communication quality index at time t. Q(t) is a function consisting of one or more quality indexes that should be taken into consideration by a network operator. However, Q(t) may be set so that the larger Q(t) is, the better the communication quality will be, and may be a function expressed as, for example, a weighted linear sum of multiple indexes.
[0057] Here, the quality indicator is set as the average throughput, and the UE n The throughput of T n (t), the average throughput Q(t) at time t can be expressed by the following equation:
[0058] The set quality target is that the throughput of each UE is T target However, other indicators such as delay may be added as quality indicators or may be substituted, and statistics other than the average may be used.
[0059] <Operations Related to Reinforcement Learning> The optimization control unit 120 in the control device 100 outputs control actions using a policy learned by reinforcement learning. Here, an example of the operation of reinforcement learning in this embodiment will be described.
[0060] Action a at time t t The antenna set A that can be stopped is capacity Let us consider a vector that indicates the antenna state, either active or inactive, at time t. For example, if active is 1 and inactive is 0, then the action a t Is A capacity It should be noted that if a certain antenna is in an activated state at time t and is also in an activated state at time t+1, and transitions to the same state, no operation is performed.
[0061] State s at time t t Let be a vector that represents the utilization rate of all antennas and the combination of activation / deactivation of all antennas. Also, let be the reward r t is the power reduction amount P(t) and the average throughput Q(t) with the weight parameter w 1 In addition, for all UEs, T n ≧T target Based on these, the control device 100 determines the action a at time t+1. t+1 Here, the problem to be solved is expressed as the following equation.
[0062] where γ∈(0,1] denotes the discount rate for future rewards, and E s,π [*] represents the expected value for the state and policy. s.t. is the constraint mentioned above, and the reward r t As mentioned above, the power reduction amount P(t) and the average throughput Q(t) are calculated by the weight parameter w 1 This is the sum of the two.
[0063] To solve the above problem, the control device 100 learns a policy π that is a solution to the above problem through reinforcement learning. Calculating the policy π that satisfies the constraints and maximizes the expected value of the reward can be achieved using a general deep reinforcement learning algorithm.
[0064] More specifically, the policy π is obtained as a parameter of an optimized neural network. During operation, the optimization control unit 110 uses the policy π to optimize the behavior a t Output.
[0065] In this embodiment, as described above, the antenna set has a graph structure including a node set V, an edge set E, and an edge weight set W. on When a reconnection (handover) is performed to one of the antennas having a coverage area overlapping with the stopped antenna, the antenna to be reconnected (handover) is selected based on a probability corresponding to the weight of each edge between the source antenna and each of the candidate antennas for reconnection. By using feature quantities corresponding to such content (graph information input from the relationship information calculation unit 110) in searching for a solution, the optimization control unit 110 can directly reflect the relationship between base stations in the policy π. Furthermore, by using this graph information, the time required for learning can be reduced and the accuracy of control can be improved.
[0066] (Hardware Configuration Example) The control device 100 described in this embodiment can be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on a cloud.
[0067] That is, the control device 100 can be realized by using hardware resources such as a CPU and memory built into a computer to execute a program corresponding to the processing performed by the control device 100. The program can be recorded on a computer-readable recording medium (such as a portable memory) and can be saved or distributed. The program can also be provided via a network such as the Internet or email.
[0068] Fig. 5 is a diagram showing an example of the hardware configuration of the computer. The computer in Fig. 5 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, all of which are interconnected via a bus B. The computer may further include a GPU.
[0069] The program that realizes the processing on the computer is provided by a recording medium 1001, such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc.
[0070] The memory device 1003 reads and stores the program from the auxiliary storage device 1002 when an instruction to start the program is received. The CPU 1004 realizes functions related to the control device 100 in accordance with the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, mouse, buttons, a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the results of calculations.
[0071] (Effects of the embodiment) The technology described in the present embodiment improves the accuracy of control by inputting relationship information between communication devices in communication device sleep control that optimizes power consumption and communication quality. Furthermore, in control using deep reinforcement learning, the time required for learning can be reduced and the accuracy of control can be improved.
[0072] The following additional notes are provided regarding the above-described embodiments.
[0073] <Additional Notes> (Additional Item 1) A control device for controlling a plurality of communication devices that provide a network service, comprising: a memory; and at least one processor connected to the memory, wherein the processor calculates, for each pair of two communication devices between which reconnection may occur, a relationship between a certain communication device and other communication devices that are candidates for reconnection of a terminal that was connected to the certain communication device in response to a predetermined trigger, and a probability of reconnection to the candidate, and outputs a control action based on the calculated relationship information and an observation value acquired from an environment in which the plurality of communication devices are installed. (Additional Item 2) The control device according to Additional Item 1, wherein the predetermined trigger is a communication device being switched to a sleep state or a communication device being powered off, and the processor outputs a control action that aims to optimize power consumption reduction and average communication quality based on the relationship information and the observation value. (Supplementary Item 3) The control device according to Supplementary Item 1, wherein the processor calculates a relationship between the communication devices as a graph structure by establishing an edge between a node that is a communication device and another node that is a candidate for reconnection and setting a weight of the edge as a probability of reconnecting to the other node. (Supplementary Item 4) A control method executed by a control device for controlling a plurality of communication devices that provide network services, the control method comprising: a calculation step of calculating, for each pair of two communication devices between which reconnection may occur among the plurality of communication devices, a relationship between a certain communication device and another communication device that is a candidate for reconnection of a terminal that was connected to the certain communication device in response to a predetermined trigger, and a probability of reconnecting to the candidate, and a control step of outputting a control action based on relationship information calculated by the calculation step and observation values acquired from an environment in which the plurality of communication devices are installed.
[0074] Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
[0075] REFERENCE SIGNS LIST 100 Control device 110 Relationship information calculation unit 120 Optimization control unit 200 Environment 1000 Drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device
Claims
1. A control device for controlling a plurality of communication devices that provide network services, comprising: a calculation unit that calculates, for each pair of two communication devices among the plurality of communication devices, a relationship between a certain communication device, other communication devices that are candidates for reconnection of a terminal that was connected to the communication device in response to a predetermined trigger, and the probability of reconnection to the candidate communication devices; and a control unit that outputs control actions based on relationship information output from the calculation unit and observation values obtained from the environment in which the plurality of communication devices are installed.
2. The control device described in claim 1, wherein the specified trigger is that the communication device is switched to a sleep state or that the power of the communication device is turned off, and the control unit outputs control actions that aim to optimize the amount of power consumption reduction and average communication quality based on the relationship information and the observed value.
3. The control device according to claim 1, wherein the calculation unit calculates the relationship between the communication devices as a graph structure by establishing an edge between the node that is the communication device and another node that is a candidate for reconnection, and setting the weight of the edge as the probability of reconnecting to the other node.
4. A control method executed by a control device for controlling a plurality of communication devices that provide network services, comprising: a calculation step of calculating, for each pair of two communication devices among the plurality of communication devices, a relationship between a certain communication device and another communication device that is a candidate for reconnection of a terminal that was connected to the communication device in response to a predetermined trigger, and the probability of reconnection to the candidate; and a control step of outputting control actions based on the relationship information calculated by the calculation step and observation values obtained from the environment in which the plurality of communication devices are installed.
Citation Information
Patent Citations
Radio communication system, radio base station device, radio base station controller, and method for controlling power consumption
JP2014154997A
Base station, control device, communication system, method, and program
WO2024034524A1