Setting device and parameter updating method
The setting device employs a two-stage reinforcement learning process to optimize antenna activation/deactivation, addressing the issue of excessive sleeping in base stations, ensuring reduced power consumption without degrading communication quality.
Patent Information
- Application Number
- PCT/JP2024/027080
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional methods for reducing power consumption in base stations by putting them into sleep mode during non-peak hours can lead to excessive sleeping, degrading communication quality, and existing deep reinforcement learning approaches fail to effectively manage communication quality constraints.
A setting device that utilizes a two-stage reinforcement learning process to dynamically adjust weight parameters, balancing power consumption reduction and communication quality constraints by optimizing antenna activation/deactivation strategies in cellular networks.
Efficiently reduces power consumption while maintaining communication quality by dynamically adjusting weight parameters to satisfy predefined constraints, thus enhancing the accuracy and speed of power management.
Smart Images

Figure JP2024027080_05022026_PF_FP_ABST
Abstract
Description
Setting device and parameter update method
[0001] The present invention relates to techniques for reducing power consumption in networks.
[0002] In recent years, the rapid increase in traffic demand in cellular networks has led to constant demand for improved performance, while at the same time, the impact of environmental loads has led to the need to reduce the power consumption of communication equipment. More than half of the power consumed by communication equipment is consumed by base stations, and in order to reduce power consumption, it has been considered to temporarily put less used base stations into sleep mode.
[0003] Generally, base stations are designed based on the traffic usage volume during peak hours, and the traffic usage rate is low during non-peak hours. Non-Patent Document 1 discloses a technology that utilizes this property of base stations to reduce power consumption by putting the base station to sleep during non-peak hours.
[0004] J. Wu, Y. Zhang, M. Zukerman, and EK -N. Yung, "Energy-Efficient Base-Stations Sleep-Mode Techniques in Green Cellular Networks: A Survey," in IEEE Communications Surveys & Tutorials, vol. 17, no. 2, pp. 803-826, Secondquarter 2015, doi: 10.1109 / COMST.2015.2403395.
[0005] However, conventional technology that puts base stations to sleep can cause them to sleep excessively, which can degrade the communication quality of users who use communication services in the base station's area.
[0006] The present invention has been made in consideration of the above points, and aims to provide a technology that enables control over an environment in which network services are provided, while achieving power consumption reduction and satisfying communication quality constraints.
[0007] According to the disclosed technology, there is provided a setting device that sets parameters used to learn a model for controlling an environment having a plurality of control objects that provide network services, the setting device comprising: a learning unit that learns the model using the control action output by the model, observation values obtained from the environment, and the parameters; an evaluation unit that evaluates the model using the observation values obtained from the environment as a result of the control action output by the model learned by the learning unit; and a parameter update unit that updates the parameters based on the evaluation results by the evaluation unit.
[0008] The disclosed technology provides a technology that enables control of an environment in which network services are provided, while achieving reduction in power consumption and satisfying communication quality constraints.
[0009] 1 is a configuration diagram of a control device 100. FIG. 2 is a flowchart for explaining the operation of the control device 100. FIG. 3 is a configuration diagram of a setting device 300. FIG. 4 is a flowchart for explaining the operation of the setting device 300. FIG. 5 is a diagram showing an example of a control hardware configuration.
[0010] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment.
[0011] Hereinafter, an embodiment relating to sleep control of communication devices such as base stations in a wireless network for reducing power consumption while maintaining communication quality will be described. However, the embodiment described below is merely an example, and the network is not limited to a wireless network, and the controlled object is not limited to a base station or an antenna. For example, the controlled object may be an individual server in a data center that has many servers.
[0012] In the following, first, the conventional technology and its problems will be specifically described, and then the basic technology that is the premise of the technology according to the present embodiment will be described. After the basic technology, the problems of the basic technology and the technology according to the present embodiment will be described. Note that the technology disclosed in the documents mentioned below is publicly known, but the explanation of the problems is not publicly known.
[0013] (Regarding the prior art and its problems) As mentioned above, the technology of Non-Patent Document 1 may cause the base station to sleep excessively, which may degrade the communication quality of users who use communication services in that area.
[0014] In response to this, Li, Rongpeng, et al. "TACT: A transfer actor-critic learning framework for energy saving in cellular radio access networks," IEEE transactions on wireless communications 13.4 (2014): 2000-2011, discloses technology that aims not only to reduce power consumption but also to minimize degradation in communication quality. Furthermore, the paper uses a reinforcement learning approach that can learn from temporal fluctuations in traffic, including these fluctuations, and resolves the long learning time required in deep reinforcement learning by using a transfer learning approach.
[0015] Also, "Ye, Junhong, and Ying-Jun Angela Zhang. "DRAG: Deep reinforcement learning based base station activation in heterogeneous networks." IEEE Transactions on Mobile Computing 19.9 (2019): 2076-2087." Wu, Qiong, et al. "Deep reinforcement learning with spatio-temporal traffic forecasting for data-driven base station sleep control." IEEE / ACM Transactions on Networking 29.2 (2021): 935-948. "Li, Rongpeng, et al. "TACT: A transfer actor-critic learning framework for energy saving in cellular radio access networks." IEEE transactions on wireless communications 13.4 (2014): As with "2000-2011," a deep reinforcement learning approach is used to reduce power consumption and minimize degradation of communication quality. By adding a prediction mechanism for traffic fluctuations, it is possible to improve accuracy in response to fluctuations, demonstrating that dynamic control with high accuracy is possible even compared to conventional rule-based methods.
[0016] Generally, network operators set a target value for communication quality when providing communication services. Therefore, it is desirable to reduce power consumption while achieving this target value. In the deep reinforcement learning approach used in the above-mentioned conventional technology, although the objective function includes a term that minimizes degradation of communication quality, it cannot satisfy the constraint that the quality does not fall below the target value, and therefore cannot be used as is.
[0017] While it is desirable to reduce power consumption while satisfying the constraint of achieving a preset communication quality target, the deep reinforcement learning approach used in conventional technology does not provide a means for inputting a constraint equation. As a result, while the desired power reduction and average quality are maximized, an optimal search cannot be performed to satisfy the constraint, which may result in a degradation of the user's communication quality.
[0018] (Summary of Basic Technology) In the basic technology, the control device 100 applies deep reinforcement learning to minimize the cost and constraint violations regarding the similarity with the proposed control action, based on the results of the action (proposed control action) that maximizes the amount of power reduction and average communication quality output by the previous control. Through this processing, sleep control of the control target that reduces power consumption while satisfying the target value of communication quality is realized.
[0019] (Device Configuration Example in Basic Technology) Fig. 1 shows a configuration example of a control device 100 in the basic technology. As shown in Fig. 1, the control device 100 includes an objective optimization control unit 110 and a constraint optimization control unit 120. Fig. 1 also shows an environment 200 that is the target of control and observation by the control device 100.
[0020] In this embodiment, it is assumed that the environment 200 is a cellular network equipped with one or more base stations. Note that the objective optimization control unit 110 may be provided outside the control device 100.
[0021] The control device 100 in this embodiment performs sleep control for switching one or more control targets (e.g., base stations or antennas provided in the base stations) between a state in which power consumption can be reduced and an active state. The state in which power consumption can be reduced is, for example, a sleep mode or a state in which the power is turned off.
[0022] The operation of the control device 100 will be outlined with reference to the flowchart shown in FIG.
[0023] In S1 (step 1), the objective optimization control unit 110 calculates and outputs a control action plan that maximizes the amount of power consumption reduction and the average communication quality based on the observed values obtained from the environment 200.
[0024] In S2, the constraint optimization control unit 120 calculates and outputs a control action that minimizes the cost related to the similarity with the control action plan and the violation of the constraint that the communication quality achieves the target value, based on the control action plan output from the objective optimization control unit 110 and the observation value obtained from the environment 200 (e.g., communication quality at the UE).
[0025] The method for realizing the objective optimization control unit 110 and the constraint optimization control unit 120 is not limited to a specific method, but in this embodiment, it is assumed that the objective optimization control unit 110 and the constraint optimization control unit 120 are each a neural network model.
[0026] It is also assumed that the parameters of the above model are updated (optimized) by reinforcement learning. For example, each time an action is taken, the objective optimization control unit 110 and the constraint optimization control unit 120 each retain the state after the action (e.g., activation / deactivation of each antenna), and learn the policy so as to optimize the policy (measure) based on the state. The model parameters correspond to the policy. More specifically, the policy is learned so as to reduce power consumption and ensure that the communication quality at the UE obtained as an observation value satisfies the constraints.
[0027] The control device 100 may learn the policy while operating the control over the environment 200, or may operate the control over the environment 200 using the learned policy.
[0028] The operation of the control device 100 will be described in more detail below. In the following description, "A / B" means "A or B." In addition, in the text of the specification, for convenience of description, normal font is used for characters that represent a set, such as "A," but it is clear from the context that this means a set. In addition, in the text of the specification, a hat (^) intended to be written above a character is written before the character. "^ρ i (t)" is an example.
[0029] (Formulation of Basic Technology) Here, we consider the problem of selecting whether to activate or deactivate multiple antennas that provide different frequency bands and are installed at a single base station in a cellular network corresponding to environment 200. Under the control of the base station, there are one or more UEs (User Equipments) that communicate with the base station. The UEs may also be called terminals.
[0030] The antenna set is A all ={1,2,...,A all}, and from the viewpoint of maintaining connectivity, the set of antennas that cannot be stopped is designated as A coverage ={1,2,...,A coverage}, and the set of antennas that can be stopped is A capacity ={1,2,...,A capacity}. This is the antenna set A capacity The coverage area that can provide service is coverage This means that the area is covered by the coverage area of at least one antenna.
[0031] The set of antennas that are not shut down is called A. on (∈A all ) and in sleep control, the antenna set A on Determine the following. capacity When the antenna of is stopped, the UE connected to the stopped antenna will on The handover is performed to one of the antennas that has a coverage area overlapping with that of the stopped antenna.
[0032] The number of resource blocks (RB) each antenna has is B total Let RB usage rate of antenna i at time t be ρ i Let (t)∈[0,1].
[0033] Power consumption P of all antennas at time t c is expressed by the following formula:
[0034] Here, P i represents the maximum power consumption of antenna i, and q i∈(0,1) represents the proportion of the maximum power consumption that is consumed in a fixed manner. However, a different power consumption model may be used.
[0035] Antenna set A all Let N = {1, 2, ..., N} be the set of UEs served by the UE. If all antennas are not stopped (if no antennas are stopped) at the time of traffic demand at time t, the power consumption P a (t) is expressed by the following formula:
[0036] where ^ρ i (t) is the RB usage rate of antenna i at time t when all antennas are not stopped. The power reduction amount P(t) at time t is P(t) = 1 - P c (t) / P a It is expressed as (t).
[0037] Let Q(t) be the communication quality index at time t. Q(t) is a function consisting of one or more quality indexes that should be taken into consideration by a network operator. However, Q(t) may be set so that the larger Q(t) is, the better the communication quality will be, and may be a function expressed as, for example, a weighted linear sum of multiple indexes.
[0038] Here, the quality index is set as the average throughput, and the throughput of the UE at time t is T n (t), the average throughput Q(t) at time t can be expressed by the following equation:
[0039] The set quality target is that the throughput of each UE is T target However, other indicators such as delay may be added as quality indicators or may be substituted, and statistics other than the average may be used.
[0040] (Operations Related to Reinforcement Learning) The objective optimization control unit 110 and the constraint optimization control unit 120 in the control device 100 output a control action plan and a control action, respectively, using a policy learned by reinforcement learning. Here, the design of reinforcement learning in this embodiment will be described.
[0041] Action a at time tt The antenna set A that can be stopped is capacity Let us consider a vector that indicates the antenna state, either active or inactive, at time t. For example, if active is 1 and inactive is 0, then the action a t A capacity It should be noted that if a certain antenna is in an activated state at time t and is also in an activated state at time t+1, and transitions to the same state, no operation is performed.
[0042] State s at time t t Let be a vector that represents the utilization rate of all antennas and the combination of activation / deactivation of all antennas. Also, let be the reward r t is the power reduction amount P(t) and the average throughput Q(t) with the weight parameter w 1 In addition, all UEs n In this case, T n ≧T target Based on these, the control device 100 determines the action a at time t+1. t+1 Here, the problem to be solved is expressed as the following equation.
[0043] where γ∈(0,1] denotes the discount rate for future rewards, and E s,π [*] represents the expected value for the state and policy. s.t. is the constraint mentioned above, and the reward r t As mentioned above, the power reduction amount P(t) and the average throughput Q(t) are calculated by the weight parameter w 1 This is the sum of the two.
[0044] For the above problem, the control device 100 applies two-stage reinforcement learning to search for a solution that satisfies the constraints. First, the objective optimization control unit 110 generates a policy π that maximizes only the following objective: o Learn.
[0045] Next, the constraint optimization control unit 120 calculates the policy π o Action a output by t Let,be the action ^a that satisfies the constraints.t A policy π that converts c The policy π o Action a output by t corresponds to the "control action plan" shown in Figure 1, and the policy π c The action ^a output by t corresponds to the "control action" shown in Figure 1.
[0046] In this embodiment, policy π c We use two cost functions to learn the first cost function, which is the cost of the action a t and action ^a t Cost function C at time t for the similarity between s (t). For example, a distance such as the l-1 norm, the l-2 norm, or the KL divergence may be used as the cost for the similarity, or a nonlinear function for the distance may be used. As the nonlinear function, a function that assigns a larger cost when two actions are farther apart may be used.
[0047] In this embodiment, the preceding policy π o The action output by C is intended to maximize the expected value of the reward, which is the goal. However, if there is a constraint violation, the optimal point that eliminates the constraint violation is found by adding the minimum number of activated antennas. s (t) is the previous policy π o Action a output by t Number of activated antennas in |A on For |, the cost is designed as a convex function so that if the change in the number of activated antennas is small, the cost is small, and if the change is large, the cost is large. For example, C s (t) is defined as follows:
[0048] Here, ||*|| 1 represents the l-1 norm, and |A capacity | represents the number of elements in the set. The -1 in the numerator and the -1 in the denominator are s It is used to normalize (t) to 0-1.
[0049] Next, the second cost function will be described. The second cost function is a function that represents the cost related to the constraint violation at time t, and is C v (t). For example, C v As (t), the number of UEs violating the constraint or the amount of violated throughput may be used, or a nonlinear function that increases according to the number of violations may be used. Here, since a strict constraint is imposed on the number of UEs violating the constraint when the number of UEs violating the constraint is 1 or more, a function with a steep slope when the number of UEs violating the constraint changes from 0 to 1 is used. For example, as shown below, using the tanh function, C v Define (t).
[0050] Here, I(*) is an indicator function that takes the value 1 if the condition in the parentheses is met and takes the value 0 if it is not met, and α represents a parameter of the tanh function.
[0051] The two cost functions described above are weighted by the weight parameter w 2 The cost c at time t is calculated by adding t is defined as follows:
[0052] The constraint optimization control unit 120 t Using the following problem, we can solve the policy π c In other words, we learn a policy π c Learn.
[0053] (Issues of the Basic Technology) As described above, in the basic technology, the control device 100 performs two-stage control using a reinforcement learning approach. That is, the objective optimization control unit 110 in the first stage outputs action plans that maximize the amount of power consumption reduction and the average communication quality, and the constraint optimization control unit 120 performs control aimed at reducing constraint violations and selecting actions with high similarity from the action plans output in the first stage.
[0054] Here, the control performed in the second stage deals with two indices: the cost related to constraint violation (constraint violation cost) and the similarity between the input action plan and the output action. Optimization is performed by minimizing the weighted sum of these indices using a parameter representing the weight (weight parameter).
[0055] As described above, in the second-stage reinforcement learning of the basic technology, the objective to be optimized (minimized) is expressed as a weighted sum of the constraint violation cost and the similarity between the input action plan and the output action.
[0056] In use cases where the basic technology is applied, it is more important to eliminate constraint violations, and it is necessary to reduce the cost of constraint violations and output actions that are close to the action plans output by the first-stage reinforcement learning that are highly effective in reducing power consumption and improving average quality.
[0057] In this case, it is necessary to set large weight parameters for terms that minimize the constraint violation cost, but it is not clear which weight parameters can achieve the target constraint violation cost. In order to search for appropriate weight parameters, control agents (specifically, the objective optimization control unit 110 and the constraint optimization control unit 120) that have been trained with various parameters are prepared, and the performance of each control agent is evaluated to confirm whether the target constraint violation cost is achieved, and the weight parameters are determined. Therefore, it is necessary to train with various weight parameters.
[0058] Generally, reinforcement learning requires a certain amount of time to learn because it requires a certain amount of diversity depending on the dimensions of the state and the action. Therefore, learning a control agent with various weight parameters requires a certain amount of time, as the learning time is equal to the number of weight parameters to be explored.
[0059] (Overview of the Technology According to the Present Embodiment) In this embodiment, a technology will be described that solves the above-mentioned problems in the basic technology, efficiently searches for weight parameters of the objective function of reinforcement learning in the basic technology, and shortens the search time. The technology according to this embodiment is outlined below. Note that the "weight parameters" may also be called "parameters."
[0060] In this embodiment, the setting device 300 (described later) calculates the cost (C v (t)) and a cost function (C s (t)) and the objective function (c t ) weight parameters are explored.
[0061] Specifically, the setting device 300 sets up an evaluation phase at a predetermined arbitrary learning interval, and based on the evaluation results of that evaluation phase, dynamically changes the weight parameters in response to the deviation between the evaluated constraint violation and the target constraint violation cost.
[0062] For example, if the evaluated constraint violation does not achieve the target constraint violation cost, the setting device 300 updates the weight parameters to increase the weight of the constraint violation cost. By this means, weight parameters that can achieve the target constraint violation cost are searched for during learning.
[0063] As described above, by searching for weight parameters during the learning process, it is possible to shorten the search time without having to learn individual control agents according to the number of weight parameters to be searched. The operation and configuration of the setting device 300 in this embodiment will be described in detail below.
[0064] (Regarding the Operation of the Setting Device 300 in the Present Embodiment) The formulation that forms the basis of the operation of the setting device 300 is the same as that of the basic technology. Furthermore, the setting device 300 acquires observation results from the environment 200, the same as in the basic technology, and performs control actions on the environment 200.
[0065] The setting device 300 includes an objective optimization control unit 110 and a constraint optimization control unit 120 in the basic technology. Hereinafter, the objective optimization control unit 110 will be referred to as the "first-stage control agent," and the constraint optimization control unit 120 will be referred to as the "second-stage control agent." Both the "first-stage control agent" and the "second-stage control agent" are, for example, neural network models.
[0066] As in the basic technology, in this embodiment, the problem of selecting activation / deactivation of antennas that provide different frequency bands and are installed at one base station in a cellular network is considered.
[0067] As with the basic technology, the antenna set is A all ={1,2,...,A all}, and the set of antennas that are not stopped is A on (∈A all ) and in sleep control, the antenna set A on Determine.
[0068] The control agent in the first stage proposes an action plan a that aims to optimize the reduction of power consumption and the improvement of average quality. t The second-stage control agent outputs this proposed action as an action ^a that satisfies the constraints. t policy π for converting c In other words, the second-stage control agent learns the action ^a t Policy π for outputting c Learn.
[0069] As explained in the basic technology, the cost for a constraint violation is C v (t), and the action plan a is used to find an action plan that is as close as possible to the action plan of the first-stage control agent. t and action ^a t The cost representing the similarity of s Let (t) be C. v (t) and C s Specific examples of (t) are as explained in the basic technology. v (t) and C s Each of (t) is not limited to the cost functions explained in the basic technology.
[0070] The two cost functions described above are summed using weight parameters to obtain the cost c at time t. t In this embodiment, the w 2 1-w 1 Then, 1-w 2 Wow 1Define it again as follows: t Represents.
[0071] The setting device 300 calculates the policy π for the above costs by the following minimization problem, as in the basic technology. c Learn.
[0072] where T is the total time the control is performed, and γ c ∈(0,1] is the discount rate for future rewards, and E s,π_c [*] represents expected values for states and policies.
[0073] In the above problem, the setting device 300 determines the weight parameter w that achieves the target constraint violation cost. 1 This section explains how to explore the
[0074] The setting device 300 sets the policy π c In the learning process, after K episodes of learning are completed, the constraint violation is evaluated by M samplings to determine the weight parameter w 1 where K and M are integers of 1 or greater.
[0075] The setting device 300 further performs learning for K episodes and performs evaluation again by sampling M times. In the sampling, observed values are acquired from the environment 200. The observed values include, for example, the throughput of each UE. Note that the sampling may also be called a trial.
[0076] The setting device 300 calculates this trial using the weight parameter w 1 This is repeated until convergence occurs. Here, the value of the number of times of learning K may be a predetermined value (a fixed value) or may be dynamically changed. For example, K is set to a large value at the start of learning, and then gradually decreased.
[0077] An example of a procedure in which the setting device 300 evaluates constraint violations and updates parameters by sampling M times will be described below.
[0078] Here, as with the basic technology, antenna set A allLet the set of UEs being served be N = {1, 2, ..., N}, and let the throughput of the UE at time t be T n The constraint to be satisfied is that the throughput of all UEs is T target That is all.
[0079] After K learning cycles have been completed, the setting device 300 evaluates the violation status. Here, the evaluation function may be the maximum number of violating UEs among M samplings (observation values), the number of trials (number of samplings) in which one or more violating UEs occurred, or the proportion of trials (proportion among M trials) in which one or more violating UEs occurred. This may be set appropriately based on the target degree to which violations are to be reduced.
[0080] For example, when the target constraint violation cost is set to a cost indicating that "the ratio of trials in which one or more UEs violated the constraint should be 0.1% or less," the setting device 300 sets the weight parameter w as follows: 1 Update.
[0081] Among M attempts, the number of violating UEs in the mth attempt is N m The setting device 300 updates the weight parameter w 1 (j) is updated at the j+1th time as follows:
[0082] Here, α represents the update rate, and may be set to a predetermined constant value, or may be set to a value that decays with respect to the number of updates j. I(*) is an indicator function that takes 1 if the condition in the parentheses is met, and 0 if it is not met. For example, in the above formula, if M=1000, and N out of M times is m If the number of times that α≧1 is 2, the value in the parentheses for α in the above formula is 0.001. If α=1, then in the j+1th update, the weight parameter w 1 increases by 0.001 from the jth time. v The weight of (t) increases by 0.001, so C v Learning is performed to significantly lower (t) by the amount of the increased weight.
[0083] The setting device 300 repeatedly performs the above update until it approaches the target, thereby determining the weight parameters w appropriate for the problem. 1 get.
[0084] (Device Configuration, Device Operation) An example configuration of a setting device 300 that executes the above processing is shown in Fig. 3. As shown in Fig. 3, the setting device 300 includes a learning unit 310, an evaluation unit 320, and a parameter update unit 330. As shown in Fig. 3, there is an environment 200 described in the basic technology, and the learning unit 310 and the evaluation unit 320 each perform control actions on the environment 200 and can obtain observation values from the environment 200.
[0085] The "learning unit 310, evaluation unit 320, and parameter update unit 330" may be provided in one device, or the learning unit 310, evaluation unit 320, and parameter update unit 330 may each be provided in separate devices. Even when the learning unit 310, evaluation unit 320, and parameter update unit 330 are each provided in separate devices, the configuration having the "learning unit 310, evaluation unit 320, and parameter update unit 330" may be called setting device 300.
[0086] The operation of the setting device 300 shown in FIG. 3 will be described along the procedure of the processing flow shown in FIG.
[0087] <S11: Learning> The learning unit 310 has an "objective optimization control unit 110 and a constraint optimization control unit 120" (i.e., a first-stage control agent and a second-stage control agent). Each of the first-stage control agent and the second-stage control agent may be called a model.
[0088] In S11, the learning unit 310 learns the model through K episodes. That is, as explained in the basic technology, the policy π for outputting the control action is c The learning unit 310 passes the learned model to the evaluation unit 320.
[0089] The trained model here is a model that calculates a policy π based on the current state of the environment 200. c This model outputs an action according to the policy πc corresponds to the model parameters of the model.
[0090] <S12: Evaluation by Sampling> In S12, the evaluation unit 320 evaluates the model by sampling. Specifically, the evaluation is performed as follows.
[0091] The evaluation unit 320 has a trained model trained by the training unit 310. The trained model receives the current state based on the previous control action as an input and calculates the action ^a t The evaluation unit 320 outputs the action ^a t The control action is performed on the environment 200 according to the above. The evaluation unit 320 also obtains observed values from the environment 200 having a state that is a result of the execution of the control action. The observed values include the throughput of each UE.
[0092] "Executing a control action and obtaining observed values" is called sampling, and the evaluation unit 320 performs this sampling M times. The evaluation unit 320 evaluates constraint violations based on the M samplings. For example, the evaluation unit 320 calculates the second term on the right-hand side of equation 10 and passes the calculation result to the parameter update unit 330 as the evaluation result. The parameter update unit 330 receives the evaluation result.
[0093] <S13: Updating Weighting Parameters> In S13, the parameter updating unit 330 updates the weighting parameters. Specifically, the update is performed as follows.
[0094] The parameter update unit 330 updates the current weight parameter w 1 from the learning unit 310. The parameter updating unit 330 updates the current weight parameter w based on the evaluation result received from the evaluation unit 320. 1 For example, the parameter update unit 330 updates the weight parameter w by performing the calculation of "Equation 10." 1 The parameter update unit 330 updates the updated weight parameter w 1 is passed to the learning unit 310.
[0095] <S14: Convergence determination> The processes of S11 to S13 are performed to determine the weight parameter w 1If the parameter update unit 330 determines that the weight parameter update has converged, the weight parameter setting process ends, and if not, the process returns to S11.
[0096] Any method may be used to determine convergence. For example, the weight parameter w 1 It may be determined that convergence has occurred when the amount of change due to the update of becomes smaller than a threshold value. Alternatively, it may be determined that convergence has occurred and the process ends when the processes of S11 to S13 have been repeated a predetermined number of times.
[0097] (Hardware Configuration Example) Any of the devices (control device 100, setting device 300) described in this embodiment can be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud.
[0098] That is, the device can be realized by executing a program corresponding to the processing performed by the device using hardware resources such as a CPU and memory built into a computer. The program can be recorded on a computer-readable recording medium (such as a portable memory) and stored or distributed. The program can also be provided via a network such as the Internet or email.
[0099] Fig. 5 is a diagram showing an example of the hardware configuration of the computer. The computer in Fig. 5 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, all of which are interconnected via a bus B. The computer may further include a GPU.
[0100] The program that realizes the processing on the computer is provided by a recording medium 1001, such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc.
[0101] The memory device 1003 reads and stores a program from the auxiliary storage device 1002 when an instruction to start the program is received. The CPU 1004 realizes functions related to the device in accordance with the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, mouse, buttons, a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the results of calculations.
[0102] As described above, in the technology according to the present embodiment, weight parameters are searched for during the learning process, thereby reducing the search time without having to learn individual control agents according to the number of weight parameters to be searched. As a result, it becomes possible to more efficiently realize control that satisfies communication quality constraints while achieving reduced power consumption.
[0103] The following additional notes are provided regarding the above-described embodiments.
[0104] <Supplementary Notes> (Supplementary Item 1) A setting device that sets parameters used to train a model for performing control on an environment having a plurality of control objects that provide network services, comprising: a memory; and at least one processor connected to the memory, wherein the processor: trains the model using a control action output by the model, an observation value obtained from the environment, and the parameters; evaluates the model using the observation value obtained from the environment as a result of the control action output by the trained model; and updates the parameters based on the evaluation results. (Supplementary Item 2) The setting device according to Supplementary Item 1, wherein the processor trains the model so as to minimize a weighted sum, using the parameters, of a cost related to a violation of a communication quality constraint, calculated based on the observation value, and a cost related to the similarity of the control action to a control action proposal that maximizes power consumption reduction and average communication quality. (Supplementary Item 3) The setting device according to Supplementary Item 1, wherein the processor evaluates the model based on a state of violation of the communication quality constraint, calculated based on the observation value. (Addendum 4) A parameter updating method executed by a setting device that sets parameters used to learn a model for controlling an environment having a plurality of control objects that provide network services, the parameter updating method comprising: a learning step of learning the model using a control action output by the model, an observation value obtained from the environment, and the parameters; an evaluation step of evaluating the model using the observation value obtained from the environment as a result of the control action output by the model learned by the learning step; and a parameter updating step of updating the parameters based on the evaluation result by the evaluation step.
[0105] Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
[0106] REFERENCE SIGNS LIST 100 Control device 110 Objective optimization control unit 120 Constraint optimization control unit 200 Environment 300 Setting device 310 Learning unit 320 Evaluation unit 330 Parameter update unit 1000 Drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device
Claims
A setting device for setting parameters used for learning a model for performing control on an environment including a plurality of control objects that provide network services, a learning unit that learns the model using a control action output by the model, an observation value obtained from the environment, and the parameters; an evaluation unit that evaluates the model using observed values obtained from the environment as a result of a control action output by the model learned by the learning unit; a parameter update unit that updates the parameters based on the evaluation result by the evaluation unit; A setting device comprising: The learning unit learns the model so as to minimize a weighted sum using the parameters of a cost related to a violation of a constraint on communication quality, which is calculated based on the observed values, and a cost related to the similarity of the control action to a control action plan that maximizes the amount of power consumption reduction and the average communication quality. The setting device according to claim 1 . The evaluation unit evaluates the model based on a state of violation of a constraint on communication quality calculated based on the observed value. The setting device according to claim 1 . A parameter update method executed by a setting device that sets parameters used for learning a model for performing control over an environment including a plurality of control objects that provide a network service, the method comprising: a learning step of learning the model using a control action output by the model, an observation value obtained from the environment, and the parameters; an evaluation step of evaluating the model using observation values obtained from the environment as a result of control actions output by the model learned in the learning step; a parameter updating step of updating the parameters based on the evaluation result obtained by the evaluation step; A parameter update method comprising:
Citation Information
Patent Citations
Tilt angle optimizing device, tilt angle optimizing method, and program
WO2023157198A1