Electronic device, control method, and program
Patent Information
- Application Number
- PCT/JP2025/007559
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-05
- Filing Date
- 2025-03-03
- Publication Date
- 2025-10-02
Smart Images

Figure JP2025007559_02102025_PF_FP_ABST
Abstract
Description
Electronic device, control method, and program CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to Japanese Patent Application No. 2024-33185, filed on March 5, 2024, the entire disclosure of which is incorporated herein by reference.
[0002] The present disclosure relates to an electronic device, a control method, and a program.
[0003] In a wireless communication device equipped with a phased array antenna including multiple antenna elements, a beamforming (hereinafter also simply referred to as "BF") technique is known, which transmits radio waves toward or receives them from a specific direction. Known BF techniques include digital BF, analog BF, and hybrid BF. For example, Patent Document 1 discloses a receiver having an adaptive array with analog BF. Patent Document 2 discloses a process for suppressing inter-beam interference for digital signals as an example of digital BF processing. Patent Document 3 discloses suppressing the peak-to-average power ratio of electrical signals supplied to multiple transmitting antenna elements in a wireless transmitting station that performs hybrid BF.
[0004] JP 2002-135035 A JP 2017-224968 A JP 2015-34618 A
[0005] An electronic device according to one embodiment includes an antenna module capable of performing beamforming using a plurality of antenna elements, and a learning module capable of performing hierarchical reinforcement learning by a higher-level agent and a lower-level agent. The antenna module switches on or off the power supplies of a plurality of beamforming control units that control beamforming by the plurality of antenna elements. In the learning module, the higher-level agent controls the power supply states of the plurality of beamforming control units, and the lower-level agent controls the transmission power and / or reception power in the beamforming pattern by the plurality of antenna elements.
[0006] A control method according to one embodiment is a control method for an electronic device comprising: an antenna module capable of performing beamforming using a plurality of antenna elements; and a learning module capable of performing hierarchical reinforcement learning by a superior agent and a subordinate agent, wherein the antenna module includes a step of switching on or off the power of a plurality of beamforming control units that control beamforming by the plurality of antenna elements; a step of controlling the power supply states of the plurality of beamforming control units by the superior agent in the learning module; and a step of controlling the transmission power and / or reception power in the beamforming pattern by the plurality of antenna elements by the subordinate agent in the learning module.
[0007] A program according to one embodiment causes an electronic device including an antenna module capable of performing beamforming using a plurality of antenna elements, and a learning module capable of performing hierarchical reinforcement learning by a higher-level agent and a lower-level agent to execute the following steps: the antenna module switches on or off the power of a plurality of beamforming control units that control beamforming by the plurality of antenna elements; the higher-level agent in the learning module controls the power supply states of the plurality of beamforming control units; and the lower-level agent in the learning module controls the transmission power and / or reception power in the beamforming pattern by the plurality of antenna elements.
[0008] FIG. 1 is a block diagram schematically showing the functional configuration of an electronic device according to an embodiment. FIG. 2 is a block diagram showing the functional configuration of the radio section shown in FIG. 1 in more detail. FIG. 3 is a block diagram showing the functional configuration of the phased array antenna module shown in FIG. 2 in more detail. FIG. 4 is a block diagram showing an example of the functional configuration of the RF section shown in FIG. 3 in more detail. FIG. 4 is a block diagram showing another example of the functional configuration of the RF section shown in FIG. 3 in more detail. FIG. 5 is a diagram conceptually showing the operation of a learning module of an electronic device according to an embodiment. FIG. 6 is a diagram showing an example of a table stored in an electronic device according to an embodiment. FIG. 7 is a diagram showing an example of a table stored in an electronic device according to an embodiment. FIG. 8 is a diagram showing an example of a table stored in an electronic device according to an embodiment.
[0009] When performing beamforming using a wireless communication device or the like, it is desirable to reduce power consumption as much as possible while maintaining good communication quality. An object of the present disclosure is to provide an electronic device, a control method, and a program that can reduce power consumption while maintaining good wireless communication quality when performing beamforming. According to one embodiment, it is possible to provide an electronic device, a control method, and a program that can reduce power consumption while maintaining good wireless communication quality when performing beamforming.
[0010] In the present disclosure, an "electronic device" may refer to a device powered by electricity. Furthermore, a "system" may refer to a device or devices including a device powered by electricity. Furthermore, a "user" may refer to a person (typically a human) who uses a system and / or electronic device according to an embodiment. By using a system and / or electronic device according to an embodiment, a user can perform calibration with high accuracy while reducing processing costs. The system and / or electronic device according to an embodiment can reduce power consumption while maintaining good wireless communication quality when performing beamforming, for example, in a communication device using analog beamforming or hybrid beamforming.
[0011] First, the technical matters considered by the applicant when conceiving the present invention will be explained.
[0012] Base stations that use 5G (fifth generation mobile communication systems) wireless communications are subject to free space loss and other factors. Therefore, there are concerns that the radio waves transmitted from these base stations may not be able to travel as far as expected. To address this issue, analog beamforming (BF), a technology that concentrates antenna gain in a specific direction, is being adopted.
[0013] A 5G base station generally comprises an RU (Radio Unit) that constitutes the radio section including the antenna portion of the base station, a DU (Distributed Unit) also called a slave station, and a CU (Central Unit) called a master station or aggregation station. The RU has the function of transmitting and receiving radio waves to and from a UE (User Equipment), which is a terminal (mobile device) for communication, and communicating with the DU. The DU mainly has the function of modulating and demodulating signals and retransmitting lost signals. The CU mainly has the function of controlling multiple DUs and controlling RRC (Radio Resource Control), which is a communication protocol between the UE and the base station.
[0014] The RU of a base station configured as described above is equipped with a phased array antenna module (PAAM) connected to multiple antennas to support analog baseband. For this reason, the RU of a base station tends to consume more power than the CU and / or DU. In recent years, due to the influence of climate change and / or rising energy costs, there has been a strong demand for power saving in mobile communication systems. Therefore, power saving in RUs is desired in wireless communication technology.
[0015] Therefore, as a control to make the RU of the base station power-saving, it is possible to consider controls such as turning on / off the power of the components that make up the PAAM (for example, a functional unit that controls the BF (hereinafter referred to as the "BF control unit" as appropriate)) and adjusting the power for transmitting and receiving radio waves. Here, the BF control unit may be configured, for example, with an IC. Therefore, the BF control unit is also referred to as the "BF-IC" as appropriate.
[0016] However, when the power of a component constituting the PAAM (e.g., a BF control unit) is turned on / off, the analog BF pattern changes. In this case, the beam irradiation area differs for each analog BF pattern. Furthermore, the distribution of UE density in the area covered by a base station not only fluctuates over the long term, such as in the morning, evening, or night, but also fluctuates from moment to moment in the short term depending on the movement of UEs. When the distribution of UE density changes in this way, the distribution of requested traffic also fluctuates.
[0017] For these reasons, it would be desirable to be able to control the power saving of RUs in response to changes in conditions, such as fluctuations in UE density distribution. It would also be desirable to be able to control the transmission and reception power for each analog BF pattern. That is, it would be desirable to have a technology in a base station that supports analog BF that simultaneously controls the power on / off of components that make up the PAAM (e.g., a BF control unit) and the transmission and reception power for each analog BF pattern.
[0018] In view of the above circumstances, the applicant has conceived of an electronic device according to an embodiment. The electronic device according to an embodiment will be described below with reference to the drawings.
[0019] FIG. 1 is a block diagram illustrating a schematic functional configuration of an electronic device according to an embodiment.
[0020] 1 may be, for example, a base station that realizes wireless communication with a wireless communication terminal (mobile device) UE (User Equipment) by transmitting and receiving radio waves to and from the terminal UE. In FIG. 1, the wireless communication terminal (mobile device) UE that communicates with the electronic device 1 is not shown.
[0021] As shown in FIG. 1, an electronic device 1 according to an embodiment includes a wireless unit 10, a control unit 20, and a learning module 30.
[0022] The radio unit 10 may be based on the same concept as an RU in a base station that performs general 5G wireless communication, for example. That is, the radio unit 10 may have the same functions as an RU in a base station that performs general 5G wireless communication, for example.
[0023] 1, the radio unit 10 may include, for example, N (N is an integer equal to or greater than 2) antenna elements 110. Each of the N antenna elements 110 may be electrically connected to the radio unit 10. As will be described later, at least some of the N antenna elements 110 may constitute a phased array antenna.
[0024] The control unit 20 may be based on the same concept as the CU and DU in a base station that performs general 5G wireless communication, for example. That is, the radio unit 10 may have the same functions as the CU and DU in a base station that performs general 5G wireless communication, for example.
[0025] As shown in FIG. 1, the control unit 20 may have, for example, P electrical connections with the radio unit 10 (P is an integer equal to or greater than 1).
[0026] As shown in FIG. 1 , the learning module 30 may be electrically connected to the radio unit 10 and the control unit 20. The learning module 30 may control at least a portion of the operation of the radio unit 10. The learning module 30 may also receive information related to various controls performed by the control unit 20. The learning module 30 may control at least a portion of the operation of the radio unit 10 based on the results received from the control unit 20. The learning module 30 may be configured to include, for example, a hierarchical reinforcement learning module. The operation of the learning module 30 will be described further below.
[0027] In the electronic device 1 according to an embodiment, at least one of the wireless unit 10, the control unit 20, and the learning module 30 may include a memory or the like capable of storing various types of information as appropriate.
[0028] Furthermore, in the electronic device 1 according to an embodiment, at least one of the radio unit 10, the control unit 20, and the learning module 30 may provide control and processing capabilities for executing various functions. Therefore, at least one of the radio unit 10, the control unit 20, and the learning module 30 may include at least one processor, such as a central processing unit (CPU) or a digital signal processor (DSP). For example, at least one of the radio unit 10, the control unit 20, and the learning module 30 may be implemented collectively by a single processor, by several processors, or by individual processors. The processor may be implemented as a single integrated circuit. An integrated circuit is also referred to as an IC (integrated circuit). The processor may be implemented as multiple integrated circuits and discrete circuits connected to each other in a communicative manner. The processor may be implemented based on various other known technologies. At least one of the radio unit 10, the control unit 20, and the learning module 30 may be configured to include at least one of software and hardware resources, for example. In one embodiment, at least one of the radio unit 10, the control unit 20, and the learning module 30 may be configured by specific means in which software and hardware resources work together.
[0029] Furthermore, at least one of the radio unit 10, the control unit 20, and the learning module 30 may be configured as, for example, a CPU or a DSP and a program executed by the CPU or DSP. The program executed in at least one of the radio unit 10, the control unit 20, and the learning module 30, and the results of the processing executed therein, may be stored in, for example, any memory. At least one of the radio unit 10, the control unit 20, and the learning module 30 may include, as appropriate, a memory required for their operation.
[0030] The electronic device 1 according to an embodiment may not include some of the functional units shown in Fig. 1, or may include functional units other than those shown in Fig. 1. Furthermore, the system according to an embodiment may be configured to include the entire electronic device 1 or at least a part of the electronic device 1.
[0031] FIG. 2 is a block diagram showing in more detail the functional configuration of the radio unit 10 shown in FIG.
[0032] 2, the radio unit 10 may include a phased array antenna module (hereinafter referred to as "PAAM") 12, an ADC / DAC 14, and a memory 16. The radio unit 10 according to one embodiment may not include some of the functional units shown in FIG. 2, or may include functional units other than those shown in FIG. 2.
[0033] The PAAM 12 may be based on the same concept as a phased array antenna module in a base station performing general 5G wireless communication, for example. That is, the PAAM 12 may have the same functions as a phased array antenna module in a base station performing general 5G wireless communication, for example. However, in the electronic device 1 according to an embodiment, the PAAM 12 may be capable of receiving information stored in the memory 16. In this case, the PAAM 12 may operate based on the information received from the memory 16. Furthermore, in the electronic device 1 according to an embodiment, at least a portion of the operation of the PAAM 12 may be controlled by the learning module 30.
[0034] In the electronic device 1 according to one embodiment, the N antenna elements 110 may each be electrically connected to the PAAM 12 .
[0035] The ADC / DAC 14 may have the functions of an ADC (Analog to Digital Converter) and a DAC (Digital to Analog Converter). That is, the ADC / DAC 14 can convert an analog signal into a digital signal (ADC) and can also convert a digital signal into an analog signal (DAC). The ADC / DAC 14 may function as a DAC when transmitting radio waves from the antenna element 110, and may function as an ADC when receiving radio waves from the antenna element 110.
[0036] 2 , the PAAM 12 may include P ADC / DACs 14 (P is an integer equal to or greater than 1). The N antenna elements 110 may be electrically connected to the P ADC / DACs 14 via the PAAM 12. As described above, the radio unit 10 may have P electrical connections with the control unit 20. The PAAM 12 may be electrically connected to the control unit 20 via the P ADC / DACs 14.
[0037] The memory 16 may be any storage unit such as a semiconductor memory. The memory 16 may store various types of information. For example, in one embodiment, the memory 16 may store a table in which settings that can be selected using a control signal received from the learning module 30 as an argument are defined. Here, the control signal received from the learning module 30 may be, for example, actions g and a selected by each agent in hierarchical reinforcement learning. k The setting selectable using such a control signal as an argument may be, for example, a setting in which a set of BF control units that turn on the power is predefined. Also, for example, in one embodiment, the memory 16 may store a table in which analog BF patterns are predefined. Also, for example, in one embodiment, the memory 16 may store a table in which setting values of transmission and / or reception power when using each analog BF pattern are stored.
[0038] With the configuration shown in FIGS. 1 and 2, the electronic device 1 according to the embodiment can perform analog BF.
[0039] FIG. 3 is a block diagram showing in more detail the functional configuration of the PAAM 12 shown in FIG.
[0040] 3, the PAAM 12 may include RF units 120 corresponding to the N antenna elements 110, respectively. That is, the PAAM 12 may include N RF units 120 electrically connected to the N antenna elements 110, respectively. The PAAM 12 according to one embodiment may not include some of the functional units shown in FIG. 3, or may include functional units other than those shown in FIG. 3.
[0041] The RF unit 120 may function as the BF control unit (BF-IC) described above. That is, the RF unit 120 may function as a functional unit that controls the BF, and may be configured to include, for example, an IC. Here, the BF control unit can perform phase adjustment for analog BF under control from outside the PAAM 12. The BF control unit can also control the on / off of power to the PA (Power Amplifier) and / or LNA (Low Noise Amplifier) of the RF unit 120. Furthermore, the BF control unit can adjust the transmission and / or reception power. The configuration of the RF unit 120 functioning as the BF control unit (BF-IC) will be described further below. In the electronic device 1 according to an embodiment, the RF unit 120 (BF control unit (BF-IC)) may be a direct or indirect control target.
[0042] The combiner / divider 130 may be based on the same concept as at least one of a combiner and a divider of a PAAM in a base station that performs general 5G wireless communication, for example. That is, the combiner / divider 130 may have the same function as at least one of a combiner and a divider of a PAAM in a base station that performs general 5G wireless communication, for example.
[0043] As shown in Fig. 3 , K of the RF units 120 connected to the N antenna elements 110 may be electrically connected to one combiner / divider 130. The configuration shown in Fig. 3 illustrates an example in which K RF units 120 are electrically connected to P combiners / dividers 130, respectively. The P combiners / dividers 130 in the PAAM 12 may be electrically connected to their corresponding ADC / DACs 14. The P combiners / dividers 130 in the PAAM 12 may be electrically connected to the control unit 20 via their corresponding ADC / DACs 14.
[0044] In the electronic device 1 according to an embodiment, the PAAM 12 is not limited to the configuration shown in FIG. 3 . The PAAM 12 may have various configurations depending on the number of inputs and outputs of the RF unit 120 functioning as a BF control unit. As a relatively simple example, FIG. 3 illustrates a configuration in which the number of RF units 120 functioning as a BF control unit is equal to the number (N) of antenna elements 110. However, the number of RF units 120 functioning as a BF control unit can be changed depending on the number of inputs and outputs. Therefore, the number of RF units 120 functioning as a BF control unit is not limited to N.
[0045] As described above, in the electronic device 1 according to one embodiment, the PAAM 12 may be configured with a plurality of components (RF unit 120 (BF control unit (BF-IC))) that can be turned on and off. Furthermore, the electronic device 1 according to one embodiment may implement analog BF or hybrid BF by using the PAAM 12.
[0046] Fig. 4 is a block diagram showing in more detail an example of the functional configuration of the RF unit 120 shown in Fig. 3. As described above, the RF unit 120 may function as a BF control unit (BF-IC).
[0047] 4, the RF unit 120 may include a power amplifier (PA) 122, a low noise power amplifier (LNA) 124, a phase shifter 126A, a phase shifter 126B, a switch 128A, and a switch 128B. The RF unit 120 according to one embodiment may not include some of the functional units shown in FIG. 4, or may include functional units other than those shown in FIG. 4.
[0048] The PA 122 amplifies the power of the transmission signal supplied from the phase shifter 126A based on information stored in, for example, an arbitrary memory, etc. The technology itself, such as an amplifier that amplifies the power of the transmission signal, is already known, so a detailed description thereof will be omitted.
[0049] The LNA 124 amplifies, with low noise, a received signal based on radio waves received by the antenna element 110, based on information stored in a memory, for example. The LNA 124 may be a low noise amplifier, and amplifies, with low noise, the received signal supplied from the antenna element 110. The technology itself for amplifying a received signal with low noise is already known, so a detailed description thereof will be omitted.
[0050] In this way, the PA 122 and the LNA 124 control the amplitude of the signal. Therefore, when performing calibration in the electronic device 1 according to an embodiment, for example, the square of the amplitude component of the correction value may be set in the PA 122 and / or the LNA 124.
[0051] The phase shifter 126A controls the phase of a signal transmitted by the antenna element 110. The phase shifter 126B controls the phase of a signal received by the antenna element 110. Specifically, the phase shifter 126A or the phase shifter 126B may adjust the phase of the transmission signal or the reception signal by appropriately advancing or delaying the phase of the transmission signal or the reception signal based on information stored in a memory or the like. By the multiple phase shifters 126A or the phase shifters 126B appropriately adjusting the phase of each transmission signal or reception signal, radio waves transmitted or received from the multiple antenna elements 110 constructively interact with each other in a predetermined direction to form a beam (beamforming).
[0052] Phase shifter 126A and phase shifter 126B apply phase rotation to an analog signal. The precision of the phase rotation amount of the phase shifter is expressed by the number of bits. For example, if the precision of the phase rotation amount of the phase shifter is 2 bits, the signal can be rotated by {0, 90, 180, 360 degrees}. In this way, phase shifter 126A and phase shifter 126B control the phase of the signal.
[0053] The switches 128A and 128B are each capable of switching between a plurality of electrical connections.
[0054] 4 , the antenna element 110 and the RF unit 120 may be electrically connected. The antenna element 110 may be electrically connected to the PA 122 or the LNA 124 via a switch 128A. That is, by switching the switch 128A, the antenna element 110 is selectively connected to the PA 122 or the LNA 124.
[0055] The PA 122 may be electrically connected to a phase shifter 126A. The LNA 124 may be electrically connected to a phase shifter 126B. Also, as shown in FIG. 4 , the phase shifter 126A connected to the PA 122 and the phase shifter 126B connected to the LNA 124 may be electrically connected to the combiner / divider 130 ( FIG. 3 ) via a switch 128B. That is, by switching the switch 128B, the combiner / divider 130 and the ADC / DAC 14 are selectively connected to the phase shifter 126A connected to the PA 122 or the phase shifter 126B connected to the LNA 124.
[0056] In the RF section 120, switching between the switch 128A and the switch 128B can switch between reception and transmission of the antenna element 110. As shown in Fig. 4, when transmitting radio waves from the antenna element 110, the signal is transmitted via the PA 122. On the other hand, when receiving radio waves from the antenna element 110, the signal is received via the LNA 124.
[0057] To perform analog BF in the electronic device 1 according to an embodiment, a phase rotation amount is set in the phase shifter 126A and / or the phase shifter 126B. The electronic device 1 according to an embodiment may control the on / off of the power supply for each RF unit 120 by using control based on hierarchical reinforcement learning with the learning module 30. Furthermore, the electronic device 1 according to an embodiment may control the power of the PA 122 and / or the LNA 124 by using control based on hierarchical reinforcement learning with the learning module 30.
[0058] In the electronic device 1 according to an embodiment, the RF unit 120 is not limited to the configuration shown in Fig. 4. In the electronic device 1 according to an embodiment, the RF unit 120 may be configured as shown in Fig. 5, for example, depending on the configuration of the PAAM 12. The RF unit 120' shown in Fig. 5 is a block diagram showing in more detail another example of the functional configuration of the RF unit 120 shown in Fig. 3.
[0059] The RF section 120' shown in Fig. 5 may include, for example, a combiner 132 and a distributor 134 in addition to the functional sections constituting the RF section 120 shown in Fig. 4. The RF section 120' shown in Fig. 5 may include not only phase shifters 126A and 126B but also phase shifters 126C and 126D. Furthermore, the RF section 120' shown in Fig. 5 may include not only switches 128A and 128B but also switch 128C.
[0060] As shown in FIG. 5 , the PA 122 may be electrically connected to the combiner 132. The LNA 124 may be electrically connected to the divider 134. The combiner 132 may be electrically connected to the phase shifter 126A and the phase shifter 126C. The divider 134 may be electrically connected to the phase shifter 126B and the phase shifter 126D. In the configuration shown in FIG. 5 , the switch 128B switches the path between the phase shifter 126A and the phase shifter 126B. Similarly, the switch 128C switches the path between the phase shifter 126C and the phase shifter 126D.
[0061] The RF unit 120' shown in Figure 5 is in a state during transmission. When the electronic device 1 transmits a signal, the transmission signal passes through phase shifter 126A and phase shifter 126C and is supplied to combiner 132. When the RF unit 120' is receiving, switches 128A, 128B, and 128C are switched from the states shown in Figure 5. When the electronic device 1 receives a signal, the received signal passes through distributor 134 and is supplied to phase shifter 126B and phase shifter 126D.
[0062] Next, the hierarchical reinforcement learning executed by the learning module 30 in the electronic device 1 according to an embodiment will be further described.
[0063] The electronic device 1 according to an embodiment may control the function of the RU based on information about reception quality determined by the above-described function of the CU and / or function of the DU, thereby controlling the combination of active components in the PAAM 12. At the same time, the electronic device 1 according to an embodiment may control the function of the RU based on information about reception quality determined by the above-described function of the CU and / or function of the DU, thereby controlling the transmission and / or reception power for each analog BF pattern.
[0064] The learning module 30 uses hierarchical reinforcement learning to switch on / off the power of components constituting the PAAM 12 (e.g., the BF control unit (RF unit 120)) and adjust the transmission and / or reception power of the analog BF formed thereby. In this way, the electronic device 1 according to one embodiment can achieve power saving of the RU function of the base station. In one embodiment, the learning module 30 may be implemented within the CU function and / or the DU function, or may be implemented as a device external to the electronic device 1, for example.
[0065] Next, the hierarchical reinforcement learning performed by the learning module 30 will be described.
[0066] 6 is a diagram conceptually illustrating the operation of the learning module 30 of the electronic device 1 according to an embodiment. FIG. 6 conceptually illustrates hierarchical reinforcement learning as the operation of the learning module 30.
[0067] Reinforcement learning is an algorithm that searches for optimal behavior by repeating trial and error between an agent and the environment. Here, the agent may be something that receives a state and a reward from the environment and determines its behavior. The environment may be something that feeds back to the agent the state and reward that are determined anew by the agent's behavior. Q-learning is one type of reinforcement learning algorithm that has been studied. Q-learning optimizes the Q-value, which takes the state and behavior as arguments, by updating it with the reward received from the environment.
[0068] In addition, hierarchical reinforcement learning is a method in which a plurality of agents form a hierarchical structure. As an example, as shown in Figure 6, a case will be assumed in which two levels of agents form a hierarchical structure (a higher-level agent 32 and lower-level agents 34-1 to 34-k). In this case, as shown in Figure 6, the agent receives a state (s, s) from the environment 36. k ) and reward (r, r k ) to take the action (g, a k 6, the environment 36 receives the action decided by the agent and feeds back the resulting state and reward to the agent. The cost function is designed to correspond to the expected value of the reward.
[0069] Each agent in the hierarchy may have a dedicated Q table in which the Q value is stored. The Q value in the upper agent may be updated, for example, as shown in the following equation (1).
[0070]
[0071] Here, time t is the update count number of the Q value. The update cycle of the Q value is set in consideration of the calculation accuracy of the state and the reward, and the trade-off between the control speed. The update cycle of the Q value may be set appropriately so as to obtain good results.
[0072] In the above formula (1), s(t) represents the state received by the higher-ranking agent at time t. s(t+1) represents the state received by the higher-ranking agent at time t+1. g represents the action selected by the higher-ranking agent at time t. r(t+1) represents the reward received by the higher-ranking agent at time t+1. α and γ represent hyperparameters of Q-learning.
[0073] Furthermore, the Q value of the k-th lower-level agent may be updated as shown in equation (2).
[0074]
[0075] In the above formula (2), s k(t) represents the state received by the kth subordinate agent at time t. k (t+1) represents the state received by the kth lower agent at time t+1. g represents the action selected by the upper agent at time t. g' represents the action selected by the upper agent at time t+1. r k (t+1) represents the reward received by the kth subordinate agent at time t+1. α and γ represent hyperparameters of Q-learning. Here, α and γ may be set to different values for the superior agent and / or other subordinate agents.
[0076] The agent then determines its next action with reference to the Q value. In one embodiment, for example, the agent may select the action that maximizes the Q value in the current state.
[0077] The actions, states, and rewards in the above-mentioned hierarchical reinforcement learning will be further explained below. The applicant has confirmed that by appropriately defining actions in hierarchical reinforcement learning, it is possible to reduce power consumption while maintaining good wireless communication quality when performing beamforming.
[0078] Fig. 7 is a diagram showing an example of a table stored in the electronic device 1 according to an embodiment. Fig. 7 is a table showing a set of RF units 120 (BF control units) used in an action selected by the upper agent 32. Fig. 7 may show an example of a table stored in, for example, the learning module 30 and / or the memory 16.
[0079] In one embodiment, the action in the upper agent 32 may be selected from a table of predefined sets of RF units 120 (BF control units) to be powered on, as shown in FIG. 7 . In FIG. 7 , g may be the number of the set of RF units 120 (BF control units) to be powered on. Also, in FIG. 7 , RF unit 120#1 indicates the first of the N RF units 120. RF unit 120#2 indicates the second of the N RF units 120. RF unit 120#N indicates the Nth of the N RF units 120.
[0080] For example, as shown in Fig. 7, when the upper agent 32 selects action g = 0, the power supplies of the RF units 120#1 to 120#4 and the RF unit 120#N may be turned on. Also, as shown in Fig. 7, when the upper agent 32 selects action g = 1, the power supplies of the RF units 120#1 and 120#3 may be turned on, and the power supplies of the RF units 120#2, 120#4, and 120#N may be turned off.
[0081] In one embodiment, the behavior of the subordinate agent may be to adjust the transmission or reception power for each predefined analog BF pattern. As shown in FIG. 6, there may be multiple subordinate agents. Here, each of the multiple subordinate agents may be associated with a number that identifies a specific analog BF pattern. The number that identifies a specific analog BF pattern may be called an SSB (Synchronization Signal Block) index or a beam index.
[0082] A wireless communication terminal UE communicating with the electronic device 1 searches for an appropriate analog BF pattern during cell search. The SSB index is the number of the SSB signal. The SSB signal is analog-BFed using the analog BF pattern. Therefore, the terminal UE is essentially searching for an analog BF pattern. The terminal UE can establish a cell connection through a cell search using the SSB signal. The terminal UE can proceed with subsequent procedures, such as data communication, by associating the SSB index with the SSB index. Therefore, by associating the subagent with the SSB index, the learning module 30 can perform Q-learning. For example, the kth subagent may control the analog BF used with SSB index #k as the behavior control target in Q-learning. Furthermore, for example, the kth subagent may calculate the state and reward in Q-learning based on information on communication quality from the terminal UE connected to the cell.
[0083] Action a selected by the subordinate agent kmay be a predetermined power value that is reduced from the maximum transmission or reception power. k = {0, 1, 2, 3}, the power value to be reduced from the maximum power may be {0 dB, -3 dB, -6 dB, -9 dB}. An increase or decrease in transmission or reception power corresponds to an increase or decrease in power consumption of the RF unit 120 (BF control unit), which leads to power saving control of the RU function in the electronic device 1.
[0084] In one embodiment, one subordinate agent may adjust the transmit or receive power of one analog BF pattern.
[0085] Fig. 8 is a diagram showing an example of a table stored in the electronic device 1 according to an embodiment. Fig. 8 is a table showing the correspondence between analog BF patterns controlled by each lower-level agent. Fig. 8 may show an example of a table stored in, for example, the learning module 30 and / or the memory 16.
[0086] As shown in Fig. 8, the patterns (analog BF patterns) controlled by each of the lower agents (lower agent #1 to lower agent #k) may be defined in advance. For example, in Fig. 8, when the action g selected by the higher agent is 0, the analog BF pattern controlled by the lower agent #1 is pattern A_1. On the other hand, when the action g selected by the higher agent is 1, the analog BF pattern controlled by the lower agent #1 is pattern B_1. In this way, when the value of the action g selected by the higher agent is different, the analog BF patterns controlled by the same lower agent (e.g., lower agent #1) may be different.
[0087] The analog BF pattern changes depending on the number of antenna elements 110 and / or the arrangement of the antenna elements 110 (for example, the spacing between the antenna elements 110). Therefore, the analog BF pattern also changes depending on the set of RF units 120 (BF control units) that are powered on. For this reason, the analog BF patterns controlled by the same lower-level agent differ. On the other hand, depending on the set of RF units 120 (BF control units) that are powered on, the analog BF pattern controlled by the same lower-level agent may not change.
[0088] In one embodiment, one subordinate agent may collectively adjust the transmission and / or reception power of multiple analog BF patterns. For example, the electronic device 1 according to one embodiment may collectively adjust the transmission and / or reception power of the analog BF patterns based on the similarity of the analog BF patterns. Here, the similarity of the analog BF patterns may refer to pattern characteristics such as the radiation direction and / or half-width of the main lobe. For example, in one embodiment, an analog BF pattern with a radiation direction of 0° horizontally and 10° vertically and an analog BF pattern with a radiation direction of 0° horizontally and 20° vertically may be pre-grouped as analog BF patterns that are similar to each other. In this case, in one embodiment, one subordinate agent may be assigned to determine the transmission and / or reception power of the pre-grouped analog BF patterns.
[0089] As an example, a table showing the correspondence relationship of analog BF patterns controlled by each lower-level agent in such a case is shown in Fig. 9. Fig. 9 is a diagram showing an example of a table stored in the electronic device 1 according to an embodiment. Fig. 9 may show an example of a table stored in, for example, the learning module 30 and / or the memory 16. Fig. 9 may be viewed in the same way as Fig. 8.
[0090] In the hierarchical reinforcement learning shown in FIG. 6, the environment 36 is a mobile communication system including the electronic device 1, which is a base station. k ) and reward (r, r k) is calculated by the CU function and / or the DU function of the electronic device 1. The result of such calculation is fed back to each agent.
[0091] In hierarchical reinforcement learning, the state may be, for example, information about reception quality, such as received signal strength indicator (RSSI), reference signal received power (RSRP), reference signal received quality (RSRQ), signal to interference plus noise ratio (SINR), or channel quality indicator (CQI).
[0092] In hierarchical reinforcement learning, the state may be calculated on the side of the electronic device 1, which is a base station. Alternatively, the state may be calculated on the side of the UE and fed back to the electronic device 1, which is a base station. The state in the lower agent may be a value obtained by averaging information on reception quality calculated in the electronic device 1, which is a base station (or calculated in the terminal UE), when the analog BF pattern is applied. The state in the upper agent may be a value obtained by averaging information on reception quality calculated in the electronic device 1, which is a base station (or calculated in the terminal UE), without distinguishing between analog BF patterns.
[0093] In the hierarchical reinforcement learning shown in Fig. 6, the reward may be a function of the required communication quality (e.g., required cell throughput) and the power consumption due to the function of the RU. In one embodiment, the reward may be calculated for each analog BF pattern in the learning module 30. The result calculated in this way is fed back to each agent.
[0094] Specifically, when calculating the reward, for example, the learning module 30 may determine whether the sum of the actual throughputs in the communication between the electronic device 1, which is the base station, and the terminal UE exceeds a preset required throughput.
[0095] If the sum of the actual throughputs does not exceed the required throughput, the learning module 30 may set the rewards of the higher-level agents and the lower-level agents to a predetermined penalty value. In this case, the learning module 30 may set the penalty value to a negative value so that the rewards are lowered.
[0096] On the other hand, if the sum of the actual throughputs exceeds the required throughput, the learning module 30 may set the reward for the higher-level agent to the reciprocal of the power consumption as a function of the RU of the wireless unit 10. In this case, the learning module 30 may set the reward for the lower-level agent to the reciprocal of the power consumed for transmitting or receiving the analog BF pattern. In this case, by setting the reward to the reciprocal of the power, the reward may be designed to be high when the power consumption is relatively low.
[0097] In hierarchical reinforcement learning, by defining the actions, states, and rewards as described above, the learning module 30 can utilize the general Q-learning algorithm.
[0098] Next, the distinction between the MIMO (Multi-Input Multi-Output) method and DL (Down Link) / UL (Up Link) communication in the hierarchical reinforcement learning by the learning module 30 will be further explained.
[0099] The learning module 30 may calculate states and rewards in hierarchical reinforcement learning under these conditions, regardless of the distinction between TDD and FDD, the MIMO system (SU-MIMO or MU-MIMO), and the distinction between DL and UL communication. Therefore, the electronic device 1 according to an embodiment can operate independently of the distinction between TDD and FDD, the MIMO system, and the distinction between DL and UL communication. The electronic device 1 according to an embodiment does not depend on the distinction between DL and UL communication. Therefore, in the learning module 30, the subordinate agent may control transmission and / or reception power.
[0100] In addition, in the case of a TDD system, DL / UL communications are symmetrical, so the learning module 30 may apply, for example, behavior determined under DL communication conditions to UL communications as well as DL communications.
[0101] In order to realize the above-described hierarchical reinforcement learning, in the electronic device 1 according to an embodiment, various necessary information may be stored in, for example, the learning module 30 and / or the memory 16. For example, the learning module 30 and / or the memory 16 may receive control signals (g, a which are actions selected by each agent) from the hierarchical reinforcement learning module. k The learning module 30 and / or the memory 16 may store a table in which a set of BF control units to turn on the power is predefined, and which can be selected using the parameter (value of the parameter) as an argument. Also, the learning module 30 and / or the memory 16 may store a table in which, for example, analog BF patterns are predefined. Also, the learning module 30 and / or the memory 16 may store a table in which setting values of transmission and / or reception power when using each analog BF pattern are stored.
[0102] As described above, the electronic device 1 according to one embodiment includes the PAAM 12 configured with a plurality of components (RF unit 120 (BF control unit (BF-IC))) that can be individually controlled to be powered on / off. The electronic device 1 according to one embodiment uses the PAAM 12 to enable analog BF or hybrid BF while controlling the RU function of the radio unit 10 to save power. More specifically, the electronic device 1 according to one embodiment uses hierarchical reinforcement learning by the learning module 30 to control the power on / off of the components (e.g., the RF unit 120) that make up the PAAM 12. At the same time, the electronic device 1 according to one embodiment also controls the transmission and / or reception power of individual beams (analog BF patterns).
[0103] Furthermore, as described above, in the electronic device 1 according to an embodiment, the learning module 30 may employ two-stage hierarchical reinforcement learning involving a higher-level agent and a lower-level agent. In the learning module 30, the target controlled by the higher-level agent (target selected by action) may be a set of RF units 120 (BF control units (BF-IC)) that are predefined and turned on, as shown in FIG. 7 . In the learning module 30, the lower-level agent may control the transmission and / or reception power for individual analog BF patterns. The learning module 30 may predefine the analog BF patterns, for example, as shown in at least one of FIGS. 7 , 8 , and 9 , so that the analog BF patterns are determined in association with the action (g) of the higher-level agent, as shown in FIGS. 8 and 9 .
[0104] As such, the electronic device 1 according to one embodiment may include an antenna module such as the PAAM 12 and a learning module such as the learning module 30. The antenna module may be configured to be capable of performing beamforming using multiple antenna elements. The learning module may be configured to be capable of performing hierarchical reinforcement learning by a higher-level agent and a lower-level agent. The antenna module also switches on or off the power of multiple beamforming control units that control beamforming using the multiple antenna elements. In the learning module, the higher-level agent may control the power states of the multiple beamforming control units. In the learning module, the lower-level agent may control the transmission power and / or reception power in the beamforming pattern using the multiple antenna elements.
[0105] Furthermore, in the learning module 30 according to one embodiment, information defined as an object to be controlled by a lower-level agent may be associated with information defined as an object to be controlled by a higher-level agent. In this case, the information defined as an object to be controlled by a higher-level agent may be information relating to the power on or off of a plurality of beamforming control units. Furthermore, the information defined as an object to be controlled by a lower-level agent may be information relating to the beamforming pattern of a plurality of antenna elements.
[0106] Furthermore, in the learning module 30 according to one embodiment, the control target of the lower agent may be defined in accordance with the action g of the upper agent, or such a definition may be made in advance. In this case, the control target of the upper agent may be defined as the control related to the power on / off of a plurality of beamforming control units. Furthermore, the control target of the lower agent may be defined as the control related to the beamforming pattern of a plurality of antenna elements.
[0107] According to an embodiment of the electronic device 1, the power consumed by the RU function of the radio unit 10 can be minimized while satisfying a statistical value, such as the average value or the sum of the average and variance, of the communication quality set / required for the electronic device 1, which is a base station, over a predetermined time period. Here, the communication quality set / required for the electronic device 1 may be set / required by an operator based on some kind of guideline, for example. Furthermore, the communication quality set / required for the electronic device 1 may be based on, for example, cell throughput, packet delay time, RSRP, RSRQ, or SINR. Thus, the electronic device 1 according to an embodiment employs a learning module 30 that performs hierarchical reinforcement learning, thereby realizing power saving for the RU function of the radio unit 10.
[0108] Below, we will further explain points to note regarding the electronic device 1 according to an embodiment. The electronic device 1 according to an embodiment can be realized as a device such as a base station that performs analog BF. On the other hand, hybrid BF is a method that performs both analog BF and digital BF. Therefore, the electronic device 1 according to an embodiment can also be realized as a part that performs analog BF in a device that performs hybrid BF. Furthermore, for the sake of simplicity, the RF unit 120 (BF control unit (BF-IC)) described above is assumed to include not only the phase shifter 126A and the phase shifter 126B, but also the PA 122 and the LNA 124. In such a case, the power consumption of the PA 122 and the LNA 124 is relatively large. Therefore, in the electronic device 1 according to an embodiment, it is possible to assume that the power is turned on / off as the RF unit 120 (BF control unit (BF-IC)) that collectively includes these components.
[0109] Although the embodiments of the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art would easily be able to make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of the present disclosure. For example, functions included in each component or step can be rearranged so as not to cause logical inconsistencies, and multiple components or steps can be combined or divided into one. Although the embodiments of the present disclosure have been described primarily in terms of an apparatus, the embodiments of the present disclosure can also be realized as a method including steps executed by each component of the apparatus. The embodiments of the present disclosure can also be realized as a method, a program executed by a processor included in an electronic device, or a storage medium or recording medium on which a program is recorded. It should be understood that these are also encompassed within the scope of the present disclosure.
[0110] The above-described embodiments are not limited to implementation as the electronic device 1. For example, the above-described embodiments may be implemented as a system including the electronic device 1. Furthermore, the above-described embodiments may be implemented as, for example, a control method for the electronic device 1 or a control method for a device such as a system including the electronic device 1. Furthermore, the above-described embodiments may be implemented as, for example, a program executed by a device such as the electronic device 1 or a system including the electronic device 1, or an information processing device (e.g., a computer). Furthermore, in the technology disclosed herein, all of the components of the electronic device 1 and / or a system including the electronic device 1 do not need to reside in a single housing. For example, the controllers and / or memory units of the components of the electronic device 1 and / or a system including the electronic device 1 may be connected to each other via a network that is wired, wireless, or a combination thereof.
[0111] A power control system, a power control device, and / or a power control method according to an embodiment may be implemented, for example, as follows. [Supplementary Note 1] An electronic device comprising: an antenna module capable of performing beamforming using a plurality of antenna elements; and a learning module capable of performing hierarchical reinforcement learning by a superior agent and a subordinate agent, wherein the antenna module switches on or off the power of a plurality of beamforming control units that control beamforming by the plurality of antenna elements, and in the learning module, the superior agent controls the power state of the plurality of beamforming control units, and the subordinate agent controls the transmission power and / or reception power in a beamforming pattern by the plurality of antenna elements. [Supplementary Note 2] The electronic device according to Supplementary Note 1, wherein, in the learning module, information defined as an object to be controlled by the subordinate agent is associated with information defined as an object to be controlled by the superior agent. [Supplementary Note 3] The electronic device according to Supplementary Note 2, wherein the information defined as an object to be controlled by the superior agent is information regarding the power on or off of the plurality of beamforming control units. [Supplementary Note 4] An electronic device according to Supplementary Note 2 or 3, wherein the information defined as the object to be controlled by the lower agent is information relating to a beamforming pattern by the multiple antenna elements. [Supplementary Note 5] An electronic device according to any of Supplements 1 to 4, wherein the object to be controlled by the lower agent is defined in the learning module according to the action of the upper agent. [Supplementary Note 6] An electronic device according to Supplementary Note 5, wherein control relating to turning on or off power to the multiple beamforming control units is defined as the object to be controlled by the upper agent. [Supplementary Note 7] An electronic device according to Supplementary Note 5 or 6, wherein control relating to a beamforming pattern by the multiple antenna elements is defined as the object to be controlled by the lower agent.[Supplementary Note 8] A control method for an electronic device comprising: an antenna module capable of performing beamforming using a plurality of antenna elements; and a learning module capable of performing hierarchical reinforcement learning by a higher-level agent and a lower-level agent, the control method including the steps of: in the antenna module, switching on or off the power supplies of a plurality of beamforming control units that control beamforming by the plurality of antenna elements; in the learning module, using the higher-level agent, setting the power supply states of the plurality of beamforming control units as control targets; and in the learning module, setting the transmission power and / or reception power in a beamforming pattern by the plurality of antenna elements as control targets. [Supplementary Note 9] A program for causing an electronic device comprising: an antenna module capable of performing beamforming using a plurality of antenna elements; and a learning module capable of performing hierarchical reinforcement learning by a higher-level agent and a lower-level agent, to execute the following steps: by the antenna module switching on or off the power supplies of a plurality of beamforming control units that control beamforming by the plurality of antenna elements; by the higher-level agent in the learning module setting the power supply states of the plurality of beamforming control units as control targets; and by the lower-level agent in the learning module setting the transmission power and / or reception power in the beamforming pattern by the plurality of antenna elements as control targets.
[0112] 1 Electronic device 10 Radio section 12 Phased array antenna module (PAAM) 14 ADC / DAC 16 Memory 20 Control section 30 Learning module 110 Antenna element 120 RF section 122 PA 124 LNA 126A, 168B Phase shifter 128A, 128B Switch 130 Combiner / Divider 140A, 140B Combiner
Claims
1. An electronic device comprising: an antenna module capable of performing beamforming using multiple antenna elements; and a learning module capable of performing hierarchical reinforcement learning by a higher-level agent and a lower-level agent, wherein the antenna module switches on or off the power of multiple beamforming control units that control beamforming by the multiple antenna elements; and in the learning module, the higher-level agent controls the power state of the multiple beamforming control units, and the lower-level agent controls the transmission power and / or reception power in the beamforming pattern by the multiple antenna elements.
2. The electronic device according to claim 1, wherein in said learning module, information defined as an object to be controlled by said lower-level agent is associated with information defined as an object to be controlled by said higher-level agent.
3. The electronic device according to claim 2, wherein the information defined as the object to be controlled by the upper agent is information regarding the power on or off of the plurality of beamforming control units.
4. The electronic device according to claim 2 or 3, wherein the information defined as the object to be controlled by the lower-level agent is information relating to a beamforming pattern by the plurality of antenna elements.
5. The electronic device according to any one of claims 1 to 4, wherein the learning module defines an object to be controlled by the lower agent in accordance with the behavior of the upper agent.
6. The electronic device according to claim 5, wherein control of turning on or off the power of the plurality of beamforming control units is defined as an object to be controlled by the upper agent.
7. The electronic device according to claim 5 or 6, wherein control of a beamforming pattern by the plurality of antenna elements is defined as an object to be controlled by the lower-level agent.
8. A control method for an electronic device comprising: an antenna module capable of performing beamforming using a plurality of antenna elements; and a learning module capable of performing hierarchical reinforcement learning by a higher-level agent and a lower-level agent, the control method including the steps of: the antenna module switching on or off the power supplies of a plurality of beamforming control units that control beamforming by the plurality of antenna elements; the higher-level agent in the learning module controlling the power supply states of the plurality of beamforming control units; and the lower-level agent in the learning module controlling the transmission power and / or reception power in the beamforming pattern by the plurality of antenna elements.
9. A program for causing an electronic device comprising an antenna module capable of performing beamforming using a plurality of antenna elements, and a learning module capable of performing hierarchical reinforcement learning by a higher-level agent and a lower-level agent, to execute the following steps: the antenna module switches on or off the power of a plurality of beamforming control units that control beamforming by the plurality of antenna elements; the higher-level agent in the learning module controls the power supply states of the plurality of beamforming control units; and the lower-level agent in the learning module controls the transmission power and / or reception power in the beamforming pattern by the plurality of antenna elements.