Beam management using adaptive learning
Optimizing beam management through adaptive learning algorithms, the problems of low beam management efficiency and high energy consumption in wireless communication systems are solved, and more efficient beam selection and signal transmission are achieved.
Patent Information
- Application Number
- CN202080031295.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-01
- Filing Date
- 2020-04-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-04-10
AI Technical Summary
Existing wireless communication systems have problems with inefficiency and high energy consumption in beam management, especially in millimeter wave communications, where beam management procedures require extensive measurement and adjustment to adapt to changing communication environments.
Adaptive learning algorithms are adopted, using artificial neural networks and machine learning technology, and continuously optimize beam management procedures through training and feedback, and select the optimal beam pairing to reduce measurement times and power consumption.
It improves the efficiency of beam management and reduces energy consumption, enables more intelligent beam selection, adapts to different communication environments, and improves signal transmission quality and throughput.
Smart Images

Figure CN113785503B_ABST
Abstract
Description
[0001] Priority claim
[0002] This patent application claims priority to U.S. non-provisional application No. 16 / 400,864, filed on May 1, 2019, entitled “BEAM MANAGEMENT USING ADAPTIVE LEARNING,” which is assigned to the assignee of the present application and is hereby expressly incorporated herein by reference.
[0003] introduction
[0004] Aspects of the present disclosure relate generally to wireless communications and, more particularly, to techniques for beam management.
[0005] Wireless communication systems are widely deployed to provide various telecommunication services such as telephony, video, data, messaging, broadcast, etc. These wireless communication systems may employ multiple-access technologies capable of supporting communication with multiple users by sharing available system resources (e.g., bandwidth, transmit power, etc.). Examples of such multiple-access systems include Third Generation Partnership Project (3GPP) Long Term Evolution (LTE) systems, LTE-Advanced (LTE-A) systems, Code Division Multiple Access (CDMA) systems, Time Division Multiple Access (TDMA) systems, Frequency Division Multiple Access (FDMA) systems, Orthogonal Frequency Division Multiple Access (OFDMA) systems, Single Carrier Frequency Division Multiple Access (SC-FDMA) systems, and Time Division Synchronous Code Division Multiple Access (TD-SCDMA) systems, to name just a few.
[0006] In some examples, a wireless multiple access communication system may include several base stations (BS), each of which is capable of supporting communication of multiple communication devices (also referred to as user equipment (UE)) simultaneously. In an LTE or LTE-A network, a set of one or more base stations may define an evolved B node (eNB). In other examples (e.g., in a next generation, new radio (NR), or 5G network), a wireless multiple access communication system may include several distributed units (DUs) (e.g., edge units (EUs), edge nodes (ENs), radio heads (RHs), smart radio heads (SRHs), transmission reception points (TRPs), etc.) in communication with several central units (CUs) (e.g., central nodes (CNs), access node controllers (ANCs), etc.), wherein a set of one or more DUs in communication with a CU may define an access node (e.g., which may be referred to as a BS, a next generation B node (gNB or g B node), a TRP, etc.). A BS or DU may communicate with a set of UEs on a downlink channel (e.g., for transmission from a BS or DU to a UE) and an uplink channel (e.g., for transmission from a UE to a BS or DU).
[0007] These multiple access technologies have been adopted in various telecommunication standards to provide a common protocol that enables different wireless devices to communicate at a city, country, region, and even global level. New radio (e.g., 5G NR) is an example of an emerging telecommunication standard. NR is an enhancement set of the LTE mobile standard promulgated by 3GPP. NR is designed to better support mobile broadband Internet access by improving spectrum efficiency, reducing costs, improving services, utilizing new spectrum, and better integrating with other open standards using OFDMA with cyclic prefix (CP) on downlink (DL) and uplink (UL). To this end, NR supports beamforming, multiple-input multiple-output (MIMO) antenna technology, and carrier aggregation.
[0008] However, as the demand for mobile broadband access continues to grow, there is a need for further improvements to NR and LTE technologies. Preferably, these improvements should be applicable to other multiple access technologies and the telecommunication standards that employ these technologies. Summary of the invention
[0009] The systems, methods, and apparatus of the present disclosure each have several aspects, no single one of which is solely responsible for its desirable attributes. Without limiting the scope of the disclosure as expressed in the appended claims, some features will now be briefly discussed. After considering this discussion, and particularly after reading the section entitled "Detailed Description," it will be understood how the features of the present disclosure provide advantages including an improved beam management procedure using adaptive learning.
[0010] Certain aspects provide a method for wireless communication by a node. The method generally includes determining one or more beams to be used for a beam management procedure using adaptive learning. The method generally includes performing a beam management procedure using the determined one or more beams.
[0011] In some examples, the node is a base station (BS).
[0012] In some examples, the node is a user equipment (UE).
[0013] In some examples, the method includes updating an adaptive learning algorithm used for adaptive learning. In some examples, the adaptive learning algorithm is updated based on feedback and / or training information. In some examples, the method includes performing another beam management procedure using the updated adaptive learning algorithm.
[0014] In some examples, the feedback includes feedback associated with a beam management procedure.
[0015] In some examples, the training information includes one or more of the following: training information acquired by deploying one or more UEs in one or more simulated communication environments prior to network deployment of the one or more UEs; training information acquired through feedback previously received when the one or more UEs were deployed in one or more communication environments; training information from the network, one or more UEs and / or the cloud; and / or training information received when the node is online and / or idle.
[0016] In some examples, the training information includes training information received from one or more UEs different from the node after the node is deployed. In some examples, the training information includes information associated with the beam. In some examples, the training information includes measurements made by one or more UEs or feedback associated with one or more beam management procedures performed by the one or more UEs.
[0017] In some examples, using the adaptive learning algorithm includes outputting an action based on one or more inputs. In some examples, feedback is associated with the action. In some examples, updating the adaptive learning algorithm based on the feedback includes adjusting one or more weights applied to the one or more inputs.
[0018] In some examples, the adaptive learning algorithm includes an adaptive machine learning algorithm; an adaptive reinforcement learning algorithm; an adaptive deep learning algorithm; an adaptive continuous infinite learning algorithm; and / or an adaptive policy optimization reinforcement learning algorithm.
[0019] In some examples, the adaptive learning algorithm is modeled as a partially observable Markov decision process (POMDP).
[0020] In some examples, the adaptive learning algorithm is implemented by an artificial neural network.
[0021] In some examples, the artificial neural network includes a deep Q network (DQN) including one or more deep neural networks (DNNs). In some examples, using adaptive learning to determine one or more beams includes passing state parameters and action parameters through one or more DNNs; for each state parameter, outputting a value of each action parameter; and selecting an action associated with a maximum output value.
[0022] In some examples, updating the adaptive learning algorithm includes adjusting one or more weights associated with one or more neuronal connections in the artificial neural network.
[0023] In some examples, using adaptive learning to determine one or more beams to be used for a beam management procedure includes determining one or more beams to be included in a codebook based on the adaptive learning and selecting the one or more beams to be used for the beam management procedure from the codebook.
[0024] In some examples, determining one or more beams to be used for the beam management procedure includes using adaptive learning to select one or more beams to be used for the beam management procedure from a codebook.
[0025] In some examples, adaptive learning uses state parameters associated with channel measurements, reward parameters associated with received signal throughput or spectral efficiency, and action parameters associated with selection of beam pairs corresponding to the channel measurements.
[0026] In some examples, channel measurements include reference signal received power (RSRP); spectral efficiency, channel flatness, and / or signal-to-noise ratio (SNR).
[0027] In some examples, the received signal includes a physical downlink shared channel (PDSCH) transmission.
[0028] In some examples, the reward parameter is deducted by a penalty amount.
[0029] In some examples, the amount of penalty depends on the number of one or more beams being measured for the beam management procedure.
[0030] In some examples, the amount of penalty depends on the amount of power consumption associated with the beam management procedure.
[0031] In some examples, the beams include one or more beams for transmission and / or reception of one or more synchronization signal blocks (SSBs).
[0032] In some examples, performing a beam management procedure using the determined one or more beams includes measuring a channel using the determined one or more beams based on an SSB transmission from a BS, the SSB transmission being associated with one or more transmit beams of the BS; and selecting one or more beam pair links (BPLs) associated with one or more channel measurements that are above a channel measurement threshold and / or are one or more strongest channel measurements of all channel measurements associated with the SSB transmissions.
[0033] In some examples, the determined one or more beams include a subset of available receive beams.
[0034] In some examples, the method includes receiving a PDSCH using one of one or more selected BPLs; determining a throughput associated with the PDSCH; updating an adaptive learning algorithm based on the determined throughput; and using the updated adaptive learning algorithm to determine another one or more beams to be used to execute another beam management procedure for selecting another one or more BPLs.
[0035] Certain aspects provide a node configured for wireless communication. The node generally includes means for determining one or more beams to be used for a beam management procedure using adaptive learning. The node generally includes means for performing a beam management procedure using the determined one or more beams.
[0036] Certain aspects provide a node configured for wireless communication. The node generally includes a memory. The node generally includes a processor coupled to the memory and configured to use adaptive learning to determine one or more beams to be used for a beam management procedure. The processor and memory are generally configured to perform a beam management procedure using the determined one or more beams.
[0037] Certain aspects provide a computer-readable medium. The computer-readable medium generally stores computer-executable code. The computer-executable code generally includes code for determining one or more beams to be used for a beam management procedure using adaptive learning. The computer-executable code generally includes code for performing a beam management procedure using the determined one or more beams.
[0038] To achieve the foregoing and related ends, one or more aspects include features fully described below and particularly pointed out in the claims. The following description and the accompanying drawings set forth in detail certain illustrative features of one or more aspects. However, these features are only indicative of several of the various ways in which the principles of the various aspects can be employed. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to understand in detail the manner in which the above-stated features of the present disclosure are used, a more particular description of the content briefly summarized above may be made with reference to various aspects, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only certain typical aspects of the present disclosure and are not to be considered limiting of its scope, as the description may admit to other equally effective aspects.
[0041] Figure 1 is a block diagram conceptually illustrating an example telecommunications system in accordance with certain aspects of the present disclosure.
[0042] Figure 2
[0013] Example beam management procedures in accordance with certain aspects of the present disclosure are illustrated.
[0043] Figure 3 Illustrated are example synchronization signal block (SSB) locations within an example half-frame in accordance with certain aspects of the present disclosure.
[0044] Figure 4 Example transmit and receive beams for SSB measurements are illustrated in accordance with certain aspects of the present disclosure.
[0045] Figure 5An example networking environment is illustrated in which a predictive model is used for beam management in accordance with certain aspects of the present disclosure.
[0046] Figure 6 An example reinforcement learning model in accordance with certain aspects of the present disclosure is conceptually illustrated.
[0047] Figure 7 An example deep Q-network (DQN) learning model in accordance with certain aspects of the present disclosure is conceptually illustrated.
[0048] Figure 8 is a flow diagram illustrating example operations for wireless communications by a node in accordance with certain aspects of the present disclosure.
[0049] Fig. 9 is an example call flow diagram illustrating example signaling for beam management using adaptive learning in accordance with certain aspects of the present disclosure.
[0050] Fig.10 is an example call flow diagram illustrating example signaling for a BPL discovery procedure using adaptive learning, in accordance with certain aspects of the present disclosure.
[0051] Fig.11 Illustrated are communications devices that may include various components configured to perform operations for the techniques disclosed herein in accordance with aspects of the present disclosure.
[0052] Fig.12 is a block diagram conceptually illustrating designs of example base stations (BSs) and user equipment (UEs) in accordance with certain aspects of the present disclosure.
[0053] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the figures. It is contemplated that elements disclosed in one aspect may be beneficially utilized in other aspects without specific recitation.
[0054] Detailed Description
[0055] Aspects of the present disclosure provide apparatus (equipment), methods, processing systems, and computer-readable media for beam management using adaptive learning.
[0056] Some systems, such as new radio systems (e.g., 5G NR), support millimeter wave (mmW) communications. In mmW communications, signals used for communication between devices (referred to as mmW signals) may have a high carrier frequency (e.g., 25 GHz or higher, such as within a 30 to 300 GHz band) and may have a wavelength in the range of 1 mm to 10 mm. Based on such characteristics of mmW signals, mmW communications can provide high-speed (e.g., gigabit speed) communications between devices. However, compared to lower frequency signals, mmW signals may experience atmospheric effects and may not propagate well through materials. Therefore, compared to lower frequency signals, mmW signals may experience relatively high path losses (e.g., attenuation or reduction of the power density of the waves corresponding to the mmW signal) as they propagate.
[0057] To overcome path loss, mmW communication systems utilize directional beamforming. Beamforming may involve the use of transmit (TX) beams and / or receive (RX) beams. The TX beam corresponds to a transmitted mmW signal that is directed to have more power in a specific direction relative to other directions, such as toward a receiver. By directing the transmitted mmW signal to the receiver, more energy of the mmW signal is directed to the receiver, thereby overcoming higher path losses. The RX beam corresponds to a technique performed at the receiver to apply gain to a signal received in a specific direction while attenuating signals received in other directions. The use of RX beams also helps to overcome higher path losses, for example by improving the signal-to-noise ratio (SNR) of receiving the desired mmW signal at the receiver. In some aspects, hybrid beamforming (e.g., signal processing in analog and digital domains) may be used.
[0058] Thus, in some aspects, for a particular transmitter to communicate with a particular receiver, the transmitter needs to select a TX beam to use, and the receiver needs to select an RX beam to use. The TX beam and RX beam used for communication are referred to as a beam pairing. In some aspects, the RX and TX beams in a beam pairing are selected so as to provide adequate communication coverage and / or capacity.
[0059] In certain aspects, a beam management procedure may be used to select (e.g., initially select, update select, refine to a narrower beam within a previously selected beam, etc.) a beam pairing. Figure 2-4 Discussed in more detail, the beam management procedure may involve measuring signals using different RX and / or TX beams for reception / transmission and selecting a beam for beam pairing based on the measurements. For example, the beam having the highest measured channel or link quality (e.g., throughput, SNR, etc.) among those measured beams may be selected.
[0060] In some cases, as described below with reference to Figure 2-4 As discussed in more detail, there are a large number of RX and / or TX beams supported at the transmitter and / or receiver, which may mean that there are a large number of measurements that may be performed for beam management procedures. In addition, the communication environment between the transmitter and the receiver may be different at different times, such as due to obstructions (e.g., when a user's hand blocks the TX / RX beam at the transmitter / receiver (e.g., user equipment (UE)), and / or an object blocks the line of sight (LOS) path between the transmitter and the receiver), movement and / or rotation of the transmitter / receiver, etc.
[0061] To account for such factors, in some cases, beam management procedures are based on heuristics. Heuristics-based beam management procedures attempt to predict real-world deployment scenarios for transmitters and receivers and typically update the beam management procedures used by the transmitter and receiver (such as using downloaded software patches) based on problems encountered (or anticipated) over time when the transmitter and receiver communicate. For example, a heuristics-based beam management procedure may measure only certain RX and / or TX beams of a transmitter and receiver based on parameters of the transmitter and / or receiver, rather than all beams.
[0062] In order to further improve the beam management procedure, various aspects of the present disclosure provide the use of adaptive learning as part of the beam management procedure. For example, a UE (and / or BS) acting as a transmitter and / or receiver may use an adaptive learning-based beam management algorithm that adapts over time based on learning. Specifically, the learning may be based on feedback associated with previous beam selections for the UE and / or BS. The feedback may include an indication of the previous beam selection and parameters associated with the previous beam selection. The algorithm may initially be trained based on feedback in a laboratory environment and then updated (e.g., continuously) using feedback while the UE and / or BS is in deployment. In some examples, the algorithm is a beam management algorithm based on deep reinforcement learning that uses machine learning and artificial neural networks to update and apply a predictive model for beam selection during a beam management procedure. In this way, the beam management algorithm based on adaptive learning learns from user behavior (e.g., frequently traversed paths, how the user holds the UE, etc.) and is therefore also personalized for the user.
[0063] The following description provides an example of using adaptive learning as part of a beam management procedure, without limiting the scope, applicability, or examples set forth in the claims. Changes may be made to the functions and arrangements of the elements discussed without departing from the scope of the present disclosure. Various examples may appropriately omit, replace, or add various procedures or components. For example, the described method may be performed in an order different from the order described, and various steps may be added, omitted, or combined. Moreover, the features described with reference to some examples may be combined in some other examples. For example, any number of aspects set forth herein may be used to implement an apparatus or practice method. In addition, the scope of the present disclosure is intended to cover such equipment or methods practiced using other structures, functionalities, or structures and functionalities as a supplement to or in addition to the various aspects of the present disclosure set forth herein. It should be understood that any aspect of the present disclosure disclosed herein may be implemented by one or more elements of the claims. The wording "exemplary" is used herein to mean "used as an example, instance, or explanation". Any aspect described herein as "exemplary" need not be interpreted as superior to or superior to other aspects.
[0064] Figure 1 An example wireless communication network 100 is illustrated in which various aspects of the present disclosure may be performed. For example, the wireless communication network 100 may be a new radio system (e.g., a 5G NR network). The wireless communication network 100 may support mmW communication through beamforming. Nodes (e.g., wireless nodes) such as UE 120a and / or base station (BS) 110a in the wireless communication network 100 may be configured to perform a beam management procedure to select a beam pairing for communicating with another node. For example, the UE 120a and the BS 110a may perform a beam management procedure to determine a receive beam of the UE 120a and a transmit beam of the BS 110a as a beam pairing to be used for communication (e.g., downlink communication), also referred to as a beam pair link (BPL). As will be described in more detail herein, the UE 120a and / or the BS 110a may use a beam management procedure based on adaptive learning. The UE 120a and / or the BS 110a may use adaptive learning to determine one or more beams to be used for the beam management procedure. As Figure 1 As shown in FIG. 1 , UE 120a has a beam selection manager 122. According to one or more aspects described herein, beam selection manager 122 may be configured to use an adaptive learning-based algorithm to determine / select beams to be used for beam management procedures. Figure 1As shown, additionally or alternatively, BS 110a may have a beam selection manager 112. According to various aspects described herein, beam selection manager 112 may be configured to use an adaptive learning algorithm to determine / select beams to be used for beam management procedures. UE 120a and / or BS 110a may then use the determined one or more beams to perform beam management procedures.
[0065] It should be noted that although certain aspects are described with respect to the beam management procedure being performed by a wireless node, certain aspects of this beam management procedure may be performed by other types of nodes, such as nodes connected to a BS through a wired connection.
[0066] like Figure 1 As illustrated in , the wireless communication network 100 may include several BSs 110a-z (each also individually referred to herein as BS 110 or collectively referred to herein as BS 110) and other network entities. BS 110 may communicate with UEs 120a-y (each also individually referred to herein as UE 120 or collectively referred to herein as UE 120) in the wireless communication network 100. Each BS 110 may provide communication coverage for a particular geographic area. In some examples, BSs 110 may be interconnected to each other and / or to one or more other BSs or network nodes (not shown) in the wireless communication network 100 via various types of backhaul interfaces, such as direct physical connections, wireless connections, virtual networks, or the like using any suitable transport networks. Figure 1 In the example shown in FIG. 1 , BSs 110a, 110b, and 110c may be macro BSs for macro cells 102a, 102b, and 102c, respectively. BS 110x may be a pico BS for pico cell 102x. BSs 110y and 110z may be femto BSs for femto cells 102y and 102z, respectively. A BS may support one or more (e.g., three) cells.
[0067] The wireless communication network 100 may also include a relay station. A relay station is a station that receives transmissions of data and / or other information from an upstream station (e.g., a BS or a UE) and sends transmissions of the data and / or other information to a downstream station (e.g., a UE or a BS). A relay station may also be a UE that relays transmissions for other UEs. Figure 1 In the example shown in , a relay station 110r may communicate with a BS 110a and a UE 120r to facilitate communication between the BS 110a and the UE 120r. A relay station may also be referred to as a relay BS, a relay, or the like.
[0068] UEs 120 (eg, 120x, 120y, etc.) may be dispersed throughout wireless communication network 100, and each UE may be stationary or mobile.
[0069] A network controller 130 may be coupled to a set of BSs and provide coordination and control for these BSs. Network controller 130 may communicate with BSs 110 via a backhaul. BSs 110 may also communicate with each other via a wireless or wired backhaul (eg, directly or indirectly).
[0070] In some examples, the wireless communication network 100 (e.g., a 5G NR network) may support mmW communications. As discussed above, such systems using mmW communications may use beamforming to overcome high path loss and may perform beam management procedures to select beams for beamforming.
[0071] The BS beam (e.g., TX or RX) and the UE beam (e.g., the other of TX or RX) form a BPL. Both the BS (e.g., BS 110a) and the UE (e.g., UE 120a) can determine (e.g., find / select) at least one qualified beam to form a communication link. For example, on the downlink, the BS 110a uses a transmit beam to transmit a downlink transmission, and the UE 120a uses a receive beam to receive the downlink transmission. The combination of the transmit beam and the receive beam forms a BPL. The UE 120a and the BS 110a establish at least one BPL to the wireless communication network 100 for the UE 120a. In some examples, multiple BPLs (e.g., a group of BPLs) can be configured for communication between the UE 120a and one or more BSs 110. Different BPLs can be used for different purposes, such as for communicating different channels, for communicating with different BSs, and / or for being used as a fallback BPL in the event of failure of an existing BPL.
[0072] In some examples, for initial cell acquisition, a UE (e.g., UE 120a) may search for the strongest signal corresponding to a cell associated with a BS (e.g., BS 110a) and the associated UE receive beam and BS transmit beam corresponding to a BPL for receiving / transmitting a reference signal. After initial acquisition, UE 120a may perform new cell detection and measurement. For example, UE 120a may measure a primary synchronization signal (PSS) and a secondary synchronization signal (SSS) to detect a new cell. As described below with reference to Figure 3 As discussed in more detail, the PSS / SSS may be transmitted by a BS (e.g., BS 110a) in different synchronization signal blocks (SSBs) across one or more synchronization signal (SS) burst sets. The UE 120a may measure different SSBs within an SS burst set to perform beam management procedures, as further discussed herein.
[0073] In 5G NR, the beam management procedure for determining BPL may be referred to as the P1 procedure. Figure 2An example P1 procedure 202 is illustrated. A BS 210 (e.g., such as BS 110a) may send a measurement request to a UE 220 (e.g., such as UE 120a) and may subsequently transmit one or more signals (sometimes referred to as "P1 signals") to the UE 220 for measurement. In the P1 procedure 202, the BS 210 transmits signals by beamforming in different spatial directions (corresponding to transmit beams 211, 212, ... 217) in each symbol so as to reach several (e.g., most or all) relevant spatial locations of the cell of the BS 210. In this way, the BS 210 transmits signals using different transmit beams in different directions over time. In some examples, SSB is used as the P1 signal. In some examples, a channel state information reference signal (CSI-RS), a demodulation reference signal (DMRS), or another downlink signal may be used as the P1 signal.
[0074] In the P1 procedure 202, in order to successfully receive at least one codeword of the P1 signal, the UE 220 finds (e.g., determines / selects) an appropriate receive beam (221, 222, ... 226). Signals (e.g., SSBs) from multiple BSs can be measured simultaneously for a given signal index (e.g., SSB index) corresponding to a given time period. The UE 220 can apply a different receive beam during each occurrence (e.g., each codeword) of the P1 signal. Once the UE 220 successfully receives the codeword of the P1 signal, the UE 220 and the BS 210 have found the BPL (i.e., the UE RX beam for receiving the P1 signal in the codeword and the BS TX beam for transmitting the P1 signal in the codeword). In some cases, the UE 220 does not search all of its possible UE RX beams until it finds the best UE RX beam because this causes additional delay. UE 220 may instead select an RX beam once it is "good enough" (e.g., has a quality (e.g., SNR) that satisfies a threshold (e.g., a predefined threshold). UE 220 may not know which beam BS 210 used to transmit the P1 signal in a symbol; however, UE 220 may report to BS 210 the time at which it observed the signal. For example, UE 220 may report to BS 210 the index of the symbol at which the P1 signal was successfully received. BS 210 may receive the report and may determine which BS TX beam the BS 210 used at the indicated time. In some examples, UE 220 measures the signal quality of the P1 signal, such as reference signal received power (RSRP) or another signal quality parameter (e.g., SNR, channel flatness, etc.). UE 220 may report the measured signal quality (e.g., RSRP) to BS 210 along with the symbol index. In some cases, UE 220 may report to BS 210 multiple symbol indices corresponding to multiple BS TX beams.
[0075] As part of the beam management procedure, the BPL used between the UE 220 and the BS 110 may be refined / changed. For example, the BPL may be periodically refined to adapt to changing channel conditions, such as attenuation due to movement of the UE 220 or other objects, Doppler spread, etc. The UE 220 may monitor the quality of the BPL (e.g., the BPL found / selected during the P1 procedure and / or the previously refined BPL) to refine the BPL when the quality degrades (e.g., when the BPL quality drops below a threshold or when another BPL has a higher quality). In 5G NR, the beam management procedure for beam refinement of the BPL may be referred to as P2 and P3 procedures to refine the BS beam and UE beam of an individual BPL, respectively.
[0076] Figure 2 An example P2 procedure 204 and a P3 procedure 206 are illustrated. Figure 2 As shown, for P2 protocol 204, BS 210 transmits symbols of a signal using different BS beams (e.g., TX beams 215, 214, 213) that are spatially close to the BS beam of the current BPL. For example, BS 210 transmits signals in different symbols using adjacent TX beams around the TX beam of the current BPL (e.g., beam sweeping). Figure 2 As shown, the TX beam used by BS 210 for P2 procedure 204 may be different from the TX beam used by BS 210 for P1 procedure 202. For example, the TX beam used by BS 210 for P2 procedure 204 may be spaced closer together and / or may be more focused (e.g., narrower) than the TX beam used by BS 210 for P1 procedure 202. During P2 procedure 204, UE 220 keeps its RX beam (e.g., RX beam 224) unchanged. UE 220 may measure the signal quality (e.g., RSRP) of the signal in different symbols and indicate the symbol in which the highest signal quality is measured. Based on the indication, BS 210 may determine the strongest (e.g., best or associated with the highest signal quality) TX beam (i.e., the TX beam used in the indicated symbol). The BPL may be refined accordingly to use the indicated TX beam.
[0077] like Figure 2As shown, for P3 procedure 206, BS 220 maintains a constant TX beam (e.g., the TX beam of the current BPL) and uses the constant TX beam (e.g., TX beam 214) to transmit the codeword of the signal. During P3 procedure 206, UE 220 uses different RX beams (e.g., RX beams 223, 224, 225) in different codewords to scan the signal. For example, UE 220 can use an RX beam adjacent to the RX beam in the current BPL (i.e., the BPL being refined) to perform the sweep. UE 220 can measure the signal quality (e.g., RSRP) of the signal for each RX beam and identify the strongest UE RX beam. UE 220 can use the identified RX beam for the BPL. UE 220 can report the signal quality to BS 210.
[0078] As discussed above, in some examples, SSB measurements may be used for beam management. Figure 3 An example SSB position within an example NR radio frame format 302 is illustrated. The transmission timeline for each of the downlink and uplink may be divided into units of radio frames. Figure 3 As shown, the example 10ms NR radio frame format 302 may include ten 1ms subframes (subframes with indices 0, 1...9). In NR, the basic transmission time interval (TTI) may be referred to as a time slot. In NR, a subframe may contain a variable number of time slots (e.g., 1, 2, 4, 8, 16...time slots), depending on the subcarrier spacing (SCS). NR may support a base SCS of 15KHz, and other SCSs may be defined relative to the base SCS (e.g., 30kHz, 60kHz, 120kHz, 240kHz, etc.). Figure 3 In the example shown, the SCS is 120kHz. Figure 3 As shown, subframe 304 (subframe 0) includes 8 time slots (time slots 0, 1, ... 7) with a duration of 0.125 ms. The symbol and time slot lengths scale with the subcarrier spacing. Each time slot may include a variable number of symbol (e.g., OFDM symbol) periods (e.g., 7 or 14 symbols), depending on the SCS. For Figure 3 For the 120 kHz SCS shown, each of slots 306 (slot 0) and slots 308 (slot 1) includes 14 symbol periods (slots with indices 0, 1 ... 13) having a duration of 0.25 ms.
[0079] In some examples, the SSB may be transmitted up to sixty-four times in up to sixty-four different beam directions. The up to sixty-four transmissions of the SSB are referred to as an SS burst set. The SSBs in an SS burst set may be transmitted in the same frequency region, and the SSBs in different SS burst sets may be transmitted in different frequency regions. Figure 3In the example shown, in subframe 304, SSB is transmitted in every time slot (time slots 0, 1, ... 6). Figure 3 In the example shown, in time slot 306 (time slot 0), SSB 310 is transmitted in codewords 4, 5, 6, 7, and SSB 312 is transmitted in codewords 8, 9, 10, 11 and in time slot 308 (time slot 1), SSB 314 is transmitted in codewords 2, 3, 4, 5, and SSB 316 is transmitted in codewords 6, 7, 8, 9, and so on. The SSB may include PSS, SSS, and a physical broadcast channel (PBCH) of two codes. PSS and SSS may be used by UEs for cell search and acquisition. For example, PSS may provide half-frame timing, SSS may provide control protocol (CP) length and frame timing, and PSS and SSS may provide cell identity. PBCH carries some basic system information, such as downlink system bandwidth, timing information within a radio frame, SS burst set periodicity, system frame number, etc.
[0080] like Figure 4 As shown, SSB can be used to perform measurements using different transmit and receive beams, for example according to beam management procedures such as Figure 2 P1 procedure 202 is shown. Figure 4 An example of a BS 410 (e.g., such as BS 110a) using 4 TX beams and a UE 420 (e.g., such as UE 120a) using 2 RX beams is illustrated. For each SSB, BS 410 transmits the SSB using a different TX beam. Figure 4 As shown, UE 420 may scan its RX beam 422 while BS 410 transmits SSBs 310, 312, 314, 316 that sweep its four TX beams 412, 414, 416, 418, respectively. BPLs may be identified and used for data communications during a period of time, as discussed. Figure 4 As shown, BS 410 uses TX beam 414 and UE 420 uses RX beam 422 for data communications during a period of time. UE 410 may then scan its RX beam 424 while BS 410 transmits SSBs 426, 428 sweeping its TX beams 412, 414, and so on.
[0081] As can be seen, as the number of TX / RX beams increases, the number of scans that the UE needs to perform on each TX beam to scan each of its RX beams can become larger. Power consumption can scale linearly with the number of measured SSBs. Therefore, the time and power overhead associated with beam management can become larger if all beams are actually scanned.
[0082]
[0013] Accordingly, aspects of the present disclosure provide techniques for helping nodes perform measurements of other nodes when using beamforming, such as by using adaptive learning, which can reduce the number of measurements used for beam management procedures and thereby reduce power consumption.
[0083] Example beam management procedure using adaptive learning
[0084] A non-adaptive algorithm is deterministic and varies with its input. If the algorithm is faced with the exact same input at different times, its output will be exactly the same. An adaptive algorithm is one that changes its behavior based on its past experience. This means that different devices using an adaptive algorithm can end up having different algorithms over time.
[0085] According to certain aspects, a beam management procedure may be performed using a beam management algorithm based on adaptive learning. Thus, the beam algorithm changes (e.g., adapts, updates) over time based on new learning. The beam management procedure may be used for initial acquisition, cell discovery after initial acquisition, and / or determining a BPL for the strongest cell detected by the UE. For example, adaptive learning may be used to construct a UE codebook that indicates beams to be used (e.g., measured) for a beam management procedure. In some examples, adaptive learning may be used to select UE receive beams to be used to discover BPLs. Adaptive learning may be used to intelligently select which UE receive beams will be used to measure signals based on training and experience so that fewer beams may be measured while still finding a suitable BPL (e.g., satisfying a threshold signal quality).
[0086] In some examples, adaptive learning-based beam management involves training a model, such as a prediction model. The model can be used during a beam management procedure to select which UE receive beams will be used to measure signals. The model can be trained based on training data (e.g., training information), which can include feedback, such as feedback associated with the beam management procedure. Figure 5 Illustrated is an example networking environment 500 in which a prediction model 524 is used for beam management in accordance with certain aspects of the present disclosure.
[0087] like Figure 5 As shown, networked environment 500 includes node 520, training system 530, and training repository 515 that are communicatively connected via network 505. Node 520 may be a UE (e.g., such as UE 120a in wireless communication network 100) or a BS (e.g., such as BS 110a in wireless communication network 100). Network 505 may be a wireless network, such as wireless communication network 100, which may be a 5G NR network. Although training system 530, node 520, and training repository 515 are in Figure 5Although illustrated as separate components in the illustration, those skilled in the art will recognize that training system 530, nodes 520, and training repository 515 may be implemented on any number of computing systems, as one or more stand-alone systems, or in a distributed environment.
[0088] The training system 530 generally includes a prediction model training manager 532 that uses the training data to generate a prediction model 524 for beam management. The prediction model 524 may be determined based on information in the training repository 515.
[0089] The training repository 515 may include training data acquired before and / or after the node 520 is deployed. The node 520 may be trained in a simulated communication environment (e.g., in a field test, a drive test) before the node 520 is deployed. For example, various beam management procedures (e.g., various selections of UE RX beams for measuring signals) may be tested in various scenarios, such as at different UE speeds, with the UE stationary, with various rotations of the UE, with various BS deployments / geometries, etc., to obtain training information related to the beam management procedures. This information may be stored in the training repository 515. After deployment, the training repository 515 may be updated to include feedback associated with the beam management procedures performed by the node 520. The training repository may also be updated with information from other BSs and UEs, for example, based on experiences learned by other BSs and / or other UEs that may be associated with the beam management procedures performed by these BSs and / or UEs.
[0090] The prediction model training manager 532 can use the information in the training repository 515 to determine a prediction model 524 (e.g., an algorithm) for beam management, such as to select a UE RX beam for measuring a signal. As discussed in more detail herein, the prediction model training manager 532 can use various different types of adaptive learning, such as machine learning, deep information, reinforcement learning, etc. to form the prediction model 524. The training system 530 can adapt (e.g., update / improve) the prediction model 524 over time. For example, when the training repository is updated with new training information (e.g., feedback), the model 524 is updated based on the new learning / experience.
[0091] Training system 530 may be located at node 520, a BS in network 505, or a different entity that determines prediction model 524. If located at a different entity, prediction model 524 is provided to node 520.
[0092] Training repository 515 may be a storage device, such as a memory. Training repository 515 may be located on node 520, training system 530, or another entity in network 505. Training repository 515 may be in cloud storage. Training repository 515 may receive training information from node 520, an entity in network 505 (e.g., a BS or UE in network 505), a cloud, or other source.
[0093] As described above, the node 520 is provided with (or generates, e.g., if the training system 530 is implemented in the node 520) a prediction model. As illustrated, the node 520 may include a beam selection manager 522 configured to use the prediction model 524 for beam management (e.g., such as described above with reference to Figure 2 In some examples, node 520 utilizes prediction model 524 to construct a UE codebook and / or determine / select a beam from the UE codebook to be used for a beam management procedure. Prediction model 524 is updated when training system 530 adapts prediction model 524 with new learning.
[0094] Thus, the beam management algorithm of node 520 (using predictive model 524) is based on adaptive learning in that the algorithm used by node 520 changes over time, even after deployment, based on experience / feedback gained by node 520 in deployment scenarios (and / or through training information also provided by other entities).
[0095] According to certain aspects, adaptive learning may use any appropriate learning algorithm. As described above, the learning algorithm may be used by a training system (e.g., such as training system 530) to train a prediction model (e.g., such as prediction model 524) to obtain an adaptive learning-based beam management algorithm for use by a device (e.g., such as node 520) for a beam management procedure. In some examples, the adaptive learning algorithm is an adaptive machine learning algorithm, an adaptive reinforcement learning algorithm, an adaptive deep learning algorithm, an adaptive continuous infinite learning algorithm, or an adaptive policy optimization reinforcement learning algorithm (e.g., a proximal policy optimization (PPO) algorithm, a policy gradient, a trust region policy optimization (TRPO) algorithm, etc.). In some examples, the adaptive learning algorithm is modeled as a partially observable Markov decision process (POMDP). In some examples, the adaptive learning algorithm is implemented by an artificial neural network (e.g., a deep Q network (DQN) including one or more deep neural networks (DNNs)).
[0096] In some examples, adaptive learning (e.g., used by training system 530) is performed using a neural network. Neural networks can be designed to have various connectivity patterns. In a feedforward network, information is passed from a lower layer to a higher layer, where each neuron in a given layer communicates to a neuron in a higher layer. A hierarchical representation can be constructed in successive layers of a feedforward network. Neural networks can also have reflow or feedback (also known as top-down) connections. In a reflow connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. The reflow architecture can help identify patterns that span more than one block of input data that is sequentially delivered to the neural network. The connection from a neuron in a given layer to a neuron in a lower layer is called a feedback (or top-down) connection. A network with many feedback connections may be helpful when the recognition of high-level concepts can assist in discerning specific low-level features of the input.
[0097] In some examples, adaptive learning (e.g., used by training system 530) is performed using a deep belief network (DBN). DBN is a probabilistic model including multiple layers of hidden nodes. DBN can be used to extract a hierarchical representation of a training data set. DBN can be obtained by stacking multiple layers of restricted Boltzmann machines (RBM). RBM is a class of artificial neural networks that can learn probability distributions on input sets. Since RBM can learn probability distributions without information about which class each input can be classified into, RBM is often used for unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBM of the DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the input and target category from the previous layer) and can be used as a classifier.
[0098] In some examples, adaptive learning (e.g., used by training system 530) is performed using a deep convolutional network (DCN). A DCN is a network in a convolutional network that is configured with additional pooling and normalization layers. DCN has achieved the most advanced performance on many tasks. DCN can be trained using supervised learning, where both input and output targets are known to many paradigms and are used to modify the weights of the network using a gradient descent method. DCN can be a feedforward network. In addition, as described above, the connection from the neurons in the first layer of the DCN to the neuron group in the next higher layer is shared across the neurons in the first layer. The feedforward and shared connections of the DCN can be utilized for fast processing. The computational burden of the DCN can be much smaller than the computational burden of a neural network of similar size, such as one including a reflow or feedback connection.
[0099] An artificial neural network, which may include a group of interconnected artificial neurons (e.g., a neuron model), is a computing device or represents a method performed by a computing device. These neural networks can be used in various applications and / or devices, such as Internet Protocol (IP) cameras, Internet of Things (IoT) devices, autonomous vehicles, and / or service robots. Individual nodes in an artificial neural network can mimic biological neurons by taking input data and performing simple operations on the data. The results of the simple operations performed on the input data are selectively passed to other neurons. Weight values are associated with each vector and node in the network, and these values constrain how the input data is related to the output data. For example, the input data of each node can be multiplied by a corresponding weight value, and the products can be summed. The sum of these products can be adjusted by an optional bias, and an activation function can be applied to the result, thereby generating an output signal or "output activation" of the node. The weight values can be initially determined by the iterative flow of training data in the network (e.g., the weight values are established during the training phase, in which the network learns how to identify a specific category by the typical input data characteristics of each category).
[0100] Different types of artificial neural networks (such as recurrent neural networks (RNNs), multilayer perceptron (MLP) neural networks, convolutional neural networks (CNNs), etc.) can be used to implement adaptive learning (e.g., used by the training system 530). The working principle of RNNs is to save the output of a layer and feed the output back to the input to help predict the results of the layer. In an MLP neural network, data can be fed into an input layer, and one or more hidden layers provide several levels of abstraction of the data. Predictions can then be made for the output layer based on the abstracted data. MLPs can be particularly suitable for classification prediction problems, where inputs are assigned classes or labels. Convolutional neural networks (CNNs) are a type of feedforward artificial neural network. Convolutional neural networks can include a collection of artificial neurons, each of which has a receptive field (e.g., a spatial local area of the input space) and together spell out an input space. Convolutional neural networks have many applications. In particular, CNNs have been widely used in the fields of pattern recognition and classification. In a layered neural network architecture, the output of the first layer of artificial neurons becomes the input of the second layer of artificial neurons, the output of the second layer of artificial neurons becomes the input of the third layer of artificial neurons, and so on. Convolutional neural networks can be trained to recognize feature hierarchies. Computations in convolutional neural network architectures can be distributed across a population of processing nodes, which can be configured in one or more computational chains. These multi-layer architectures can be trained one layer at a time and can be fine-tuned using backpropagation.
[0101] In some examples, when using an adaptive machine learning algorithm, the training system 530 generates a vector from the information in the training repository 515. In some examples, the training repository 515 stores the vector. In some examples, the vector maps one or more features to a label. For example, the features may correspond to various deployment scenario patterns discussed herein, such as UE mobility, speed, rotation, channel conditions, BS deployment / geometry in the network, and the like. The label may correspond to the predicted optimal beam selection (e.g., for RX beams) associated with the features used to perform the beam management procedure. The prediction model training manager 532 may use the vector to train the prediction model 524 for the node 520. As discussed above, the vector may be associated with a weight in the adaptive learning algorithm. When the learning algorithm is adapted (e.g., updated), the weight applied to the vector may also be changed. Thus, when the beam management procedure is executed again under the same features (e.g., under the same set of conditions), the model may give the node 520 different results (e.g., different beam selections).
[0102] Figure 6 An example reinforcement learning model is conceptually explained. Reinforcement learning can be a semi-supervised learning model in machine learning. Reinforcement learning allows an agent 604 (e.g., a node 520 and / or a training system 530) to take actions (e.g., beam selection) based on states observed by an interpreter 602 (e.g., such as a node 520) (e.g., RSPR of SSBs using different beams) and interact with an environment 606 (e.g., a current deployment scenario) to maximize the total reward (e.g., physical downlink shared channel (PDSCH) throughput using the selected beam) that can be observed by the interpreter 602 and fed back to the agent 604 as reinforcement. In some examples, the agent 604 and the interpreter 602 can be implemented as the same or separate component devices that can perform various functions of the node 520, the training system 530, and / or the training repository 515.
[0103] In some examples, reinforcement learning is modeled as a Markov decision process (MDP). An MDP is a discrete, time-stochastic control process. MDPs provide a mathematical framework for modeling decision making in situations where outcomes may be partially random and partially under the control of the decision maker. In an MDP, at each time step, the process is in a state in a finite set of states S, and the decision maker can choose any action in a finite set of actions A available in that state. The process responds at the next time step by randomly moving to a new state and giving the decision maker a corresponding reward, where R α (s, s') is the immediate reward (or expected immediate reward) after transitioning from state s to state s'. The probability that the process moves to its new state is affected by the action chosen, for example, according to the state transition function. The state transition can be represented by P α(s, s′) = Pr(s t+1 =s′|s t =s,α t =α) is given.
[0104] An MDP seeks to find a policy for decision making: a function of π that specifies the action π(s) that the decision maker will choose when in state s. The goal is to choose a policy π that maximizes the reward. For example, a policy that maximizes a cumulative function of the reward, such as the discounted sum. An example function is shown below:
[0105] in
[0106] α t =π(s t ), that is, the action given by the strategy, and γ is the discount factor and satisfies 0≤γ≤1.
[0107] The solution to an MDP is a policy that describes the best action (eg, maximizing the expected discounted reward) for each state in the MDP.
[0108] In some examples, a partially observable MDP (POMDP) is used. A POMDP can be used when the state may not be known when an action is taken, and therefore the probability and / or reward may be unknown. For a POMDP, reinforcement learning can be used. The following function can be defined:
[0109] Q(s,a)=∑ s′ P α (s,s′)(R α (s, s′) + γV(s′)).
[0110] Experience during learning can be based on (s, a) pairs and results s'. For example, where the node was previously in state s and made beam selection a, and achieved throughput s'. In this example, the node can update the array Q directly based on the learned experience. This can be referred to as Q-learning. In some examples, the learning algorithm can be continuous.
[0111] In some examples, for an adaptive learning based beam management algorithm, a state may correspond to the M strongest beam quality measurements (e.g., reference signal received power (RSRP) of SSBs on different beams) in an environment (e.g., the current deployment scenario of the UE), including conditions including UE mobility, BS deployment pattern (e.g., geometry), obstructions, etc. discussed herein. An action may correspond to a beam selection. The reward may be the throughput achieved using the beam selection, such as PDSCH throughput. The reward may be another parameter, such as, for example, spectral efficiency. Thus, by using this MDP at a given time in a given state, a node may adopt a strategy to find a beam selection that specifies maximum throughput. As discussed above, the reward may be discounted. For beam management, the reward may be offset by a penalty as a function of the measured SSBs, such as to optimize for minimum power.
[0112] Return to reference Figure 5 The example network environment in 500 and Figure 6 In the reinforcement learning model 600 in, in some examples, the prediction model training manager 532 or agent 604 can use reinforcement learning to determine a policy (e.g., an MDP solution) for a prediction model (e.g., prediction model 524). The node 520 or agent 604 can take actions based on the policy given by the prediction model (e.g., prediction model 524) in the environment (e.g., environment 606) at a given time for the current state (e.g., observed by the node 520 or interpreter 602), such as beam selection for a beam management procedure. The reinforcement learning algorithm and the prediction model can be updated / adapted based on the learned experience (e.g., which can be stored in the training repository 515).
[0113] The framework of reinforcement learning provides tools to optimally solve POMDPs. This learning changes the weights of a multilayer perceptron (e.g., a neural net) that determines the next action to take. The algorithms in deep ML are encoded in the neural net weights. Thus, changing the weights changes the algorithm.
[0114] In some examples, the adaptive learning based beam management uses an adaptive deep learning algorithm. The adaptive deep learning algorithm can be a deep Q network (DQN) implemented by a neural network. Figure 7 An example DQN learning model 700 is conceptually illustrated in accordance with certain aspects of the present disclosure. Figure 7 As shown, agent 706 (e.g., such as agent 604 or node 520) includes an artificial neural network, such as in Figure 7708. For a current environment 702 (e.g., such as environment 606), which may be a real deployment scenario involving a UE (e.g., UE 120a) and a BS (e.g., BS 110a) and various conditions described herein, an agent 706 observes a state 704(s). For example, the observed state may be the M strongest RSRPs corresponding to SSBs measured using different beams for a beam management procedure.
[0115] In some examples, the adaptive learning algorithm is modeled as a POMDP through reinforcement learning. POMDP can be used when the state may not be known when taking an action, and therefore the probability and / or reward may be unknown. For POMDP, reinforcement learning can be used. The Q array can be defined as:
[0116] Q i+1 (s, a) = E{r+γmaxQ i (s′, a′)|s, a}.
[0117] like Figure 7 As shown, given a state 704s (e.g., RSRP) and a possible action a are input to a DNN 708, the DNN may execute an algorithm to output a value (e.g., parameter θ) according to the possible action a, so as to determine a strategy (e.g., π) based on the maximum value. θ (s, a)) The policy and corresponding actions are taken and applied to the environment. For example, agent 706 makes a beam selection and then uses the selected beam in environment 702. Figure 7 As shown, the reward for the action is fed back to the agent 706 to update the algorithm. For example, the throughput achieved by the selected beam can be fed back. Based on the feedback, the agent 706 updates the DNN 708 (e.g., by changing the weights associated with the vector).
[0118] According to certain aspects, adaptive learning-based beam management allows for continuous infinite learning. In some examples, learning can be augmented by federated learning. For example, while some machine learning methods use centralized training data on a single machine or in a data center, with federated learning, learning can be collaborative, involving multiple devices to form a predictive model. With federated learning, model training can be done on the device through collaborative learning from multiple devices. For example, back to reference Figure 5-7 , node 520, agent 604, and agent 706 may receive training information and / or updated trained models from a variety of different devices.
[0119] In an illustrative example, beam management algorithms for multiple different UEs may be trained in multiple different operating scenarios, for example using deep reinforcement learning. The outputs from the training of different UEs may be combined to train a beam management algorithm for a UE. Once the beam management algorithm is trained, the algorithm may continue to learn based on actual deployment scenarios. As discussed above, the state may be the best M RSRP metrics at the current time; the reward may be the measured PDSCH throughput for the current best beam pair; and the action may be a selection of which beam pair / pairs to measure.
[0120] According to certain aspects, adaptive learning-based beam management allows user personalization as well as design robustness. In some examples, adaptive learning-based beam management may be optimized. For example, when a user (e.g., such as node 520) visits / traverses a path (e.g., an environment), the adaptive algorithm learns and optimizes for the environment. In addition, different BS vendors may have different beam management implementations, such as how to transmit SSB. For example, some BS vendors transmit many narrow TX beams that will also be used as data beams; while other vendors transmit a few wide beams and use beam refinement (e.g., P2 and / or P3 procedures) to narrow and track data beams. In some examples, adaptive learning-based beam management may be optimized for a specific beam management implementation for a vendor. In some examples, adaptive learning-based beam management may be optimized for a user, such as the way the user holds / uses the UE affects the possible blocking of the UE's beam.
[0121] Figure 8 800 for wireless communication in accordance with certain aspects of the present disclosure. Operations 800 may be performed, for example, by a node (e.g., such as node 520, which may be a wireless node, such as BS 110a or UE 120a in wireless communication network 100). Operations 800 may be implemented as a processor (e.g., Fig.12 In addition, signal transmission and reception by the node in operation 800 may be performed by one or more antennas (e.g., Fig.12 In some aspects, the transmission and / or reception of signals by a node may be implemented via a bus interface of one or more processors (e.g., controller / processor 1240, 1280) that obtain and / or output signals.
[0122] Operations 800 may begin, at 805, by determining, using adaptive learning, one or more beams to be used for a beam management procedure.
[0123] At 810, the node performs a beam management procedure using the determined one or more beams.
[0124] According to certain aspects, adaptive learning uses an adaptive learning algorithm. The adaptive learning algorithm may be updated (e.g., adapted) based on feedback and / or training information. The node may perform another beam management procedure using the updated adaptive learning algorithm. The feedback may be feedback associated with a beam management procedure. For example, after performing a beam management procedure using the determined one or more beams, the node may receive feedback about the throughput achieved, and the beam management algorithm may be updated based on the feedback. In some examples, the feedback may be associated with beam management performed by different devices, such as different nodes.
[0125] Fig. 9 is an example call flow diagram illustrating example signaling 900 for beam management using adaptive learning in accordance with certain aspects of the present disclosure. Fig. 9 As shown, at 908, UE 902 (e.g., such as UE 120a) may have an initial learning algorithm (e.g., including a prediction model). In some examples, UE 902 may train the initial learning algorithm or the learning algorithm may be trained and then provided to UE 902. At 910, UE 902 performs a beam management procedure (e.g., such as P1 procedure 202) with one or more BSs 904. For example, UE 902 may use an adaptive learning algorithm to determine a beam to use and / or measure. At 912, UE 902 receives additional training information and / or feedback. For example, UE 902 may receive feedback from BS 904 (e.g., such as BS 110a) about the beam management procedure performed at 910, such as a PDSCH throughput achieved using a selected beam. Additionally or alternatively, UE 902 may receive additional training information from BS 904 and / or another UE 906. At 914, UE 902 determines an updated adaptive learning algorithm based on the additional training information and / or feedback. At 916, UE 902 may perform another beam management with BS 904 (or another BS) through the updated adaptive learning algorithm.
[0126] In some examples, the training information (and / or feedback) includes training information acquired by deploying one or more UEs in one or more simulated communication environments prior to network deployment of the one or more UEs; training information acquired through feedback previously received when the one or more UEs were deployed in one or more communication environments (e.g., based on measurements and / or beam management procedures performed by the UE); training information from the network, one or more UEs and / or the cloud; and / or training information received when the node is online and / or idle.
[0127] In some examples, using the adaptive learning algorithm at 805 includes the node outputting an action based on one or more inputs; wherein feedback is associated with the action; and updating the adaptive learning algorithm based on the feedback includes adjusting one or more weights applied to the one or more inputs.
[0128] In some examples, the adaptive learning algorithm used by the node at 805 includes an adaptive machine learning algorithm; an adaptive reinforcement learning algorithm; an adaptive deep learning algorithm; an adaptive continuous infinite learning algorithm; and / or an adaptive policy optimization reinforcement learning algorithm. Figure 6-7 As discussed, the adaptive learning algorithm may be modeled as a POMDP. The adaptive learning algorithm may be implemented by an artificial neural network. In some examples, the artificial neural network may be a DQN that includes one or more DNNs. Using adaptive learning to determine one or more beams may include passing state parameters and action parameters through one or more DNNs; for each state parameter, outputting a value for each action parameter; and selecting an action associated with a maximum output value. Updating the adaptive learning algorithm may include adjusting one or more weights associated with one or more neuron connections in the artificial neural network.
[0129] In some examples, using adaptive learning at 805 to determine one or more beams to be used for the beam management procedure includes determining one or more beams to be included in a codebook based on the adaptive learning and selecting one or more beams to be used for the beam management procedure from the codebook.
[0130] In some examples, using adaptive learning to determine at 805 one or more beams to be used for the beam management procedure includes using adaptive learning to select one or more beams to be used for the beam management procedure from a codebook.
[0131] In some examples, adaptive learning is used to select the BPL.
[0132] In some examples, adaptive learning uses state parameters associated with channel measurements, reward parameters associated with received signal throughput or spectral efficiency, and action parameters associated with selection of beam pairs corresponding to the channel measurements. In some examples, the channel measurements include RSRP; spectral efficiency, channel flatness, and / or signal-to-noise ratio (SNR). In some examples, the received signal is a PDSCH transmission.
[0133] In some examples, the reward parameter is deducted by a penalty amount. In some examples, the penalty amount depends on the number of beams measured for beam management procedures (e.g., beams used for transmission and / or reception of SSB). In some examples, the penalty amount depends on the amount of power consumption associated with the beam management procedure.
[0134] In some examples, performing a beam management procedure using the determined one or more beams at 810 includes measuring a channel based on an SSB transmission from a BS, the SSB transmission being associated with a plurality of different transmit beams of the BS, using the determined one or more beams; and selecting one or more BPLs associated with the channel measurements, which are channel measurements above a channel measurement threshold and / or are one or more strongest channel measurements of all channel measurements associated with the measured SSB transmissions. In some examples, the determined one or more beams are a subset of available receive beams. In some examples, the node receives a PDSCH using one of the one or more selected BPLs; determines a throughput associated with the PDSCH; updates an adaptive learning algorithm based on the determined throughput; and uses the updated adaptive learning algorithm to determine another one or more beams to be used to perform another beam management procedure for selecting another one or more BPLs.
[0135] Fig.10 is an example call flow diagram illustrating example signaling 1000 for a BPL discovery procedure (eg, such as P1 procedure 202) using adaptive learning in accordance with certain aspects of the present disclosure. Fig.10 As shown, at 1008, UE 1002 (e.g., such as UE 120a) may perform initial training in a simulation environment prior to deployment at 1010. The initial training at 1008 may train an initial learning algorithm (e.g., including a prediction model) at UE 1002. At 1010, UE 1002 may be deployed in a network having at least one BS 1004 (e.g., such as BS 110a). UE 1002 may perform beam management procedures (e.g., such as P1 procedures 202) with one or more BSs 1004 in the network. For example, as shown in FIG. Fig.10As shown, at 1012, the UE 1002 may use an adaptive learning algorithm to select a beam or RX / TX beam pair. At 1016, the UE 1002 uses the beam(s) selected at 1012 to measure the SSB transmission(s) received from the BS 1004 at 1014. At 1018, the UE 1002 reports the measurement and / or BPL selection to the BS 1004. Then, at 1020, the BS 1004 transmits the PDSCH to the UE 1002 using the BPL indicated by the UE 1002 (or selected based on the measurement reported by the UE 1002). At 1022, the UE 1002 may determine the PDSCH throughput. The PDSCH throughput may serve as feedback or reinforcement for the adaptive learning. At 1026, the UE 1002 updates the adaptive learning algorithm based on the feedback. Optionally, UE 1002 may receive additional training information and / or feedback from BS 1004 and / or another UE 1006 (e.g., UE 2) that UE 1002 may use to update the adaptive learning algorithm at 1026. UE 1002 may then perform another beam management with BS 1004 (or another BS) using the updated adaptive learning algorithm.
[0136] Fig.11 The description may include operations configured to perform the techniques disclosed herein (such as, Figure 8 1100 includes a communication device 1100 including various components (e.g., corresponding to means-plus-function components) of the operations illustrated in the foregoing. The communication device 1100 includes a processing system 1102 coupled to a transceiver 1108. The transceiver 1108 is configured to transmit and receive signals (such as various signals as described herein) for the communication device 1100 via an antenna 1110. The processing system 1102 may be configured to perform processing functions for the communication device 1100, including processing signals received and / or to be transmitted by the communication device 1100.
[0137] The processing system 1102 includes a processor 1104 coupled to a computer readable medium / memory 1112 via a bus 1106. In some aspects, the computer readable medium / memory 1112 is configured to store programs that, when executed by the processor 1104, cause the processor 1104 to execute Figure 81104 includes circuitry 1118 for determining one or more beams to be used for a beam management procedure using adaptive learning; and circuitry 1120 for performing a beam management procedure using the determined one or more beams.
[0138] In some examples, the communication device 1100 may include a system on a chip (SOC) (not shown) that may include a central processing unit (CPU) or a multi-core CPU configured to perform adaptive learning-based beam management according to certain aspects of the present disclosure. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., a neural network with weights), delays, frequency slot information, and task information may be stored in a memory block associated with a neural processing unit (NPU), in a memory block associated with the CPU, in a memory block associated with a digital signal processor (DSP), in a different memory block, or may be distributed across multiple memory blocks. Instructions executed at the CPU may be loaded from a program memory associated with the CPU or may be loaded from a different memory block.
[0139] In some examples, the adaptive learning-based beam management described herein may allow for improved P1 procedures by adaptively updating the beam management algorithm so that beam selection may be refined to more intelligently select beams to measure based on learning. Thus, the UE may find the BPL while measuring fewer beams.
[0140] Each method disclosed herein includes one or more steps or actions for implementing the method. These method steps and / or actions can be interchangeable with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions can be changed without departing from the scope of the claims.
[0141] As used herein, a phrase referring to "at least one" of a list of items refers to any combination of those items, including single members. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, ab, ac, bc, and abc, as well as any combination with multiple identical elements (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other ordering of a, b, and c).
[0142] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, a database, or another data structure), ascertaining, and the like. Also, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. Also, "determining" may include resolving, selecting, choosing, establishing, and the like.
[0143] Fig.12 1 and 120a (eg, in the example embodiment of the present invention) that can be used to implement various aspects of the present disclosure. Figure 1 For example, antenna 1252, processors 1266, 1258, 1264, and / or controller / processor 1280 of UE 120a and / or antenna 1234, processors 1220, 1230, 1238, and / or controller / processor 1240 of BS 110a may be used to perform the various techniques and methods described herein. Fig.12 As shown, the controller / processor 1280 of the UE 120a has a beam selection manager 1281, which can be configured to use adaptive learning to determine, for example, a beam to be used for a beam management procedure in accordance with various aspects described herein. Fig.12 As shown, additionally or alternatively, the controller / processor 1240 of BS 110a may have a beam selection manager 1241, which may be configured to determine beams using adaptive learning in accordance with aspects described herein.
[0144] At BS 110a, a transmit processor 1220 may receive data from a data source 1212 and control information from a controller / processor 1240. The control information may be for a physical broadcast channel (PBCH), a physical control format indicator channel (PCFICH), a physical hybrid ARQ indicator channel (PHICH), a physical downlink control channel (PDCCH), a group common PDCCH (GC PDCCH), etc. The data may be for a physical downlink shared channel (PDSCH), etc. The processor 1220 may process (e.g., encode and symbol map) the data and the control information to obtain data symbols and control symbols, respectively. The transmit processor 1220 may also generate reference symbols (such as for a primary synchronization signal (PSS), a secondary synchronization signal (SSS), and a cell-specific reference signal (CRS)). A transmit (TX) multiple-input multiple-output (MIMO) processor 1230 may perform spatial processing (e.g., precoding) on data symbols, control symbols, and / or reference symbols, where applicable, and may provide an output symbol stream to modulators (MODs) 1232a-1232t. Each modulator 1232 may process a respective output symbol stream (e.g., for OFDM, etc.) to obtain an output sample stream. Each modulator may further process (e.g., convert to analog, amplify, filter, and up-convert) the output sample stream to obtain a downlink signal. The downlink signals from modulators 1232a-1232t may be transmitted via antennas 1234a-1234t, respectively.
[0145] At UE 120a, antennas 1252a-1252r may receive downlink signals from BS 110a and may provide received signals to demodulators (DEMODs) in transceivers 1254a-1254r, respectively. Each demodulator 1254 may condition (e.g., filter, amplify, downconvert, and digitize) a respective received signal to obtain input samples. Each demodulator may further process the input samples (e.g., for OFDM, etc.) to obtain received symbols. A MIMO detector 1256 may obtain received symbols from all demodulators 1254a-1254r, perform MIMO detection on the received symbols where applicable, and provide detected symbols. A receive processor 1258 may process (e.g., demodulate, deinterleave, and decode) the detected symbols, provide decoded data for UE 120a to a data sink 1260, and provide decoded control information to a controller / processor 1280.
[0146] On the uplink, at the UE 120a, a transmit processor 1264 may receive and process data from a data source 1262 (e.g., for a physical uplink shared channel (PUSCH)) and control information from a controller / processor 1280 (e.g., for a physical uplink control channel (PUCCH)). The transmit processor 1264 may also generate reference symbols for a reference signal (e.g., a sounding reference signal (SRS)). The symbols from the transmit processor 1264 may be precoded by a TX MIMO processor 1266, if applicable, further processed by a demodulator 1254a-1254r in the transceiver (e.g., for SC-FDM, etc.), and transmitted to the base station 110. At BS 110a, the uplink signal from UE 120a may be received by antenna 1234, processed by modulator 1232, detected by MIMO detector 1236 if applicable, and further processed by receive processor 1238 to obtain decoded data and control information sent by UE 120a. Receive processor 1238 may provide decoded data to data sink 1239 and decoded control information to controller / processor 1240.
[0147] Controllers / processors 1240 and 1280 may direct the operation at BS 110a and UE 120a, respectively. Controller / processor 1240 and / or other processors and modules at BS 110a may perform or direct the execution of processes for the techniques described herein. Memories 1242 and 1282 may store data and program codes for BS 110a and UE 120a, respectively. Scheduler 1244 may schedule UEs for data transmission on the downlink and / or uplink.
[0148] The technology described herein can be used for various wireless communication technologies, such as 3GPP Long Term Evolution (LTE), Advanced LTE (LTE-A), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single Carrier Frequency Division Multiple Access (SC-FDMA), Time Division Synchronous Code Division Multiple Access (TD-SCDMA), and other networks. The terms "network" and "system" are often used interchangeably. A CDMA network can implement radio technologies such as Universal Terrestrial Radio Access (UTRA), cdma2000, etc. UTRA includes Wideband CDMA (WCDMA) and other variants of CDMA. cdma2000 covers IS-2000, IS-95, and IS-856 standards. A TDMA network can implement radio technologies such as Global System for Mobile Communications (GSM). OFDMA networks can implement radio technologies such as NR (e.g., 5G RA), Evolved UTRA (E-UTRA), Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), IEEE 802.20, Flash-OFDMA, etc. UTRA and E-UTRA are parts of Universal Mobile Telecommunications System (UMTS). LTE and LTE-A are versions of UMTS that use E-UTRA. UTRA, E-UTRA, UMTS, LTE, LTE-A, and GSM are described in documents from an organization named "3rd Generation Partnership Project" (3GPP). cdma2000 and UMB are described in documents from an organization named "3rd Generation Partnership Project 2" (3GPP2). NR is an emerging wireless communication technology being developed in collaboration with the 5G Technology Forum (5GTF). NR access (e.g., 5G NR) can support various wireless communication services, such as mmW. NR can utilize OFDM with CP on uplink and downlink and includes support for half-duplex operation using TDD. Beamforming can be supported and the beam direction can be dynamically configured. MIMO transmission with precoding can also be supported. In some examples, the MIMO configuration in the DL can support up to 8 transmit antennas (with multi-layer DL transmission of up to 8 streams) and up to 2 streams per UE. In some examples, multi-layer transmission of up to 2 streams per UE can be supported.
[0149] In 3GPP, the term "cell" may refer to the coverage area of a Node B (NB) and / or a NB subsystem serving the coverage area, depending on the context in which the term is used. In NR systems, the terms "cell", BS, next generation Node B (gNB or g Node B), access point (AP), distributed unit (DU), carrier, or transmit receive point (TRP) may be used interchangeably. In some examples, a cell may not necessarily be stationary, and the geographic area of a cell may move depending on the location of a mobile BS.
[0150] UE may also be referred to as a mobile station, terminal, access terminal, subscriber unit, station, customer premises equipment (CPE), cellular phone, smart phone, personal digital assistant (PDA), wireless modem, wireless communication device, handheld device, laptop computer, cordless phone, wireless local loop (WLL) station, tablet computer, camera, gaming device, netbook, smartbook, ultrabook, appliance, medical device or medical equipment, biometric sensor / device, wearable device (such as smart watch, smart clothing, smart glasses, smart wristband, smart jewelry (e.g., smart ring, smart bracelet, etc.)), entertainment device (e.g., music device, video device, satellite radio, etc.), transportation component or sensor, smart meter / sensor, industrial manufacturing equipment, global positioning system device, or any other suitable device configured to communicate via wireless or wired medium. Some UEs may be considered machine type communication (MTC) devices or evolved MTC (eMTC) devices. MTC and eMTC UEs include, for example, robots, drones, remote devices, sensors, meters, monitors, location tags, etc., which can communicate with a BS, another device (e.g., a remote device), or some other entity. Nodes such as wireless nodes can provide connectivity for or to a network (e.g., a wide area network (such as the Internet) or a cellular network), for example, via a wired or wireless communication link. Some UEs may be considered Internet of Things (IoT) devices, which may be narrowband IoT (NB-IoT) devices.
[0151] The techniques described herein may be used for the wireless networks and radio technologies mentioned above as well as other wireless networks and radio technologies. For clarity, although various aspects may be described herein using terms commonly associated with 3G, 4G and / or 5G wireless technologies, various aspects of the present disclosure may be applied in communication systems based on other generations.
[0152] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be easily understood by those skilled in the art, and the universal principles defined herein can be applied to other aspects. Thus, the claims are not intended to be limited to the various aspects shown herein, but should be granted the full scope consistent with the language of the claims, wherein the singular reference to the element is not intended to mean "there is and only one" (unless specifically stated) but "one or more". Unless otherwise specifically stated, the term "some / some" refers to one or more. The elements of the various aspects described throughout this disclosure are all structural and functional equivalents currently or hereafter known to ordinary technicians in the art and are expressly incorporated herein by reference, and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be donated to the public, regardless of whether such disclosure is explicitly recorded in the claims. Any element of the claim should not be interpreted under the provisions of 35 USC§112(f), unless the element is explicitly stated using the phrase "device for..." or in the case of a method claim, the element is stated using the phrase "step for..."
[0153] The various operations of the methods described above may be performed by any suitable device capable of performing the corresponding functions. These devices may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, where there are operations illustrated in the accompanying drawings, these operations may have corresponding paired device-plus-function components with similar numbers.
[0154] The various illustrative logical blocks, modules, and circuits described in conjunction with the present disclosure may be implemented or executed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0155] If implemented in hardware, an example hardware configuration may include a processing system in a node. The processing system may be implemented using a bus architecture. Depending on the specific application of the processing system and the overall design constraints, the bus may include any number of interconnecting buses and bridges. The bus may link together various circuits including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect a network adapter, etc., to the processing system via the bus. The network adapter may be used to implement the signal processing functions of the PHY layer. In UE 120a (see Figure 1 ), a user interface (e.g., a keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits, such as timing sources, peripherals, voltage regulators, power management circuits, and the like, which are well known in the art and will not be described further. The processor may be implemented with one or more general and / or special purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuit systems capable of executing software. Those skilled in the art will recognize how to best implement the functionality described with respect to the processing system, depending on the specific application and the overall design constraints imposed on the overall system.
[0156] If implemented in software, each function may be stored on or transmitted by a computer-readable medium as one or more instructions or codes. Software should be broadly interpreted as meaning instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or other. Computer-readable media include both computer storage media and communication media, which include any media that facilitate the transfer of computer programs from one place to another. The processor may be responsible for managing the bus and general processing, including executing software modules stored on a machine-readable storage medium. A computer-readable storage medium may be coupled to a processor so that the processor can read and write information from / to the storage medium. In an alternative, a storage medium may be integrated into a processor. As an example, a machine-readable medium may include a transmission line, a carrier modulated by data, and / or a computer-readable storage medium having instructions stored thereon separated from a node, all of which may be accessed by a processor through a bus interface. Alternatively or additionally, a machine-readable medium or any part thereof may be integrated into a processor, such as a cache and / or a general register file, which may be the case. As an example, examples of machine-readable storage media may include RAM (random access memory), flash memory, ROM (read only memory), PROM (programmable read only memory), EPROM (erasable programmable read only memory), EEPROM (electrically erasable programmable read only memory), registers, magnetic disks, optical disks, hard drives, or any other suitable storage media, or any combination thereof. Machine-readable media may be implemented in a computer program product.
[0157] A software module may include a single instruction, or many instructions, and may be distributed over several different code segments, distributed between different programs, and distributed across multiple storage media. A computer-readable medium may include several software modules. These software modules include instructions that cause a processing system to perform various functions when executed by an apparatus such as a processor. These software modules may include a transmission module and a receiving module. Each software module may reside in a single storage device or may be distributed across multiple storage devices. As an example, when a triggering event occurs, a software module may be loaded into a RAM from a hard drive. During the execution of a software module, a processor may load some instructions into a cache to increase access speed. One or more cache lines may then be loaded into a general register file for execution by the processor. When describing the functionality of a software module as described below, it will be understood that such functionality is implemented by the processor when the processor executes instructions from the software module.
[0158] Any connection is also properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared (IR), radio, and microwave), then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray Disks, where disks often reproduce data magnetically, and discs reproduce data optically with lasers. Thus, in some aspects, computer-readable media may include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.
[0159] Thus, certain aspects may include a computer program product for performing the operations presented herein. For example, such a computer program product may include a computer-readable medium having stored (and / or encoded) instructions thereon, which instructions can be executed by one or more processors to perform the operations described herein, such as for performing the operations described herein and in Figure 8 Instructions for the operations explained in .
[0160] In addition, it should be appreciated that modules and / or other appropriate means for performing the methods and techniques described herein may be downloaded and / or otherwise obtained by a user terminal and / or base station where applicable. For example, such a device can be coupled to a server to facilitate the transfer of means for performing the methods described herein. Alternatively, the various methods described herein can be provided via a storage device (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or a floppy disk, etc.) so that once the storage device is coupled to or provided to a user terminal and / or base station, the device can obtain the various methods. In addition, any other suitable technology suitable for providing the methods and techniques described herein to a device may be utilized.
[0161] It will be understood that the claims are not limited to the precise configuration and components illustrated above. Various changes, substitutions and variations may be made in the arrangement, operation and details of the methods and apparatus described above without departing from the scope of the claims.
Claims
1. A method for wireless communication by a node, comprising: using adaptive learning to determine one or more beams to be used for a beam management procedure; as well as The beam management procedure is performed using the determined one or more beams, wherein the adaptive learning uses a state parameter associated with a channel measurement, a reward parameter associated with a received signal throughput or spectral efficiency, and an action parameter associated with a selection of a beam pair corresponding to the channel measurement.
2. The method of claim 1, further comprising: updating an adaptive learning algorithm used for the adaptive learning based on feedback or training information or a combination thereof; as well as Another beam management procedure is performed using the updated adaptive learning algorithm.
3. The method of claim 2, wherein the feedback comprises feedback associated with the beam management procedure.
4. The method of claim 2, wherein the training information comprises: training information obtained by deploying one or more user equipment UEs in one or more simulated communication environments prior to network deployment of the one or more user equipment UEs; or training information obtained through feedback previously received when the one or more UEs were deployed in one or more communication environments; or training information from at least one of the network, one or more UEs, or a cloud; or training information received while the node is at least one of online or idle; or Its combination.
5. A method as claimed in claim 4, wherein the training information includes training information received from one or more UEs different from the node after the node is deployed, wherein the training information includes information associated with beam measurements performed by the one or more UEs, or feedback associated with one or more beam management procedures performed by the one or more UEs, or a combination thereof. The method of claim 4 , wherein the node comprises a UE.
7. The method of claim 2, wherein: Using the adaptive learning algorithm includes outputting an action based on one or more inputs; The feedback is associated with the action; and Updating the adaptive learning algorithm based on the feedback includes adjusting one or more weights applied to the one or more inputs.
8. The method of claim 2, wherein the adaptive learning algorithm comprises an adaptive machine learning algorithm; Adaptive reinforcement learning algorithm; Adaptive deep learning algorithms; An adaptive continuous infinite learning algorithm; or an adaptive policy optimization reinforcement learning algorithm, or a combination thereof.
9. The method of claim 2, wherein the adaptive learning algorithm is modeled as a partially observable Markov decision process (POMDP).
10. The method of claim 2, wherein the adaptive learning algorithm is implemented by an artificial neural network.
11. The method of claim 10, wherein: The artificial neural network comprises a deep Q network DQN comprising one or more deep neural networks DNN; and Using the adaptive learning to determine the one or more beams comprises: passing one or more state parameters and one or more action parameters through the one or more DNNs; For each state parameter, output the value of each action parameter; and Select the action associated with the maximum output value.
12. The method of claim 10, wherein updating the adaptive learning algorithm comprises adjusting one or more weights associated with one or more neuronal connections in the artificial neural network.
13. The method of claim 1 , wherein using the adaptive learning to determine the one or more beams to be used for the beam management procedure comprises: determining one or more beams to include in a codebook based on the adaptive learning; as well as One or more beams are selected from the codebook to be used for the beam management procedure.
14. The method of claim 1, wherein determining the one or more beams to be used for the beam management procedure comprises using the adaptive learning to select one or more beams to be used for the beam management procedure from a codebook.
15. The method of claim 1, wherein the channel measurement comprises reference signal received power (RSRP); spectrum efficiency, channel flatness, or signal-to-noise ratio (SNR); or a combination thereof.
16. The method of claim 1, wherein the received signal comprises a Physical Downlink Shared Channel (PDSCH) transmission. The method of claim 1 , wherein the reward parameter is deducted by a penalty amount.
18. The method of claim 17, wherein the penalty amount depends on a number of one or more beams measured for use in the beam management procedure.
19. The method of claim 17, wherein the amount of penalty depends on an amount of power consumption associated with the beam management procedure.
20. The method of claim 1, wherein the one or more beams are used for transmission, reception, or both of one or more synchronization signal blocks (SSBs).
21. The method of claim 1 , wherein performing the beam management procedure using the determined one or more beams comprises: Using the determined one or more beams, measuring the channel based on a synchronization signal block (SSB) transmission from a base station (BS), the SSB transmission being associated with one or more transmit beams of the BS; as well as One or more beam-pair links (BPLs) are selected, the one or more BPLs being associated with one or more channel measurements above a channel measurement threshold; or one or more strongest channel measurements among all channel measurements associated with the SSB transmission; or a combination thereof.
22. The method of claim 21, wherein the determined one or more beams comprise a subset of available receive beams.
23. The method of claim 21, further comprising: receiving a physical downlink shared channel (PDSCH) using one of the one or more BPLs; determining a throughput associated with the PDSCH; updating the adaptive learning algorithm based on the determined throughput; as well as The updated adaptive learning algorithm is used to determine another one or more beams to be used to perform another beam management procedure to select another one or more BPLs.
24. A node configured for wireless communication, comprising: means for using adaptive learning to determine one or more beams to be used for a beam management procedure; as well as Means for performing the beam management procedure using the determined one or more beams, wherein the adaptive learning uses state parameters associated with channel measurements, reward parameters associated with received signal throughput or spectral efficiency, and action parameters associated with selection of beam pairs corresponding to the channel measurements.
25. The node of claim 24, further comprising: means for updating an adaptive learning algorithm used for said adaptive learning based on feedback or training information or a combination thereof; as well as Means for performing another beam management procedure using the updated adaptive learning algorithm.
26. The node of claim 25, wherein the feedback comprises feedback associated with the beam management procedure.
27. The node of claim 25, wherein the training information comprises: training information obtained by deploying one or more user equipment UEs in one or more simulated communication environments prior to network deployment of the one or more user equipment UEs; or training information obtained through feedback previously received when the one or more UEs were deployed in one or more communication environments; or training information from at least one of the network, one or more UEs, or a cloud; or training information received while the node is at least one of online or idle; or Its combination.
28. A node as described in claim 27, wherein the training information includes training information received from one or more UEs different from the node after the node is deployed, wherein the training information includes information associated with beam measurements performed by the one or more UEs, or feedback associated with one or more beam management procedures performed by the one or more UEs, or a combination thereof.
29. The node of claim 27, wherein the node comprises a UE.
30. The node of claim 25, wherein: Using the adaptive learning algorithm includes outputting an action based on one or more inputs; The feedback is associated with the action; and Updating the adaptive learning algorithm based on the feedback includes adjusting one or more weights applied to the one or more inputs.
31. The node of claim 25, wherein the adaptive learning algorithm comprises an adaptive machine learning algorithm; Adaptive reinforcement learning algorithm; Adaptive deep learning algorithms; An adaptive continuous infinite learning algorithm; or an adaptive policy optimization reinforcement learning algorithm, or a combination thereof.
32. The node of claim 25, wherein the adaptive learning algorithm is modeled as a Partially Observable Markov Decision Process (POMDP).
33. The node of claim 25, wherein the adaptive learning algorithm is implemented by an artificial neural network.
34. A node as claimed in claim 33, wherein: The artificial neural network includes a deep Q network DQN including one or more deep neural networks DNN; and Means for using the adaptive learning to determine one or more beams to be used for a beam management procedure include: means for communicating one or more state parameters and one or more action parameters through the one or more DNNs; means for outputting, for each state parameter, the value of each action parameter; and Means for selecting an action associated with a maximum output value.
35. A node as claimed in claim 33, wherein the means for updating the adaptive learning algorithm used for the adaptive learning based on feedback or training information or a combination thereof includes means for adjusting one or more weights associated with one or more neuronal connections in the artificial neural network.
36. The node of claim 24, wherein the means for using the adaptive learning to determine the one or more beams to be used for the beam management procedure comprises: means for determining one or more beams to include in a codebook based on the adaptive learning; as well as Means for selecting one or more beams from the codebook to be used for the beam management procedure.
37. The node of claim 24, wherein the means for determining, using adaptive learning, one or more beams to be used for a beam management procedure comprises means for selecting, from a codebook, one or more beams to be used for the beam management procedure using the adaptive learning.
38. The node of claim 24, wherein the channel measurement comprises reference signal received power (RSRP); spectral efficiency, channel flatness, or signal-to-noise ratio (SNR); or a combination thereof.
39. The node of claim 24, wherein the received signal comprises a Physical Downlink Shared Channel (PDSCH) transmission.
40. The node of claim 24, wherein the reward parameter is offset by a penalty amount.
41. The node of claim 40, wherein the penalty amount depends on a number of one or more beams measured for use in the beam management procedure.
42. The node of claim 40, wherein the penalty amount depends on an amount of power consumption associated with the beam management procedure.
43. The node of claim 24, wherein the one or more beams are used for transmission, reception, or both of one or more synchronization signal blocks (SSBs).
44. The node of claim 24, wherein the means for performing the beam management procedure using the determined one or more beams comprises: means for measuring the channel based on a synchronization signal block SSB transmission from a base station BS using the determined one or more beams, said SSB transmission being associated with one or more transmit beams of said BS; as well as Means for selecting one or more beam pair links (BPLs) associated with one or more channel measurements above a channel measurement threshold; or one or more strongest channel measurements of all channel metrics associated with the SSB transmission; or a combination thereof.
45. The node of claim 44, wherein the determined one or more beams comprise a subset of available receive beams.
46. The node of claim 44, further comprising: means for receiving a physical downlink shared channel (PDSCH) using one of the one or more BPLs; means for determining a throughput associated with the PDSCH; means for updating the adaptive learning algorithm based on the determined throughput; as well as Means for using the updated adaptive learning algorithm to determine another one or more beams to be used to perform another beam management procedure to select another one or more BPLs.
47. A node configured for wireless communication, comprising: one or more memories; as well as one or more processors coupled to the one or more memories and configured to: using adaptive learning to determine one or more beams to be used for a beam management procedure; as well as The beam management procedure is performed using the determined one or more beams, wherein the adaptive learning uses a state parameter associated with a channel measurement, a reward parameter associated with a received signal throughput or spectral efficiency, and an action parameter associated with a selection of a beam pair corresponding to the channel measurement.
48. The node of claim 47, wherein the one or more processors are further configured to: updating an adaptive learning algorithm used for the adaptive learning based on feedback or training information or a combination thereof; and Another beam management procedure is performed using the updated adaptive learning algorithm.
49. The node of claim 48, wherein the feedback comprises feedback associated with the beam management procedure.
50. The node of claim 48, wherein the training information comprises: training information obtained by deploying one or more user equipment UEs in one or more simulated communication environments prior to network deployment of the one or more user equipment UEs; or training information obtained through feedback previously received when the one or more UEs were deployed in one or more communication environments; or training information from at least one of the network, one or more UEs, or a cloud; or training information received while the node is at least one of online or idle; or Its combination.
51. A node as described in claim 50, wherein the training information includes training information received from one or more UEs different from the node after the node is deployed, wherein the training information includes information associated with beam measurements performed by the one or more UEs, or feedback associated with one or more beam management procedures performed by the one or more UEs, or a combination thereof.
52. The node of claim 50, wherein the node comprises a UE.
53. The node of claim 48, wherein: Using the adaptive learning algorithm includes outputting an action based on one or more inputs; The feedback is associated with the action; and Updating the adaptive learning algorithm based on the feedback includes adjusting one or more weights applied to the one or more inputs.
54. The node of claim 48, wherein the adaptive learning algorithm comprises an adaptive machine learning algorithm; Adaptive reinforcement learning algorithm; Adaptive deep learning algorithms; An adaptive continuous infinite learning algorithm; or an adaptive policy optimization reinforcement learning algorithm, or a combination thereof.
55. The node of claim 48, wherein the adaptive learning algorithm is modeled as a Partially Observable Markov Decision Process (POMDP).
56. The node of claim 48, wherein the adaptive learning algorithm is implemented by an artificial neural network.
57. A node as claimed in claim 56, wherein: The artificial neural network comprises a deep Q network DQN comprising one or more deep neural networks DNN; and To determine the one or more beams using the adaptive learning, the one or more processors are configured to: passing one or more state parameters and one or more action parameters through the one or more DNNs; For each state parameter, output the value of each action parameter; and Select the action associated with the maximum output value.
58. The node of claim 56, wherein to update the adaptive learning algorithm, the one or more processors are configured to adjust one or more weights associated with one or more neuronal connections in the artificial neural network.
59. The node of claim 47, wherein to determine the one or more beams to be used for the beam management procedure using the adaptive learning, the one or more processors are configured to: determining one or more beams to be included in a codebook based on the adaptive learning; and One or more beams are selected from the codebook to be used for the beam management procedure.
60. The node of claim 47, wherein to determine the one or more beams to be used for the beam management procedure, the one or more processors are configured to use the adaptive learning to select one or more beams to be used for the beam management procedure from a codebook.
61. The node of claim 47, wherein the channel measurement comprises reference signal received power (RSRP); spectral efficiency, channel flatness, or signal-to-noise ratio (SNR); or a combination thereof.
62. The node of claim 47, wherein the received signal comprises a Physical Downlink Shared Channel (PDSCH) transmission.
63. The node of claim 47, wherein the reward parameter is offset by a penalty amount.
64. The node of claim 63, wherein the penalty amount depends on a number of one or more beams measured for use in the beam management procedure.
65. The node of claim 63, wherein the penalty amount depends on an amount of power consumption associated with the beam management procedure.
66. The node of claim 47, wherein the one or more beams are used for transmission, reception, or both of one or more synchronization signal blocks (SSBs).
67. The node of claim 47, wherein to perform the beam management procedure using the determined one or more beams, the one or more processors are configured to: using the determined one or more beams to measure the channel based on a synchronization signal block SSB transmission from a base station BS, the SSB transmission being associated with one or more transmit beams of the BS; and One or more beam-pair links (BPLs) are selected, the one or more BPLs being associated with one or more channel measurements above a channel measurement threshold; or one or more strongest channel measurements among all channel measurements associated with the SSB transmission; or a combination thereof.
68. The node of claim 67, wherein the determined one or more beams comprise a subset of available receive beams.
69. The node of claim 67, wherein the one or more processors are further configured to: receiving a physical downlink shared channel (PDSCH) using one of the one or more BPLs; determining a throughput associated with the PDSCH; updating the adaptive learning algorithm based on the determined throughput; and The updated adaptive learning algorithm is used to determine another one or more beams to be used to perform another beam management procedure to select another one or more BPLs.
70. A non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to: using adaptive learning to determine one or more beams to be used for a beam management procedure; and The beam management procedure is performed using the determined one or more beams, wherein the adaptive learning uses a state parameter associated with a channel measurement, a reward parameter associated with a received signal throughput or spectral efficiency, and an action parameter associated with a selection of a beam pair corresponding to the channel measurement.
Citation Information
Patent Citations
Methods, Network Node and Wireless Terminal for Beam Tracking when Beamforming is Employed
US20190052341A1
System and method for antenna beam selection
WO2019029802A1