Beam management using adaptive learning
Optimizing beam management through adaptive learning algorithms, the problems of low beam management efficiency and high power consumption in wireless communication systems are solved, and more efficient beam selection and communication quality improvement are achieved.
Patent Information
- Application Number
- CN202510546206.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-01
- Filing Date
- 2020-04-10
- Publication Date
- 2025-08-12
AI Technical Summary
Existing wireless communication systems have problems of inefficiency and high power consumption in beam management, especially in millimeter wave communication, and a more efficient beam management method is needed to overcome path loss and environmental changes.
Adaptive learning algorithms are used to determine and select beams, and adaptive reinforcement learning and deep learning technologies are used to update beam management algorithms through training and feedback, optimize the beam selection process, reduce unnecessary measurements and reduce power consumption.
Improves the efficiency of beam management and reduces power consumption, adapts to changes in different communication environments, provides personalized beam selection, and improves communication quality and throughput.
Smart Images

Figure CN120474589A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application with an application date of April 10, 2020, application number 202080031295.2, PCT international application number PCT / US2020 / 027648, and titled “Beam Management Using Adaptive Learning”.
[0002] Priority claim
[0003] This patent application claims priority to U.S. non-provisional application No. 16 / 400,864, filed on May 1, 2019, entitled “BEAM MANAGEMENT USING ADAPTIVE LEARNING,” which is assigned to the assignee of the present application and is hereby expressly incorporated herein by reference. Technical Field
[0004] Aspects of the present disclosure relate generally to wireless communications and, more particularly, to techniques for beam management. Background Art
[0005] Wireless communication systems are widely deployed to provide various telecommunication services such as telephony, video, data, messaging, broadcasts, etc. These wireless communication systems may employ multiple-access technologies capable of supporting communication with multiple users by sharing available system resources (e.g., bandwidth, transmit power, etc.). Examples of such multiple-access systems include 3rd Generation Partnership Project (3GPP) Long Term Evolution (LTE) systems, LTE-Advanced (LTE-A) systems, Code Division Multiple Access (CDMA) systems, Time Division Multiple Access (TDMA) systems, Frequency Division Multiple Access (FDMA) systems, Orthogonal Frequency Division Multiple Access (OFDMA) systems, Single Carrier Frequency Division Multiple Access (SC-FDMA) systems, and Time Division Synchronous Code Division Multiple Access (TD-SCDMA) systems, to name a few.
[0006] In some examples, a wireless multiple-access communication system may include several base stations (BSs), each of which is capable of simultaneously supporting communication for multiple communication devices (also referred to as user equipment (UE)). In an LTE or LTE-A network, a set of one or more base stations may define an evolved Node B (eNB). In other examples (e.g., in next-generation, new radio (NR), or 5G networks), a wireless multiple-access communication system may include several distributed units (DUs) (e.g., edge units (EUs), edge nodes (ENs), radio heads (RHs), smart radio heads (SRHs), transmission reception points (TRPs), etc.) in communication with several central units (CUs) (e.g., central nodes (CNs), access node controllers (ANCs), etc.), wherein a set of one or more DUs in communication with a CU may define an access node (e.g., which may be referred to as a BS, next-generation Node B (gNB or g Node B), TRP, etc.). A BS or DU may communicate with a set of UEs on downlink channels (e.g., for transmissions from the BS or DU to the UE) and uplink channels (e.g., for transmissions from the UE to the BS or DU).
[0007] These multiple access technologies have been adopted in various telecommunication standards to provide a common protocol that enables different wireless devices to communicate at a city, country, region, and even global level. New radio (e.g., 5G NR) is an example of an emerging telecommunication standard. NR is an enhancement to the LTE mobile standard promulgated by 3GPP. NR is designed to better support mobile broadband Internet access by improving spectrum efficiency, reducing costs, improving services, utilizing new spectrum, and better integrating with other open standards using OFDMA with cyclic prefix (CP) on downlink (DL) and uplink (UL). To this end, NR supports beamforming, multiple-input multiple-output (MIMO) antenna technology, and carrier aggregation.
[0008] However, as the demand for mobile broadband access continues to grow, there is a need for further improvements to NR and LTE technologies. Preferably, these improvements should also apply to other multiple access technologies and the telecommunication standards that employ them. Summary of the Invention
[0009] The systems, methods, and devices of the present disclosure each have several aspects, no single one of which is solely responsible for its desirable attributes. Without limiting the scope of the disclosure as expressed in the appended claims, some features will now be briefly discussed. After considering this discussion, and particularly after reading the section entitled "Detailed Description," one will understand how the features of the present disclosure provide advantages, including improved beam management procedures using adaptive learning.
[0010] Certain aspects provide a method for wireless communication by a node. The method generally includes using adaptive learning to determine one or more beams to use for a beam management procedure. The method generally includes performing a beam management procedure using the determined one or more beams.
[0011] In some examples, the node is a base station (BS).
[0012] In some examples, the node is a user equipment (UE).
[0013] In some examples, the method includes updating an adaptive learning algorithm used for adaptive learning. In some examples, the adaptive learning algorithm is updated based on feedback and / or training information. In some examples, the method includes performing another beam management procedure using the updated adaptive learning algorithm.
[0014] In some examples, the feedback includes feedback associated with a beam management procedure.
[0015] In some examples, the training information includes one or more of the following: training information obtained by deploying one or more UEs in one or more simulated communication environments prior to network deployment of the one or more UEs; training information obtained through feedback previously received when the one or more UEs were deployed in the one or more communication environments; training information from the network, one or more UEs and / or the cloud; and / or training information received when the node is online and / or idle.
[0016] In some examples, the training information includes training information received from one or more UEs other than the node after the node is deployed. In some examples, the training information includes information associated with a beam. In some examples, the training information includes measurements made by the one or more UEs or feedback associated with one or more beam management procedures performed by the one or more UEs.
[0017] In some examples, using the adaptive learning algorithm includes outputting an action based on one or more inputs. In some examples, feedback is associated with the action. In some examples, updating the adaptive learning algorithm based on the feedback includes adjusting one or more weights applied to the one or more inputs.
[0018] In some examples, the adaptive learning algorithm includes an adaptive machine learning algorithm; an adaptive reinforcement learning algorithm; an adaptive deep learning algorithm; an adaptive continuous infinite learning algorithm; and / or an adaptive policy optimization reinforcement learning algorithm.
[0019] In some examples, the adaptive learning algorithm is modeled as a partially observable Markov decision process (POMDP).
[0020] In some examples, the adaptive learning algorithm is implemented by an artificial neural network.
[0021] In some examples, the artificial neural network includes a deep Q-network (DQN) comprising one or more deep neural networks (DNNs). In some examples, determining the one or more beams using adaptive learning includes passing state parameters and action parameters through the one or more DNNs; for each state parameter, outputting a value for each action parameter; and selecting the action associated with the maximum output value.
[0022] In some examples, updating the adaptive learning algorithm includes adjusting one or more weights associated with one or more neuronal connections in the artificial neural network.
[0023] In some examples, using adaptive learning to determine one or more beams to use for the beam management procedure includes determining one or more beams to include in a codebook based on the adaptive learning and selecting the one or more beams to use for the beam management procedure from the codebook.
[0024] In some examples, determining one or more beams to be used for the beam management procedure includes using adaptive learning to select one or more beams to be used for the beam management procedure from a codebook.
[0025] In some examples, adaptive learning uses state parameters associated with channel measurements, reward parameters associated with received signal throughput or spectral efficiency, and action parameters associated with selection of beam pairs corresponding to the channel measurements.
[0026] In some examples, channel measurements include reference signal received power (RSRP); spectral efficiency, channel flatness, and / or signal-to-noise ratio (SNR).
[0027] In some examples, the received signal includes a physical downlink shared channel (PDSCH) transmission.
[0028] In some examples, the reward parameter is deducted by a penalty amount.
[0029] In some examples, the amount of penalty depends on the number of one or more beams being measured for the beam management procedure.
[0030] In some examples, the amount of penalty depends on the amount of power consumption associated with the beam management procedure.
[0031] In some examples, the beam includes one or more beams for transmission and / or reception of one or more synchronization signal blocks (SSBs).
[0032] In some examples, performing a beam management procedure using the determined one or more beams includes measuring a channel using the determined one or more beams based on an SSB transmission from a BS, the SSB transmission being associated with one or more transmit beams of the BS; and selecting one or more beam pair links (BPLs) associated with the one or more channel measurements that are above a channel measurement threshold and / or are one or more strongest channel measurements of all channel measurements associated with the SSB transmissions.
[0033] In some examples, the determined one or more beams comprise a subset of available receive beams.
[0034] In some examples, the method includes receiving a PDSCH using one of one or more selected BPLs; determining a throughput associated with the PDSCH; updating an adaptive learning algorithm based on the determined throughput; and using the updated adaptive learning algorithm to determine another one or more beams to be used to perform another beam management procedure for selecting another one or more BPLs.
[0035] Certain aspects provide a node configured for wireless communication. The node generally includes means for using adaptive learning to determine one or more beams to be used for a beam management procedure. The node generally includes means for performing a beam management procedure using the determined one or more beams.
[0036] Certain aspects provide a node configured for wireless communication. The node generally includes a memory. The node generally includes a processor coupled to the memory and configured to use adaptive learning to determine one or more beams to be used for a beam management procedure. The processor and memory are generally configured to perform the beam management procedure using the determined one or more beams.
[0037] Certain aspects provide a computer-readable medium. The computer-readable medium generally stores computer-executable code. The computer-executable code generally includes code for using adaptive learning to determine one or more beams to be used for a beam management procedure. The computer-executable code generally includes code for performing the beam management procedure using the determined one or more beams.
[0038] To accomplish the foregoing and related ends, one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and accompanying drawings set forth in detail certain illustrative features of the one or more aspects. However, these features are indicative of but a few of the various ways in which the principles of the various aspects may be employed. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order that the manner in which the above-recited features of the present disclosure may be understood in detail, a more particular description of the content briefly summarized above may be obtained by reference to various aspects, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only certain typical aspects of the disclosure and are therefore not to be considered limiting of its scope, for the description may admit to other equally effective aspects.
[0040] Figure 1 is a block diagram conceptually illustrating an example telecommunications system in accordance with certain aspects of the present disclosure.
[0041] Figure 2 Illustrated are example beam management procedures in accordance with certain aspects of the present disclosure.
[0042] Figure 3 Illustrated are example synchronization signal block (SSB) locations within an example half-frame in accordance with certain aspects of the present disclosure.
[0043] Figure 4 Illustrated are example transmit and receive beams for SSB measurements in accordance with certain aspects of the present disclosure.
[0044] Figure 5 An example networking environment is illustrated in which predictive models are used for beam management in accordance with certain aspects of the present disclosure.
[0045] Figure 6 An example reinforcement learning model according to certain aspects of the present disclosure is conceptually illustrated.
[0046] Figure 7 An example Deep Q-Network (DQN) learning model, in accordance with certain aspects of the present disclosure, is conceptually illustrated.
[0047] Figure 8 is a flow diagram illustrating example operations for wireless communications by a node in accordance with certain aspects of the present disclosure.
[0048] Figure 9 is an example call flow diagram illustrating example signaling for beam management using adaptive learning, in accordance with certain aspects of the present disclosure.
[0049] Figure 10 is an example call flow diagram illustrating example signaling for a BPL discovery procedure using adaptive learning, in accordance with certain aspects of the present disclosure.
[0050] Figure 11 Illustrated are communications devices that may include various components configured to perform operations for the techniques disclosed herein, in accordance with aspects of the present disclosure.
[0051] Figure 12is a block diagram conceptually illustrating designs of example base stations (BSs) and user equipment (UEs) in accordance with certain aspects of the present disclosure.
[0052] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the figures. It is contemplated that elements disclosed in one aspect may be beneficially utilized in other aspects without specific recitation. DETAILED DESCRIPTION
[0053] Aspects of the present disclosure provide apparatus (equipment), methods, processing systems, and computer-readable media for beam management using adaptive learning.
[0054] Certain systems such as new radio systems (e.g., 5G NR) support millimeter wave (mmW) communications. In mmW communications, signals used for communication between devices (referred to as mmW signals) may have a high carrier frequency (e.g., 25 GHz or higher, such as within the 30 to 300 GHz band) and may have a wavelength in the range of 1 mm to 10 mm. Based on such characteristics of mmW signals, mmW communications can provide high-speed (e.g., gigabit speed) communications between devices. However, compared to lower frequency signals, mmW signals may experience atmospheric effects and may not propagate well through materials. Therefore, compared to lower frequency signals, mmW signals may experience relatively high path loss (e.g., attenuation or reduction in the power density of the wave corresponding to the mmW signal) as they propagate.
[0055] To overcome path loss, mmW communication systems utilize directional beamforming. Beamforming may involve the use of transmit (TX) beams and / or receive (RX) beams. A TX beam corresponds to a transmitted mmW signal that is directed to have more power in a specific direction relative to other directions, such as toward a receiver. By directing the transmitted mmW signal to the receiver, more energy of the mmW signal is directed to the receiver, thereby overcoming higher path loss. RX beamforming corresponds to a technique performed at the receiver to apply gain to a signal received in a specific direction while attenuating signals received in other directions. Using RX beams also helps to overcome higher path loss, for example by improving the signal-to-noise ratio (SNR) of receiving the desired mmW signal at the receiver. In some aspects, hybrid beamforming (e.g., signal processing in analog and digital domains) may be used.
[0056] Therefore, in some aspects, for a particular transmitter to communicate with a particular receiver, the transmitter needs to select a TX beam to use, and the receiver needs to select an RX beam to use. The TX beam and RX beam used for communication are referred to as a beam pair. In some aspects, the RX and TX beams in a beam pair are selected to provide sufficient communication coverage and / or capacity.
[0057] In certain aspects, a beam management procedure may be used to select (e.g., initially select, update selection, refine to a narrower beam within a previously selected beam, etc.) a beam pairing. As will be described below with reference to Figure 2-4 As discussed in more detail, the beam management procedure may involve measuring signals using different RX and / or TX beams for reception / transmission and selecting a beam for beam pairing based on the measurements. For example, the beam with the highest measured channel or link quality (e.g., throughput, SNR, etc.) among the measured beams may be selected.
[0058] In some cases, as described below with reference to Figure 2-4 As discussed in more detail, there are a large number of RX and / or TX beams supported at the transmitter and / or receiver, which may mean that there are a large number of measurements that may be performed for the beam management procedure. In addition, the communication environment between the transmitter and the receiver may be different at different times, such as due to obstructions (e.g., when a user's hand blocks the TX / RX beam at the transmitter / receiver (e.g., user equipment (UE)) and / or an object blocks the line of sight (LOS) path between the transmitter and the receiver), movement and / or rotation of the transmitter / receiver, etc.
[0059] To account for such factors, in some cases, beam management procedures are based on heuristics. Heuristic-based beam management procedures attempt to anticipate real-world deployment scenarios for transmitters and receivers and typically update the beam management procedures used by the transmitter and receiver (such as using downloaded software patches) based on issues encountered (or anticipated) over time as the transmitter and receiver communicate. For example, a heuristic-based beam management procedure may measure only certain RX and / or TX beams of the transmitter and receiver, rather than all beams, based on transmitter and / or receiver parameters.
[0060] To further improve the beam management procedure, aspects of the present disclosure provide for the use of adaptive learning as part of the beam management procedure. For example, a UE (and / or BS) acting as a transmitter and / or receiver may use an adaptive learning-based beam management algorithm that adapts over time based on learning. Specifically, learning may be based on feedback associated with previous beam selections for the UE and / or BS. The feedback may include an indication of the previous beam selection and parameters associated with the previous beam selection. The algorithm may initially be trained based on feedback in a laboratory environment and then updated (e.g., continuously) using feedback while the UE and / or BS is deployed. In some examples, the algorithm is a deep reinforcement learning-based beam management algorithm that uses machine learning and artificial neural networks to update and apply predictive models for beam selection during the beam management procedure. In this way, the adaptive learning-based beam management algorithm learns from user behavior (e.g., frequently traversed paths, how the user holds the UE, etc.) and is therefore also personalized to the user.
[0061] The following description provides examples of using adaptive learning as part of a beam management procedure and is not intended to limit the scope, applicability, or examples set forth in the claims. Changes may be made to the functions and arrangements of the elements discussed without departing from the scope of this disclosure. Various examples may omit, substitute, or add various procedures or components as appropriate. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Moreover, features described with reference to some examples may be combined in some other examples. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects set forth herein. In addition, the scope of this disclosure is intended to cover such apparatus or methods practiced using other structures, functionalities, or both, in addition to or in addition to the various aspects of this disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be implemented by one or more elements of the claims. The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as superior to or preferable to other aspects.
[0062] Figure 1An example wireless communication network 100 is illustrated in which various aspects of the present disclosure may be performed. For example, the wireless communication network 100 may be a new radio system (e.g., a 5G NR network). The wireless communication network 100 may support mmW communication through beamforming. Nodes (e.g., wireless nodes) such as a UE 120a and / or a base station (BS) 110a in the wireless communication network 100 may be configured to perform a beam management procedure to select a beam pairing for communicating with another node. For example, the UE 120a and the BS 110a may perform a beam management procedure to determine a receive beam of the UE 120a and a transmit beam of the BS 110a as a beam pairing to be used for communication (e.g., downlink communication), also referred to as a beam pair link (BPL). As will be described in greater detail herein, the UE 120a and / or the BS 110a may use a beam management procedure based on adaptive learning. The UE 120a and / or the BS 110a may use adaptive learning to determine one or more beams to be used for the beam management procedure. As will be described in greater detail herein, the UE 120a and / or the BS 110a may use adaptive learning to determine one or more beams to be used for the beam management procedure. Figure 1 As shown in FIG, UE 120a has a beam selection manager 122. According to one or more aspects described herein, the beam selection manager 122 can be configured to use an adaptive learning-based algorithm to determine / select beams to be used for beam management procedures. Figure 1 As shown, BS 110a may additionally or alternatively include a beam selection manager 112. According to various aspects described herein, beam selection manager 112 may be configured to use an adaptive learning algorithm to determine / select beams to be used for beam management procedures. UE 120a and / or BS 110a may then use the determined beam or beams to perform beam management procedures.
[0063] It should be noted that although certain aspects are described with respect to the beam management procedure being performed by wireless nodes, certain aspects of this beam management procedure may be performed by other types of nodes, such as nodes connected to a BS through a wired connection.
[0064] like Figure 1 As illustrated in , the wireless communication network 100 may include several BSs 110a-z (each also individually referred to herein as BS 110 or collectively referred to as BS 110) and other network entities. BS 110 may communicate with UEs 120a-y (each also individually referred to herein as UE 120 or collectively referred to as UE 120) in the wireless communication network 100. Each BS 110 may provide communication coverage for a particular geographic area. In some examples, the BSs 110 may be interconnected to each other and / or to one or more other BSs or network nodes (not shown) in the wireless communication network 100 via various types of backhaul interfaces, such as direct physical connections, wireless connections, virtual networks, or the like using any suitable transport network. Figure 1In the example shown in FIG, BSs 110a, 110b, and 110c may be macro BSs for macro cells 102a, 102b, and 102c, respectively. BS 110x may be a pico BS for pico cell 102x. BSs 110y and 110z may be femto BSs for femto cells 102y and 102z, respectively. A BS may support one or more (e.g., three) cells.
[0065] The wireless communication network 100 may also include a relay station. A relay station is a station that receives transmissions of data and / or other information from an upstream station (e.g., a BS or a UE) and sends transmissions of the data and / or other information to a downstream station (e.g., a UE or a BS). A relay station may also be a UE that relays transmissions for other UEs. Figure 1 In the example shown in , a relay station 110r may communicate with a BS 110a and a UE 120r to facilitate communication between the BS 110a and the UE 120r. A relay station may also be referred to as a relay BS, a relay, or the like.
[0066] UEs 120 (eg, 120x, 120y, etc.) may be dispersed throughout wireless communication network 100, and each UE may be stationary or mobile.
[0067] Network controller 130 may be coupled to a set of BSs and provide coordination and control for these BSs. Network controller 130 may communicate with BSs 110 via a backhaul. BSs 110 may also communicate with each other (eg, directly or indirectly) via a wireless or wired backhaul.
[0068] In some examples, the wireless communication network 100 (e.g., a 5G NR network) may support mmW communications. As discussed above, such systems using mmW communications may use beamforming to overcome high path loss and may perform beam management procedures to select beams for beamforming.
[0069] A BS beam (e.g., TX or RX) and a UE beam (e.g., the other of TX or RX) form a BPL. Both the BS (e.g., BS 110a) and the UE (e.g., UE 120a) may determine (e.g., find / select) at least one eligible beam to form a communication link. For example, on the downlink, BS 110a uses a transmit beam to transmit a downlink transmission, while UE 120a uses a receive beam to receive the downlink transmission. The combination of the transmit beam and receive beam forms a BPL. UE 120a and BS 110a establish at least one BPL for UE 120a to the wireless communication network 100. In some examples, multiple BPLs (e.g., a group of BPLs) may be configured for communication between UE 120a and one or more BSs 110. Different BPLs may be used for different purposes, such as for communicating different channels, for communicating with different BSs, and / or as a fallback BPL in the event that an existing BPL fails.
[0070] In some examples, for initial cell acquisition, a UE (e.g., UE 120a) may search for the strongest signal corresponding to a cell associated with a BS (e.g., BS 110a) and the associated UE receive beam and BS transmit beam corresponding to the BPL for receiving / transmitting reference signals. After initial acquisition, UE 120a may perform new cell detection and measurements. For example, UE 120a may measure a primary synchronization signal (PSS) and a secondary synchronization signal (SSS) to detect a new cell. As described below with reference to Figure 3 As discussed in more detail, the PSS / SSS may be transmitted by a BS (e.g., BS 110a) in different synchronization signal blocks (SSBs) across one or more synchronization signal (SS) burst sets. UE 120a may measure the different SSBs within the SS burst set to perform beam management procedures, as further discussed herein.
[0071] In 5G NR, the beam management procedure for determining BPL may be referred to as the P1 procedure. Figure 2An example P1 procedure 202 is illustrated. BS 210 (e.g., such as BS 110a) may send a measurement request to UE 220 (e.g., such as UE 120a) and may subsequently transmit one or more signals (sometimes referred to as "P1 signals") to UE 220 for measurement. In P1 procedure 202, BS 210 transmits signals in different spatial directions (corresponding to transmit beams 211, 212, ... 217) in each symbol using beamforming so as to reach several (e.g., most or all) relevant spatial locations of the cell of BS 210. In this manner, BS 210 transmits signals using different transmit beams in different directions over time. In some examples, SSBs are used as P1 signals. In some examples, channel state information reference signals (CSI-RS), demodulation reference signals (DMRS), or another downlink signal may be used as P1 signals.
[0072] In the P1 procedure 202, in order to successfully receive at least one symbol of the P1 signal, the UE 220 finds (e.g., determines / selects) an appropriate receive beam (221, 222, ... 226). Signals (e.g., SSBs) from multiple BSs can be measured simultaneously for a given signal index (e.g., SSB index) corresponding to a given time period. The UE 220 can apply a different receive beam during each occurrence (e.g., each symbol) of the P1 signal. Once the UE 220 successfully receives a symbol of the P1 signal, the UE 220 and the BS 210 have discovered the BPL (i.e., the UE RX beam used to receive the P1 signal in that symbol and the BS TX beam used to transmit the P1 signal in that symbol). In some cases, the UE 220 does not search all possible UE RX beams until it finds the best UE RX beam, as this incurs additional delay. UE 220 may instead select an RX beam once it is "good enough" (e.g., has a quality (e.g., SNR) that satisfies a threshold (e.g., a predefined threshold). UE 220 may not know which beam BS 210 used to transmit the P1 signal in a symbol; however, UE 220 may report to BS 210 the time at which it observed the signal. For example, UE 220 may report to BS 210 the symbol index at which the P1 signal was successfully received. BS 210 may receive the report and may determine which BS TX beam BS 210 used at the indicated time. In some examples, UE 220 measures the signal quality of the P1 signal, such as reference signal received power (RSRP) or another signal quality parameter (e.g., SNR, channel flatness, etc.). UE 220 may report the measured signal quality (e.g., RSRP) to BS 210 along with the symbol index. In some cases, UE 220 may report multiple symbol indices corresponding to multiple BS TX beams to BS 210.
[0073] As part of the beam management procedure, the BPL used between UE 220 and BS 110 may be refined / changed. For example, the BPL may be periodically refined to adapt to changing channel conditions, such as attenuation due to movement of UE 220 or other objects, Doppler spread, and the like. UE 220 may monitor the quality of the BPL (e.g., the BPL found / selected during the P1 procedure and / or the previously refined BPL) to refine the BPL when quality degrades (e.g., when the BPL quality falls below a threshold or when another BPL has higher quality). In 5G NR, the beam management procedure for BPL beam refinement may be referred to as the P2 and P3 procedures for respectively refining the BS beam and UE beam of an individual BPL.
[0074] Figure 2 An example P2 procedure 204 and P3 procedure 206 are illustrated. Figure 2 As shown, for P2 protocol 204, BS 210 transmits symbols of a signal using different BS beams (e.g., TX beams 215, 214, 213) that are spatially close to the BS beam of the current BPL. For example, BS 210 transmits signals in different symbols using adjacent TX beams around the TX beam of the current BPL (e.g., beam sweeping). Figure 2 As shown, the TX beam used by BS 210 for P2 procedure 204 may be different from the TX beam used by BS 210 for P1 procedure 202. For example, the TX beam used by BS 210 for P2 procedure 204 may be spaced closer together and / or may be more focused (e.g., narrower) than the TX beam used by BS 210 for P1 procedure 202. During P2 procedure 204, UE 220 maintains its RX beam (e.g., RX beam 224) unchanged. UE 220 may measure the signal quality (e.g., RSRP) of the signal in different symbols and indicate the symbol in which the highest signal quality was measured. Based on this indication, BS 210 may determine the strongest (e.g., best or associated with the highest signal quality) TX beam (i.e., the TX beam used in the indicated symbol). The BPL may be refined accordingly to use the indicated TX beam.
[0075] like Figure 2As shown, for P3 procedure 206, BS 220 maintains a constant TX beam (e.g., the TX beam of the current BPL) and uses this constant TX beam (e.g., TX beam 214) to transmit the codewords of the signal. During P3 procedure 206, UE 220 uses different RX beams (e.g., RX beams 223, 224, 225) in different codewords to scan the signal. For example, UE 220 may use RX beams adjacent to the RX beam in the current BPL (i.e., the BPL being refined) to perform the sweep. UE 220 may measure the signal quality (e.g., RSRP) of the signal for each RX beam and identify the strongest UE RX beam. UE 220 may use the identified RX beam for the BPL. UE 220 may report the signal quality to BS 210.
[0076] As discussed above, in some examples, SSB measurements may be used for beam management. Figure 3 302. Example SSB positions within an example NR radio frame format 302 are illustrated. The transmission timeline for each of the downlink and uplink may be divided into units of radio frames. Figure 3 As shown, the example 10ms NR radio frame format 302 may include ten 1ms subframes (subframes with indices 0, 1...9). In NR, the basic transmission time interval (TTI) may be referred to as a time slot. In NR, a subframe may contain a variable number of time slots (e.g., 1, 2, 4, 8, 16... time slots), depending on the subcarrier spacing (SCS). NR may support a base SCS of 15KHz, and other SCSs may be defined relative to the base SCS (e.g., 30kHz, 60kHz, 120kHz, 240kHz, etc.). Figure 3 In the example shown, the SCS is 120kHz. Figure 3 As shown, subframe 304 (subframe 0) contains 8 time slots (slots 0, 1, ... 7) with a duration of 0.125 ms. The symbol and slot lengths scale with the subcarrier spacing. Each time slot may include a variable number of symbol (e.g., OFDM symbol) periods (e.g., 7 or 14 symbols), depending on the SCS. Figure 3 For the 120 kHz SCS shown, each of slot 306 (Slot 0) and slot 308 (Slot 1) includes 14 symbol periods (slots with indices 0, 1 ... 13) having a duration of 0.25 ms.
[0077] In some examples, the SSB may be transmitted up to sixty-four times in up to sixty-four different beam directions. The up to sixty-four transmissions of the SSB are referred to as an SS burst set. The SSBs in an SS burst set may be transmitted in the same frequency region, while SSBs in different SS burst sets may be transmitted in different frequency regions. Figure 3In the example shown, in subframe 304, SSB is transmitted in every time slot (time slot 0, 1...6). Figure 3 In the example shown, in slot 306 (slot 0), SSB 310 is transmitted in symbols 4, 5, 6, and 7, SSB 312 is transmitted in symbols 8, 9, 10, and 11, and in slot 308 (slot 1), SSB 314 is transmitted in symbols 2, 3, 4, and 5, and SSB 316 is transmitted in symbols 6, 7, 8, and 9, and so on. The SSB may include the PSS, SSS, and a two-symbol physical broadcast channel (PBCH). The PSS and SSS may be used by the UE for cell search and acquisition. For example, the PSS may provide half-frame timing, the SSS may provide control protocol (CP) length and frame timing, and the PSS and SSS may provide cell identity. The PBCH carries some basic system information, such as downlink system bandwidth, timing information within a radio frame, SS burst set periodicity, system frame number, etc.
[0078] like Figure 4 As shown, SSB can be used to make measurements using different transmit and receive beams, for example according to beam management procedures such as Figure 2 P1 procedure 202 is shown. Figure 4 An example is illustrated of a BS 410 (e.g., such as BS 110a) using 4 TX beams and a UE 420 (e.g., such as UE 120a) using 2 RX beams. For each SSB, BS 410 transmits the SSB using a different TX beam. Figure 4 As shown, UE 420 may scan its RX beam 422 while BS 410 transmits SSBs 310, 312, 314, 316 sweeping its four TX beams 412, 414, 416, 418, respectively. BPLs may be identified and used for data communications within a time period, as discussed. Figure 4 As shown, BS 410 communicates data for a period of time using TX beam 414 and UE 420 uses RX beam 422. UE 410 may then scan its RX beam 424 while BS 410 transmits SSBs 426, 428 that sweep its TX beams 412, 414, and so on.
[0079] As can be seen, as the number of TX / RX beams increases, the number of times a UE scans each of its RX beams on each TX beam can increase. Power consumption can scale linearly with the number of measured SSBs. Therefore, the time and power overhead associated with beam management can increase when scanning all beams.
[0080] Thus, aspects of the present disclosure provide techniques for helping nodes perform measurements on other nodes when using beamforming, such as by using adaptive learning, which can reduce the number of measurements used for beam management procedures and thereby reduce power consumption.
[0081] Example beam management procedure using adaptive learning
[0082] Non-adaptive algorithms are deterministic and depend on their inputs. If the algorithm is presented with the exact same input at different times, its output will be exactly the same. Adaptive algorithms are algorithms that change their behavior based on past experience. This means that different devices using adaptive algorithms can end up with different algorithms over time.
[0083] According to certain aspects, a beam management procedure may be performed using a beam management algorithm based on adaptive learning. Thus, the beam algorithm changes (e.g., adapts, updates) over time based on new learning. The beam management procedure may be used for initial acquisition, cell discovery after initial acquisition, and / or determining the BPL for the strongest cell detected by the UE. For example, adaptive learning may be used to construct a UE codebook that indicates beams to be used (e.g., measured) for beam management procedures. In some examples, adaptive learning may be used to select UE receive beams to be used to discover the BPL. Adaptive learning may be used to intelligently select which UE receive beams to use to measure signals based on training and experience, so that fewer beams can be measured while still finding a suitable BPL (e.g., meeting a threshold signal quality).
[0084] In some examples, adaptive learning-based beam management involves training a model, such as a prediction model. The model can be used during a beam management procedure to select which UE receive beams to use to measure signals. The model can be trained based on training data (e.g., training information), which can include feedback, such as feedback associated with the beam management procedure. Figure 5 Illustrated is an example networking environment 500 in which a prediction model 524 is used for beam management, in accordance with certain aspects of the present disclosure.
[0085] like Figure 5 As shown, networked environment 500 includes a node 520, a training system 530, and a training repository 515 that are communicatively connected via a network 505. Node 520 may be a UE (e.g., such as UE 120a in wireless communication network 100) or a BS (e.g., such as BS 110a in wireless communication network 100). Network 505 may be a wireless network, such as wireless communication network 100, which may be a 5G NR network. Although training system 530, node 520, and training repository 515 are communicatively connected, Figure 5Although illustrated as separate components in FIG, those skilled in the art will recognize that training system 530, nodes 520, and training repository 515 may be implemented on any number of computing systems, as one or more stand-alone systems, or in a distributed environment.
[0086] The training system 530 generally includes a prediction model training manager 532 that uses training data to generate a prediction model 524 for beam management. The prediction model 524 can be determined based on information in the training repository 515.
[0087] The training repository 515 may include training data acquired before and / or after the node 520 is deployed. The node 520 may be trained in a simulated communication environment (e.g., in a field test, a drive test) before the node 520 is deployed. For example, various beam management procedures (e.g., various selections of UE RX beams for measuring signals) may be tested in various scenarios, such as at different UE speeds, with the UE stationary, with various rotations of the UE, with various BS deployments / geometries, etc., to obtain training information related to the beam management procedures. This information may be stored in the training repository 515. After deployment, the training repository 515 may be updated to include feedback associated with the beam management procedures performed by the node 520. The training repository may also be updated with information from other BSs and / or other UEs, for example, based on experiences learned by these BSs and / or other UEs that may be associated with the beam management procedures performed by these BSs and / or UEs.
[0088] The prediction model training manager 532 can use the information in the training repository 515 to determine a prediction model 524 (e.g., an algorithm) for beam management, such as to select a UE RX beam for measuring a signal. As discussed in more detail herein, the prediction model training manager 532 can use various types of adaptive learning, such as machine learning, deep learning, reinforcement learning, etc., to form the prediction model 524. The training system 530 can adapt (e.g., update / improve) the prediction model 524 over time. For example, when the training repository is updated with new training information (e.g., feedback), the model 524 is updated based on the new learning / experience.
[0089] Training system 530 may be located at node 520, a BS in network 505, or a different entity that determines prediction model 524. If located at a different entity, prediction model 524 is provided to node 520.
[0090] Training repository 515 may be a storage device, such as a memory. Training repository 515 may be located on node 520, training system 530, or another entity in network 505. Training repository 515 may be in cloud storage. Training repository 515 may receive training information from node 520, an entity in network 505 (e.g., a BS or UE in network 505), the cloud, or other sources.
[0091] As described above, the node 520 is provided with (or generates, e.g., if the training system 530 is implemented in the node 520) a prediction model. As illustrated, the node 520 may include a beam selection manager 522 configured to use the prediction model 524 for beam management (e.g., such as described above with reference to FIG. Figure 2 In some examples, node 520 utilizes prediction model 524 to construct a UE codebook and / or determine / select a beam from the UE codebook to be used for the beam management procedure. Prediction model 524 is updated as training system 530 adapts prediction model 524 with new learning.
[0092] Thus, the beam management algorithm of node 520 (using the prediction model 524) is based on adaptive learning, in that the algorithm used by the node 520 changes over time, even after deployment, based on the experience / feedback gained by the node 520 in the deployment scenario (and / or through training information also provided by other entities).
[0093] According to certain aspects, adaptive learning can use any appropriate learning algorithm. As described above, the learning algorithm can be used by a training system (e.g., such as training system 530) to train a prediction model (e.g., such as prediction model 524) to obtain an adaptive learning-based beam management algorithm for use by a device (e.g., such as node 520) for a beam management procedure. In some examples, the adaptive learning algorithm is an adaptive machine learning algorithm, an adaptive reinforcement learning algorithm, an adaptive deep learning algorithm, an adaptive continuous infinite learning algorithm, or an adaptive policy optimization reinforcement learning algorithm (e.g., a proximal policy optimization (PPO) algorithm, a policy gradient, a trust region policy optimization (TRPO) algorithm, etc.). In some examples, the adaptive learning algorithm is modeled as a partially observable Markov decision process (POMDP). In some examples, the adaptive learning algorithm is implemented by an artificial neural network (e.g., a deep Q network (DQN) including one or more deep neural networks (DNNs)).
[0094] In some examples, adaptive learning (e.g., used by training system 530) is performed using neural networks. Neural networks can be designed to have various connectivity patterns. In a feedforward network, information is passed from lower layers to higher layers, with each neuron in a given layer communicating to neurons in a higher layer. Hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have reflow or feedback (also known as top-down) connections. In a reflow connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. The reflow architecture can help identify patterns that span more than one chunk of input data delivered sequentially to the neural network. The connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. A network with many feedback connections can be helpful when the recognition of high-level concepts can assist in discerning specific low-level features of the input.
[0095] In some examples, adaptive learning (e.g., used by training system 530) is performed using a deep belief network (DBN). A DBN is a probabilistic model comprising multiple layers of hidden nodes. A DBN can be used to extract a hierarchical representation of a training data set. A DBN can be obtained by stacking multiple layers of restricted Boltzmann machines (RBMs). RBMs are a class of artificial neural networks that can learn probability distributions on an input set. Since RBMs can learn probability distributions without information about which class each input can be classified into, RBMs are often used for unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBM of a DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of the input and target category from the previous layer) and can be used as a classifier.
[0096] In some examples, adaptive learning (e.g., used by training system 530) is performed using a deep convolutional network (DCN). A DCN is a network in a convolutional network that is configured with additional pooling and normalization layers. DCN has achieved state-of-the-art performance on many tasks. DCN can be trained using supervised learning, where both input and output targets are known for many paradigms and are used to modify the weights of the network using gradient descent. DCN can be a feedforward network. In addition, as described above, the connections from the neurons in the first layer of the DCN to the neuron groups in the next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of the DCN can be utilized for fast processing. The computational burden of the DCN can be much smaller than that of a neural network of similar size that includes recurrent or feedback connections, for example.
[0097] An artificial neural network, which may include a group of interconnected artificial neurons (e.g., a neuron model), is a computing device or represents a method performed by a computing device. These neural networks can be used in various applications and / or devices, such as Internet Protocol (IP) cameras, Internet of Things (IoT) devices, autonomous vehicles, and / or service robots. Individual nodes in an artificial neural network can mimic biological neurons by taking input data and performing simple operations on the data. The results of these simple operations on the input data are selectively passed to other neurons. Weight values are associated with each vector and node in the network, and these values constrain how the input data relates to the output data. For example, the input data for each node can be multiplied by a corresponding weight value, and the products can be summed. The sum of these products can be adjusted by an optional bias, and an activation function can be applied to the result, thereby generating the node's output signal or "output activation." The weight values can be initially determined by iteratively flowing training data through the network (e.g., the weight values are established during a training phase, in which the network learns how to identify specific categories based on the input data characteristics that are typical of each category).
[0098] Adaptive learning (e.g., used by training system 530) can be implemented using different types of artificial neural networks, such as recurrent neural networks (RNNs), multilayer perceptron (MLP) neural networks, and convolutional neural networks (CNNs). RNNs work by storing the output of a layer and feeding that output back into the input to help predict the outcome of that layer. In an MLP neural network, data can be fed into an input layer, and one or more hidden layers provide several levels of abstraction of the data. Predictions can then be made for the output layer based on the abstracted data. MLPs are particularly well-suited for classification prediction problems, where inputs are assigned classes or labels. A convolutional neural network (CNN) is a type of feedforward artificial neural network. A convolutional neural network can include a collection of artificial neurons, each of which has a receptive field (e.g., a spatially localized region of the input space) and collectively forms an input space. Convolutional neural networks have numerous applications. In particular, CNNs have been widely used in the fields of pattern recognition and classification. In a layered neural network architecture, the outputs of artificial neurons in the first layer become the inputs to artificial neurons in the second layer, which in turn become the inputs to artificial neurons in the third layer, and so on. Convolutional neural networks can be trained to recognize hierarchies of features. Computation in convolutional neural network architectures can be distributed across a population of processing nodes, which can be configured in one or more computational chains. These multi-layer architectures can be trained one layer at a time and fine-tuned using backpropagation.
[0099] In some examples, when using an adaptive machine learning algorithm, training system 530 generates vectors from information in training repository 515. In some examples, training repository 515 stores vectors. In some examples, the vectors map one or more features to labels. For example, the features may correspond to various deployment scenario patterns discussed herein, such as UE mobility, speed, rotation, channel conditions, BS deployment / geometry in the network, and so on. The labels may correspond to predicted optimal beam selections (e.g., for RX beams) associated with the features used to perform beam management procedures. Prediction model training manager 532 may use the vectors to train prediction model 524 for node 520. As discussed above, the vectors may be associated with weights in the adaptive learning algorithm. When the learning algorithm adapts (e.g., is updated), the weights applied to the vectors may also be changed. Thus, when the beam management procedure is executed again under the same features (e.g., under the same set of conditions), the model may provide node 520 with different results (e.g., different beam selections).
[0100] Figure 6 An example reinforcement learning model is conceptually illustrated. Reinforcement learning can be a semi-supervised learning model in machine learning. Reinforcement learning allows an agent 604 (e.g., a node 520 and / or a training system 530) to take actions (e.g., beam selection) based on the state observed by an interpreter 602 (e.g., such as a node 520) (e.g., RSPR of SSBs using different beams) and interact with an environment 606 (e.g., a current deployment scenario) to maximize the total reward (e.g., physical downlink shared channel (PDSCH) throughput using the selected beam) that can be observed by the interpreter 602 and fed back to the agent 604 as reinforcement. In some examples, the agent 604 and the interpreter 602 can be implemented as the same or separate component devices that can perform various functions of the node 520, the training system 530, and / or the training repository 515.
[0101] In some examples, reinforcement learning is modeled as a Markov decision process (MDP). An MDP is a discrete, time-stochastic control process. An MDP provides a mathematical framework for modeling decision making in situations where the outcome may be partially random and partially under the control of the decision maker. In an MDP, at each time step, the process is in a state from a finite set of states S, and the decision maker can choose any action from a finite set of actions A available in that state. The process responds at the next time step by randomly moving to a new state and giving the decision maker a corresponding reward, where R α (s, s′) is the immediate reward (or expected immediate reward) after transitioning from state s to state s′. The probability that the process moves to its new state is affected by the action chosen, for example, according to the state transition function. The state transition can be represented by P α(s,s′)=Pr(s t11 =s′|s t =s,α t =a) is given.
[0102] An MDP seeks to find a policy for decision making: a function of π that specifies the action π(s) a decision maker will choose when in state s. The goal is to choose a policy π that maximizes the reward. For example, a policy that maximizes a cumulative function of the reward (such as the discounted sum). An example function is shown below:
[0103] in
[0104] α t =π(s t ), that is, the action given by the strategy, and γ is the discount factor and satisfies 0≤γ≤1.
[0105] The solution to an MDP is a policy that describes the optimal action (e.g., maximizing the expected discounted reward) for each state in the MDP.
[0106] In some examples, a partially observable MDP (POMDP) is used. A POMDP can be used when the state may not be known when an action is taken, and therefore the probability and / or reward may be unknown. For a POMDP, reinforcement learning can be used. The following function can be defined:
[0107] Q(s,a)=∑ si P α (s,s l )(R α (s,s l )+γV(s l ))
[0108] The experience gained during learning can be based on (s, a) pairs and the result s′. For example, if a node was previously in state s and made beam selection a, achieving a throughput s′, then the node can directly update the array Q based on this learned experience. This is referred to as Q-learning. In some examples, the learning algorithm can be continuous.
[0109] In some examples, for an adaptive learning-based beam management algorithm, the state may correspond to the M strongest beam quality measurements (e.g., reference signal received power (RSRP) of SSBs on different beams) in the environment (e.g., the current deployment scenario of the UE), including conditions discussed herein including UE mobility, BS deployment pattern (e.g., geometry), obstructions, etc. The action may correspond to beam selection. The reward may be the throughput achieved using beam selection, such as PDSCH throughput. The reward may be another parameter, such as, for example, spectral efficiency. Thus, by using this MDP at a given time in a given state, a node may adopt a strategy to find the beam selection that specifies the maximum throughput. As discussed above, the reward may be discounted. For beam management, the reward may be offset by a penalty as a function of the measured SSBs, for example, to optimize for minimum power.
[0110] Return to Reference Figure 5 The example network environment 500 and Figure 6 In some examples, the prediction model training manager 532 or the agent 604 can use reinforcement learning to determine a policy (e.g., an MDP solution) for the prediction model (e.g., prediction model 524). The node 520 or the agent 604 can take actions, such as beam selection for a beam management procedure, based on the policy given by the prediction model (e.g., prediction model 524) at a given time in the environment (e.g., environment 606) for the current state (e.g., observed by the node 520 or the interpreter 602). The reinforcement learning algorithm and the prediction model can be updated / adapted based on the learned experience (e.g., which can be stored in the training repository 515).
[0111] The framework of reinforcement learning provides tools for optimally solving POMDPs. This learning modifies the weights of a multilayer perceptron (e.g., a neural network) that determines the next action to take. Algorithms in deep ML are encoded in the weights of the neural network. Thus, changing the weights changes the algorithm.
[0112] In some examples, adaptive learning-based beam management uses an adaptive deep learning algorithm, which can be a deep Q network (DQN) implemented by a neural network. Figure 7 An example DQN learning model 700 is conceptually illustrated according to certain aspects of the present disclosure. Figure 7 As shown, agent 706 (e.g., such as agent 604 or node 520) includes an artificial neural network, such as in Figure 77. A deep neural network (DNN) 708 is shown in the example of FIG. For a current environment 702 (e.g., such as environment 606), which can be a real-world deployment scenario involving a UE (e.g., UE 120a) and a BS (e.g., BS 110a) as described herein and various conditions, an agent 706 observes a state 704(s). For example, the observed state can be the M strongest RSRPs corresponding to SSBs measured using different beams for a beam management procedure.
[0113] In some examples, the adaptive learning algorithm is modeled as a POMDP using reinforcement learning. POMDP can be used when the state may not be known when an action is taken, and therefore the probability and / or reward may be unknown. For POMDP, reinforcement learning can be used. The Q array can be defined as:
[0114] Q i+1 (s, a) = Er + γmaxQ t (s′, a′)|s, a}.
[0115] like Figure 7 As shown, given a state 704s (e.g., RSRP) and a possible action a are input to a DNN 708, the DNN may execute an algorithm to output a value (e.g., parameter θ) according to the possible action a in order to determine a policy (e.g., π θ (s, a)) The policy and corresponding action are taken and applied to the environment. For example, agent 706 makes a beam selection and then uses the selected beam in environment 702. Figure 7 As shown, the reward for the action is fed back to the agent 706 to update the algorithm. For example, the throughput achieved by the selected beam can be fed back. Based on this feedback, the agent 706 updates the DNN 708 (e.g., by changing the weights associated with the vectors).
[0116] According to certain aspects, adaptive learning-based beam management allows for continuous and unlimited learning. In some examples, learning can be augmented by federated learning. For example, while some machine learning methods use centralized training data on a single machine or in a data center, with federated learning, learning can be collaborative, involving multiple devices to form a predictive model. With federated learning, model training can be done on the device through collaborative learning from multiple devices. For example, back to reference Figure 5-7 , node 520, agent 604, and agent 706 may receive training information and / or updated trained models from a variety of different devices.
[0117] In an illustrative example, beam management algorithms for multiple different UEs can be trained in multiple different operating scenarios, for example using deep reinforcement learning. The outputs from the training of different UEs can be combined to train the beam management algorithm for the UE. Once the beam management algorithm is trained, it can continue learning based on actual deployment scenarios. As discussed above, the state can be the best M RSRP metrics at the current time; the reward can be the measured PDSCH throughput for the current best beam pair; and the action can be the selection of which beam pair(s) to measure.
[0118] According to certain aspects, adaptive learning-based beam management allows for user personalization as well as design robustness. In some examples, adaptive learning-based beam management may be optimized. For example, as a user (e.g., such as node 520) visits / traverses a path (e.g., an environment), the adaptive algorithm learns and optimizes for the environment. In addition, different BS vendors may have different beam management implementations, such as how to transmit SSB. For example, some BS vendors transmit many narrow TX beams that will also be used as data beams; while other vendors transmit a few wide beams and use beam refinement (e.g., P2 and / or P3 procedures) to narrow and track the data beams. In some examples, adaptive learning-based beam management may be optimized for a specific beam management implementation for a vendor. In some examples, adaptive learning-based beam management may be optimized for a user, such as how the user holds / uses the UE, which affects the possible blocking of the UE's beam.
[0119] Figure 8 8 is a flow diagram illustrating example operations 800 for wireless communication in accordance with certain aspects of the present disclosure. Operations 800 may be performed, for example, by a node (e.g., such as node 520, which may be a wireless node, such as BS 110a or UE 120a in wireless communication network 100). Operations 800 may be implemented as a process executed on one or more processors (e.g., Figure 12 Furthermore, signal transmission and reception by the node in operation 800 may be performed by, for example, one or more antennas (e.g., Figure 12 In some aspects, transmission and / or reception of signals by a node may be implemented via a bus interface of one or more processors (e.g., controller / processor 1240, 1280) that obtain and / or output signals.
[0120] Operations 800 may begin, at 805, by determining, using adaptive learning, one or more beams to be used for a beam management procedure.
[0121] At 810, the node performs a beam management procedure using the determined one or more beams.
[0122] According to certain aspects, adaptive learning uses an adaptive learning algorithm. The adaptive learning algorithm may be updated (e.g., adapted) based on feedback and / or training information. The node may use the updated adaptive learning algorithm to perform another beam management procedure. The feedback may be feedback associated with the beam management procedure. For example, after performing a beam management procedure using the determined one or more beams, the node may receive feedback regarding the achieved throughput, and the beam management algorithm may be updated based on the feedback. In some examples, the feedback may be associated with beam management performed by different devices, such as different nodes.
[0123] Figure 9 is an example call flow diagram illustrating example signaling 900 for beam management using adaptive learning, in accordance with certain aspects of the present disclosure. Figure 9 As shown, at 908, UE 902 (e.g., such as UE 120a) may be provided with an initial learning algorithm (e.g., including a prediction model). In some examples, UE 902 may train the initial learning algorithm or the learning algorithm may be trained and then provided to UE 902. At 910, UE 902 performs a beam management procedure (e.g., such as P1 procedure 202) with one or more BSs 904. For example, UE 902 may use an adaptive learning algorithm to determine a beam to use and / or measure. At 912, UE 902 receives additional training information and / or feedback. For example, UE 902 may receive feedback from BS 904 (e.g., such as BS 110a) regarding the beam management procedure performed at 910, such as the PDSCH throughput achieved using the selected beam. Additionally or alternatively, UE 902 may receive additional training information from BS 904 and / or another UE 906. At 914, the UE 902 determines an updated adaptive learning algorithm based on the additional training information and / or feedback. At 916, the UE 902 may perform another beam management with the BS 904 (or another BS) using the updated adaptive learning algorithm.
[0124] In some examples, the training information (and / or feedback) includes training information obtained by deploying one or more UEs in one or more simulated communication environments prior to network deployment of the one or more UEs; training information obtained through feedback previously received when the one or more UEs were deployed in one or more communication environments (e.g., based on measurements and / or beam management procedures performed by the UE); training information from the network, one or more UEs and / or the cloud; and / or training information received when the node is online and / or idle.
[0125] In some examples, using the adaptive learning algorithm at 805 includes the node outputting an action based on one or more inputs; wherein feedback is associated with the action; and updating the adaptive learning algorithm based on the feedback includes adjusting one or more weights applied to the one or more inputs.
[0126] In some examples, the adaptive learning algorithm used by the node at 805 includes an adaptive machine learning algorithm; an adaptive reinforcement learning algorithm; an adaptive deep learning algorithm; an adaptive continuous infinite learning algorithm; and / or an adaptive policy optimization reinforcement learning algorithm. Figure 6-7 As discussed, the adaptive learning algorithm can be modeled as a POMDP. The adaptive learning algorithm can be implemented by an artificial neural network. In some examples, the artificial neural network can be a DQN that includes one or more DNNs. Using adaptive learning to determine one or more beams can include passing state parameters and action parameters through the one or more DNNs; for each state parameter, outputting a value for each action parameter; and selecting the action associated with the maximum output value. Updating the adaptive learning algorithm can include adjusting one or more weights associated with one or more neuronal connections in the artificial neural network.
[0127] In some examples, using adaptive learning at 805 to determine one or more beams to be used for the beam management procedure includes determining one or more beams to be included in a codebook based on the adaptive learning and selecting the one or more beams to be used for the beam management procedure from the codebook.
[0128] In some examples, using adaptive learning to determine one or more beams to use for the beam management procedure at 805 includes using adaptive learning to select one or more beams to use for the beam management procedure from a codebook.
[0129] In some examples, adaptive learning is used to select the BPL.
[0130] In some examples, adaptive learning uses a state parameter associated with a channel measurement, a reward parameter associated with received signal throughput or spectral efficiency, and an action parameter associated with selecting a beam pair corresponding to the channel measurement. In some examples, the channel measurement includes RSRP; spectral efficiency, channel flatness, and / or signal-to-noise ratio (SNR). In some examples, the received signal is a PDSCH transmission.
[0131] In some examples, the reward parameter is deducted by a penalty amount. In some examples, the penalty amount depends on the number of beams measured for the beam management procedure (e.g., beams used for transmission and / or reception of SSB). In some examples, the penalty amount depends on the amount of power consumption associated with the beam management procedure.
[0132] In some examples, performing a beam management procedure at 810 using the determined one or more beams includes measuring a channel based on an SSB transmission from a base station (BS) associated with a plurality of different transmit beams of the BS using the determined one or more beams; and selecting one or more BPLs associated with the channel measurements that are channel measurements above a channel measurement threshold and / or are one or more strongest channel measurements from all channel measurements associated with the measured SSB transmissions. In some examples, the determined one or more beams are a subset of available receive beams. In some examples, the node receives a PDSCH using one of the one or more selected BPLs; determines a throughput associated with the PDSCH; updates an adaptive learning algorithm based on the determined throughput; and uses the updated adaptive learning algorithm to determine another one or more beams to be used for performing another beam management procedure to select another one or more BPLs.
[0133] Figure 10 is an example call flow diagram illustrating example signaling 1000 for a BPL discovery procedure (e.g., such as P1 procedure 202) using adaptive learning, in accordance with certain aspects of the present disclosure. Figure 10 As shown, at 1008, UE 1002 (e.g., such as UE 120a) may perform initial training in a simulation environment prior to deployment at 1010. The initial training at 1008 may train an initial learning algorithm (e.g., including a prediction model) at UE 1002. At 1010, UE 1002 may be deployed in a network having at least one BS 1004 (e.g., such as BS 110a). UE 1002 may perform beam management procedures (e.g., such as P1 procedure 202) with one or more BSs 1004 in the network. For example, Figure 10As shown, at 1012, UE 1002 may use an adaptive learning algorithm to select a beam or RX / TX beam pair. At 1016, UE 1002 uses the beam(s) selected at 1012 to measure the SSB transmission(s) received at 1014 from BS 1004. At 1018, UE 1002 reports the measurements and / or BPL selection to BS 1004. Then, at 1020, BS 1004 transmits a PDSCH to UE 1002 using the BPL indicated by UE 1002 (or selected based on the measurements reported by UE 1002). At 1022, UE 1002 may determine the PDSCH throughput. The PDSCH throughput may serve as feedback or reinforcement for the adaptive learning. At 1026, UE 1002 updates the adaptive learning algorithm based on the feedback. Optionally, UE 1002 may receive additional training information and / or feedback from BS 1004 and / or another UE 1006 (e.g., UE 2) that UE 1002 may use to update the adaptive learning algorithm at 1026. UE 1002 may then perform another beam management with BS 1004 (or another BS) using the updated adaptive learning algorithm.
[0134] Figure 11 Illustrated may include operations configured to perform the techniques disclosed herein (such as, Figure 8 1 and 1 . The communication device 1100 includes various components (e.g., corresponding to means-plus-function components) of the present invention and the operations illustrated in the accompanying drawings. The communication device 1100 includes a processing system 1102 coupled to a transceiver 1108. The transceiver 1108 is configured to transmit and receive signals for the communication device 1100 (such as the various signals described herein) via an antenna 1110. The processing system 1102 can be configured to perform processing functions for the communication device 1100, including processing signals received and / or to be transmitted by the communication device 1100.
[0135] The processing system 1102 includes a processor 1104 coupled to a computer readable medium / memory 1112 via a bus 1106. In some aspects, the computer readable medium / memory 1112 is configured to store data that, when executed by the processor 1104, causes the processor 1104 to execute Figure 811 or instructions (e.g., computer executable code) for performing the operations illustrated in the 11 or other operations of the various techniques for adaptive learning-based beam management discussed herein. In certain aspects, the computer-readable medium / memory 1112 stores code 1114 for determining one or more beams to be used for a beam management procedure using adaptive learning; and code 1116 for performing the beam management procedure using the determined one or more beams. In certain aspects, the processor 1104 has circuitry configured to implement the code stored in the computer-readable medium / memory 1112. The processor 1104 includes circuitry 1118 for determining one or more beams to be used for a beam management procedure using adaptive learning; and circuitry 1120 for performing the beam management procedure using the determined one or more beams.
[0136] In some examples, the communication device 1100 may include a system on a chip (SOC) (not shown) that may include a central processing unit (CPU) or a multi-core CPU configured to perform adaptive learning-based beam management according to certain aspects of the present disclosure. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., a neural network with weights), delays, frequency bin information, and task information may be stored in a memory block associated with a neural processing unit (NPU), in a memory block associated with the CPU, in a memory block associated with a digital signal processor (DSP), in different memory blocks, or may be distributed across multiple memory blocks. Instructions executed at the CPU may be loaded from a program memory associated with the CPU or may be loaded from different memory blocks.
[0137] In some examples, the adaptive learning-based beam management described herein can allow for improved P1 procedures by adaptively updating the beam management algorithm so that beam selection can be refined to more intelligently select beams to measure based on learning. Thus, the UE can find the BPL while measuring fewer beams.
[0138] Each method disclosed herein includes one or more steps or actions for implementing the method. These method steps and / or actions may be interchangeable with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of the specific steps and / or actions may be modified without departing from the scope of the claims.
[0139] As used herein, a phrase referring to "at least one" of a list of items refers to any combination of those items, including individual members. By way of example, "at least one of a, b, or c" is intended to encompass: a, b, c, ab, ac, bc, and abc, as well as any combination with multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other ordering of a, b, and c).
[0140] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or another data structure), ascertaining, and the like. Furthermore, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. Furthermore, "determining" may include resolving, selecting, choosing, establishing, and the like.
[0141] Figure 12 Illustrated are BS 110a and UE 120a (eg, in Figure 1 For example, antenna 1252, processors 1266, 1258, 1264, and / or controller / processor 1280 of UE 120a and / or antenna 1234, processors 1220, 1230, 1238, and / or controller / processor 1240 of BS 110a may be used to implement the various techniques and methods described herein. Figure 12 As shown, the controller / processor 1280 of the UE 120a has a beam selection manager 1281 that can be configured to use adaptive learning to determine, for example, beams to be used for beam management procedures in accordance with various aspects described herein. Figure 12 As shown, additionally or alternatively, the controller / processor 1240 of the BS 110a may have a beam selection manager 1241 that may be configured to determine beams using adaptive learning according to aspects described herein.
[0142] At BS 110a, a transmit processor 1220 may receive data from a data source 1212 and control information from a controller / processor 1240. The control information may be for a physical broadcast channel (PBCH), a physical control format indicator channel (PCFICH), a physical hybrid ARQ indicator channel (PHICH), a physical downlink control channel (PDCCH), a group common PDCCH (GC PDCCH), etc. The data may be for a physical downlink shared channel (PDSCH), etc. The processor 1220 may process (e.g., encode and symbol map) the data and control information to obtain data symbols and control symbols, respectively. The transmit processor 1220 may also generate reference symbols, such as for a primary synchronization signal (PSS), a secondary synchronization signal (SSS), and a cell-specific reference signal (CRS). A transmit (TX) multiple-input, multiple-output (MIMO) processor 1230 may perform spatial processing (e.g., precoding) on data symbols, control symbols, and / or reference symbols, as applicable, and may provide output symbol streams to modulators (MODs) 1232a-1232t. Each modulator 1232 may process a respective output symbol stream (e.g., for OFDM, etc.) to obtain an output sample stream. Each modulator may further process (e.g., convert to analog, amplify, filter, and frequency upconvert) the output sample stream to obtain a downlink signal. The downlink signals from modulators 1232a-1232t may be transmitted via antennas 1234a-1234t, respectively.
[0143] At UE 120a, antennas 1252a-1252r may receive downlink signals from BS 110a and may provide received signals to demodulators (DEMODs) in transceivers 1254a-1254r, respectively. Each demodulator 1254 may condition (e.g., filter, amplify, downconvert, and digitize) its respective received signal to obtain input samples. Each demodulator may further process the input samples (e.g., for OFDM, etc.) to obtain received symbols. A MIMO detector 1256 may receive received symbols from all demodulators 1254a-1254r, perform MIMO detection on the received symbols where applicable, and provide detected symbols. A receive processor 1258 may process (e.g., demodulate, deinterleave, and decode) the detected symbols, provide decoded data for UE 120a to a data sink 1260, and provide decoded control information to a controller / processor 1280.
[0144] On the uplink, at the UE 120a, a transmit processor 1264 may receive and process data from a data source 1262 (e.g., for a physical uplink shared channel (PUSCH)) and control information from the controller / processor 1280 (e.g., for a physical uplink control channel (PUCCH)). The transmit processor 1264 may also generate reference symbols for reference signals (e.g., a sounding reference signal (SRS)). The symbols from the transmit processor 1264 may be precoded by a TX MIMO processor 1266, if applicable, further processed by the demodulators 1254a-1254r in the transceiver (e.g., for SC-FDM, etc.), and transmitted to the base station 110. At BS 110a, the uplink signal from UE 120a may be received by antenna 1234, processed by modulator 1232, detected by MIMO detector 1236 if applicable, and further processed by receive processor 1238 to obtain decoded data and control information sent by UE 120a. Receive processor 1238 may provide the decoded data to a data sink 1239 and the decoded control information to a controller / processor 1240.
[0145] Controllers / processors 1240 and 1280 may direct the operations at BS 110a and UE 120a, respectively. Controller / processor 1240 and / or other processors and modules at BS 110a may perform or direct the execution of processes for the techniques described herein. Memories 1242 and 1282 may store data and program codes for BS 110a and UE 120a, respectively. Scheduler 1244 may schedule UEs for data transmission on the downlink and / or uplink.
[0146] The techniques described herein can be used for various wireless communication technologies, such as 3GPP Long Term Evolution (LTE), Advanced LTE (LTE-A), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single Carrier Frequency Division Multiple Access (SC-FDMA), Time Division Synchronous Code Division Multiple Access (TD-SCDMA), and other networks. The terms "network" and "system" are often used interchangeably. A CDMA network can implement radio technologies such as Universal Terrestrial Radio Access (UTRA) and cdma2000. UTRA includes Wideband CDMA (WCDMA) and other variants of CDMA. cdma2000 covers IS-2000, IS-95, and IS-856 standards. A TDMA network can implement radio technologies such as Global System for Mobile Communications (GSM). OFDMA networks can implement radio technologies such as NR (e.g., 5G RA), Evolved UTRA (E-UTRA), Ultra Mobile Broadband (UMB), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), IEEE 802.20, Flash-OFDMA, and others. UTRA and E-UTRA are parts of the Universal Mobile Telecommunications System (UMTS). LTE and LTE-A are versions of UMTS that use E-UTRA. UTRA, E-UTRA, UMTS, LTE, LTE-A, and GSM are described in documents from an organization called the 3rd Generation Partnership Project (3GPP). cdma2000 and UMB are described in documents from an organization called the 3rd Generation Partnership Project 2 (3GPP2). NR is an emerging wireless communication technology being developed in collaboration with the 5G Technology Forum (5GTF). NR access (e.g., 5G NR) can support various wireless communication services, such as mmW. NR can utilize OFDM with CP on both the uplink and downlink and includes support for half-duplex operation using TDD. Beamforming can be supported and the beam direction can be dynamically configured. MIMO transmission with precoding can also be supported. In some examples, MIMO configurations in the DL can support up to 8 transmit antennas (with multi-layer DL transmission of up to 8 streams) and up to 2 streams per UE. In some examples, multi-layer transmission of up to 2 streams per UE can be supported.
[0147] In 3GPP, the term "cell" can refer to the coverage area of a Node B (NB) and / or the NB subsystem serving that coverage area, depending on the context in which the term is used. In NR systems, the terms "cell," BS, next-generation Node B (gNB or g-Node B), access point (AP), distributed unit (DU), carrier, or transmit reception point (TRP) can be used interchangeably. In some examples, a cell may not necessarily be stationary, and the geographic area of a cell may move depending on the location of a mobile BS.
[0148] A UE may also be referred to as a mobile station, a terminal, an access terminal, a subscriber unit, a station, a customer premises equipment (CPE), a cellular phone, a smartphone, a personal digital assistant (PDA), a wireless modem, a wireless communication device, a handheld device, a laptop computer, a cordless phone, a wireless local loop (WLL) station, a tablet computer, a camera, a gaming device, a netbook, a smartbook, an ultrabook, an appliance, a medical device or medical equipment, a biometric sensor / device, a wearable device (such as a smart watch, smart clothing, smart glasses, a smart wristband, smart jewelry (e.g., a smart ring, a smart bracelet, etc.)), an entertainment device (e.g., a music device, a video device, a satellite radio, etc.), a vehicle component or sensor, a smart meter / sensor, industrial manufacturing equipment, a global positioning system device, or any other suitable device configured to communicate via a wireless or wired medium. Some UEs may be considered machine type communication (MTC) devices or evolved MTC (eMTC) devices. MTC and eMTC UEs include, for example, robots, drones, remote devices, sensors, meters, monitors, location tags, etc., which can communicate with a base station, another device (e.g., a remote device), or some other entity. Nodes such as wireless nodes can provide connectivity to or to a network (e.g., a wide area network (such as the Internet) or a cellular network) via, for example, a wired or wireless communication link. Some UEs may be considered Internet of Things (IoT) devices, which may be narrowband IoT (NB-IoT) devices.
[0149] The techniques described herein can be used for the wireless networks and radio technologies mentioned above as well as other wireless networks and radio technologies. For clarity, although various aspects may be described herein using terms typically associated with 3G, 4G, and / or 5G wireless technologies, various aspects of the present disclosure may be applied in communication systems based on other generations.
[0150] The preceding description is provided to enable anyone skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the universal principles defined herein may be applied to other aspects. Accordingly, the claims are not intended to be limited to the aspects shown herein, but rather should be granted the full scope consistent with the claim language, wherein singular references to elements are not intended to mean "one and only one" (unless specifically stated otherwise) but rather "one or more." Unless specifically stated otherwise, the term "some" refers to one or more. All structural and functional equivalents of the elements described throughout this disclosure, now or hereafter known to those of ordinary skill in the art, are expressly incorporated herein by reference and are intended to be covered by the claims. Furthermore, nothing disclosed herein is intended to be dedicated to the public, regardless of whether such disclosure is explicitly stated in the claims. No element of a claim should be interpreted under 35 U.S.C. §112(f) unless the element is explicitly recited using the phrase "means for..." or, in the case of a method claim, the element is recited using the phrase "step for..."
[0151] The various operations of the methods described above may be performed by any suitable device capable of performing the corresponding functions. These devices may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, where there are operations illustrated in the figures, these operations may have corresponding counterpart means-plus-function components with similar numbering.
[0152] The various illustrative logical blocks, modules, and circuits described in conjunction with the present disclosure may be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof, designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0153] If implemented in hardware, an example hardware configuration may include a processing system in a node. The processing system may be implemented using a bus architecture. Depending on the specific application and overall design constraints of the processing system, the bus may include any number of interconnecting buses and bridges. The bus may link together various circuits including a processor, a machine-readable medium, and a bus interface. The bus interface may be used to connect a network adapter, etc., to the processing system via the bus. The network adapter may be used to implement signal processing functions at the PHY layer. In UE 120a (see Figure 1 ), a user interface (e.g., a keypad, display, mouse, joystick, etc.) may also be connected to the bus. The bus may also link various other circuits, such as timing sources, peripherals, voltage regulators, power management circuits, and the like, which are well known in the art and will not be described further. The processor may be implemented using one or more general and / or special purpose processors. Examples include microprocessors, microcontrollers, DSP processors, and other circuitry capable of executing software. Those skilled in the art will recognize how to best implement the functionality described with respect to the processing system, depending on the specific application and the overall design constraints imposed on the overall system.
[0154] If implemented in software, each function may be stored as one or more instructions or codes on a computer-readable medium or transmitted via it. Software should be broadly interpreted to mean instructions, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Computer-readable media include both computer storage media and communication media, including any media that facilitates the transfer of computer programs from one location to another. The processor may be responsible for managing the bus and general processing, including executing software modules stored on a machine-readable storage medium. A computer-readable storage medium may be coupled to the processor so that the processor can read and write information from / to the storage medium. In an alternative, the storage medium may be integrated into the processor. As an example, the machine-readable medium may include a transmission line, a carrier modulated by data, and / or a computer-readable storage medium having instructions stored thereon that is separate from the node, all of which may be accessed by the processor via a bus interface. Alternatively or additionally, the machine-readable medium or any portion thereof may be integrated into the processor, such as a cache and / or general register file. As examples, examples of machine-readable storage media may include RAM (random access memory), flash memory, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), registers, magnetic disks, optical disks, hard drives, or any other suitable storage media, or any combination thereof. Machine-readable media may be embodied in a computer program product.
[0155] A software module may include a single instruction or many instructions and may be distributed across several different code segments, between different programs, and across multiple storage media. A computer-readable medium may include several software modules. These software modules include instructions that, when executed by a device (such as a processor), cause a processing system to perform various functions. These software modules may include a transmitting module and a receiving module. Each software module may reside in a single storage device or be distributed across multiple storage devices. As an example, when a triggering event occurs, a software module may be loaded from a hard drive into RAM. During the execution of the software module, the processor may load some instructions into a cache to increase access speed. One or more cache lines may then be loaded into a general register file for execution by the processor. When describing the functionality of a software module below, it will be understood that such functionality is implemented by the processor when the processor executes instructions from the software module.
[0156] Any connection is also properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared (IR), radio, and microwave), then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Thus, in some aspects, computer-readable media may include non-transitory computer-readable media (e.g., tangible media). Additionally, for other aspects, computer-readable media may include transient computer-readable media (e.g., signals). Combinations of the above should also be included within the scope of computer-readable media.
[0157] Thus, certain aspects may include a computer program product for performing the operations presented herein. For example, such a computer program product may include a computer-readable medium having stored (and / or encoded) thereon instructions that can be executed by one or more processors to perform the operations described herein, such as for performing the operations described herein and in Figure 8 Instructions for the operations explained in .
[0158] In addition, it should be appreciated that the modules and / or other appropriate means for performing the methods and techniques described herein can be downloaded and / or otherwise obtained by the user terminal and / or base station where applicable. For example, such a device can be coupled to a server to facilitate the transfer of the means for performing the methods described herein. Alternatively, the various methods described herein can be provided via a storage device (e.g., RAM, ROM, a physical storage medium such as a compact disc (CD) or floppy disk, etc.) so that once the storage device is coupled to or provided to the user terminal and / or base station, the device can obtain the various methods. In addition, any other suitable technology suitable for providing the methods and techniques described herein to a device can be utilized.
[0159] It will be understood that the claims are not limited to the precise configuration and components illustrated above. Various changes, substitutions and variations may be made in the arrangement, operation and details of the methods and apparatus described above without departing from the scope of the claims.
Claims
1. A method for wireless communication by a user equipment (UE), comprising: determining one or more beams to use for a beam management procedure using an adaptive learning algorithm, wherein using the adaptive learning algorithm comprises outputting an action based on one or more inputs; performing the beam management procedure using the determined one or more beams; receiving feedback associated with the beam management procedure using the determined one or more beams, wherein the feedback is associated with the action; as well as The adaptive learning algorithm is updated based on the feedback, wherein updating the adaptive learning algorithm based on the feedback comprises adjusting one or more weights applied to the one or more inputs.
2. The method of claim 1 , wherein performing the beam management procedure using the determined one or more beams comprises: measuring a channel based on one or more synchronization signal block (SSB) transmissions from a base station (BS) using the determined one or more beams, the SSB transmissions being associated with one or more transmit beams of the BS; as well as One or more beam pair links (BPLs) are selected, the one or more BPLs being associated with one or more channel measurements above a channel measurement threshold; or one or more strongest channel measurements of all channel measurements associated with the one or more SSB transmissions; or a combination thereof.
3. The method of claim 2, further comprising: receiving a physical downlink shared channel (PDSCH) using the selected one of the one or more BPLs; determining a throughput, a spectral efficiency, or both associated with the PDSCH, wherein the feedback is the determined throughput, spectral efficiency, or both; and The updated adaptive learning algorithm is used to determine another beam or beams to be used to perform another beam management procedure to select another BPL or BPLs.
4. The method of claim 2, wherein measuring the channel comprises measuring reference signal received power (RSRP); spectral efficiency, channel flatness, or signal-to-noise ratio (SNR); or a combination thereof.
5. The method of claim 1, further comprising: The adaptive learning algorithm is determined, the adaptive learning algorithm is updated, or both, based on training information.
6. The method of claim 5, wherein the training information comprises: training information obtained by deploying the one or more UEs in one or more simulated communication environments prior to network deployment of the one or more UEs; training information obtained from feedback previously received when the one or more UEs were deployed in one or more communication environments; training information from at least one of the network, one or more UEs, or a cloud; Training information received while the UE is at least one of online and idle; or Its combination.
7. A method as claimed in claim 5, wherein the training information includes training information received from one or more UEs different from the UE after the UE is deployed, wherein the training information includes information associated with beam measurements performed by the one or more UEs, or feedback associated with one or more beam management procedures performed by the one or more UEs, or a combination thereof.
8. The method of claim 1 , wherein the adaptive learning algorithm comprises an adaptive machine learning algorithm; Adaptive reinforcement learning algorithms; Adaptive deep learning algorithms; An adaptive continuous infinite learning algorithm; or an adaptive policy optimization reinforcement learning algorithm, or a combination thereof.
9. The method of claim 1, wherein the adaptive learning algorithm is modeled as a partially observable Markov decision process (POMDP).
10. The method of claim 1, wherein the adaptive learning algorithm is implemented by an artificial neural network.
11. The method of claim 10, wherein: The artificial neural network comprises a deep Q network (DQN) comprising one or more deep neural networks (DNNs); and Using an adaptive learning algorithm to determine one or more beams includes: passing one or more state parameters and one or more action parameters through the one or more DNNs; For each state parameter, output the value of each action parameter; and Select the action associated with the maximum output value.
12. The method of claim 10, wherein updating the adaptive learning algorithm comprises adjusting one or more weights associated with one or more neuronal connections in the artificial neural network.
13. The method of claim 1 , wherein using an adaptive learning algorithm to determine one or more beams to use for a beam management procedure comprises: determining one or more beams to include in a codebook based on the adaptive learning; as well as One or more beams are selected from the codebook to be used for a beam management procedure.
14. The method of claim 1, wherein determining one or more beams to use for a beam management procedure comprises using an adaptive learning algorithm to select one or more beams to use for the beam management procedure from a codebook.
15. The method of claim 1, wherein the adaptive learning algorithm uses a state parameter associated with a channel measurement, a reward parameter associated with received signal throughput or spectral efficiency, and an action parameter associated with selecting a beam pair corresponding to the channel measurement. The method of claim 15 , wherein the reward parameter is deducted by a penalty amount.
17. The method of claim 16, wherein the penalty amount depends on the number of one or more beams measured for the beam management procedure.
18. The method of claim 16, wherein the amount of penalty depends on an amount of power consumption associated with the beam management procedure.
19. The method of claim 1, wherein the determined one or more beams comprise a subset of available receive beams.
20. An apparatus for wireless communication by a user equipment (UE), comprising: Memory; as well as a processor coupled to the memory, the processor and the memory being configured to: determining one or more beams to use for a beam management procedure using an adaptive learning algorithm that outputs an action based on one or more inputs; performing the beam management procedure using the determined one or more beams; receiving feedback associated with the beam management procedure using the determined one or more beams, wherein the feedback is associated with the action; as well as One or more weights applied to the one or more inputs are adjusted based on the feedback to update the adaptive learning algorithm.
21. The apparatus of claim 20, wherein when configured to perform a beam management procedure using the determined one or more beams, the memory and the processor are configured to: measuring a channel based on one or more synchronization signal block (SSB) transmissions from a base station (BS) using the determined one or more beams, the SSB transmissions being associated with one or more transmit beams of the BS; and One or more beam pair links (BPLs) are selected, the one or more BPLs being associated with one or more channel measurements above a channel measurement threshold; or one or more strongest channel measurements of all channel measurements associated with the one or more SSB transmissions; or a combination thereof.
22. The apparatus of claim 21 , wherein the processor and the memory are further configured to: receiving a physical downlink shared channel (PDSCH) using the selected one of the one or more BPLs; determining a throughput, a spectral efficiency, or both associated with the PDSCH, wherein the feedback is the determined throughput, spectral efficiency, or both; and The updated adaptive learning algorithm is used to determine another beam or beams to be used to perform another beam management procedure to select another BPL or BPLs.
23. The apparatus of claim 21 , wherein the processor and the memory are configured to measure a channel, wherein when configured to measure a channel, the memory and the processor are configured to measure reference signal received power (RSRP); frequency efficiency, channel flatness, or signal-to-noise ratio (SNR); or a combination thereof.
24. The apparatus of claim 20, wherein the processor and the memory are further configured to: The adaptive learning algorithm is determined, the adaptive learning algorithm is updated, or both, based on training information.
25. The apparatus of claim 24, wherein the training information comprises: training information obtained by deploying the one or more UEs in one or more simulated communication environments prior to network deployment of the one or more UEs; training information obtained from feedback previously received when the one or more UEs were deployed in one or more communication environments; training information from at least one of the network, one or more UEs, or a cloud; Training information received while the device is at least one of online or idle; or Its combination.
26. An apparatus as described in claim 24, wherein the training information includes training information received from one or more UEs different from the apparatus after the apparatus is deployed, wherein the training information includes information associated with beam measurements performed by the one or more UEs, or feedback associated with one or more beam management procedures performed by the one or more UEs, or a combination thereof.
27. An apparatus for wireless communication by a user equipment (UE), comprising: means for determining one or more beams to use for a beam management procedure using an adaptive learning algorithm, wherein using the adaptive learning algorithm comprises outputting an action based on one or more inputs; means for performing said beam management procedure using the determined one or more beams; means for receiving feedback associated with the beam management procedure using the determined one or more beams, wherein the feedback is associated with the action; as well as Means for updating the adaptive learning algorithm based on the feedback, wherein updating the adaptive learning algorithm based on the feedback comprises adjusting one or more weights applied to the one or more inputs.
28. A non-transitory computer-readable medium having stored thereon computer-executable code for wireless communication by a user equipment (UE), the computer-executable code comprising: code for determining one or more beams to use for a beam management procedure using an adaptive learning algorithm, wherein using the adaptive learning algorithm comprises outputting an action based on one or more inputs; code for performing the beam management procedure using the determined one or more beams; code for receiving feedback associated with the beam management procedure using the determined one or more beams, wherein the feedback is associated with the action; as well as Code for updating the adaptive learning algorithm based on the feedback, wherein updating the adaptive learning algorithm based on the feedback comprises adjusting one or more weights applied to the one or more inputs.