Apparatus and method for managing beam by using artificial intelligence model in wireless communication system

AI/ML models with reinforcement learning improve beam management in wireless communication systems by adapting feedback and resource allocation, enhancing efficiency and reducing overhead.

WO2026029214A1PCT designated stage Publication Date: 2026-02-05LG ELECTRONICS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/011046
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing wireless communication systems face challenges in optimizing beam management, particularly in managing overhead, determining feedback information formats, and adapting properties for efficient beam management, especially in environments demanding greater communication capacity and reliability.

Method used

The implementation of an artificial intelligence/machine learning (AI/ML) model for beam management, utilizing reinforcement learning to adaptively adjust feedback properties, determine feedback information formats, and manage resources based on required amounts, combining rewards for optimal beam management.

Benefits of technology

Enhances beam acquisition and tracking efficiency, optimizing resource utilization and reducing overhead in wireless communication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024011046_05022026_PF_FP_ABST
    Figure KR2024011046_05022026_PF_FP_ABST
Patent Text Reader

Abstract

The objective of the present disclosure is to manage a beam by using an artificial intelligence model in a wireless communication system. This method performed by a terminal may comprise the steps of: receiving, from a base station, a first message for requesting capability information; transmitting, to the base station, a second message including the capability information; receiving, from the base station, first configuration information related to training of an artificial intelligence / machine learning (AI / ML) model for determining feedback information for beam management; receiving, from the base station, second configuration information related to at least one reference signal; receiving the at least one reference signal on the basis of the second configuration information; and performing an operation related to the training on the basis of the at least one reference signal.
Need to check novelty before this filing date? Find Prior Art

Description

Device and method for managing beams using artificial intelligence models in wireless communication systems

[0001] The present disclosure relates to a wireless communication system, and to a device and method for managing a beam using an artificial intelligence model in a wireless communication system.

[0002] Wireless access systems are widely deployed to provide various types of communication services, such as voice and data. Typically, wireless access systems are multiple access systems that support communications with multiple users by sharing available system resources (e.g., bandwidth, transmission power). Examples of multiple access systems include code division multiple access (CDMA), frequency division multiple access (FDMA), time division multiple access (TDMA), orthogonal frequency division multiple access (OFDMA), and single-carrier frequency division multiple access (SC-FDMA).

[0003] In particular, as numerous communication devices demand greater communication capacity, enhanced mobile broadband (eMBB) communication technologies are being proposed, improving upon existing radio access technology (RAT). Furthermore, wireless communication systems that consider reliability and latency-sensitive services / user equipment (UE) as well as massive machine type communications (mMTC), which connects multiple devices and objects to provide diverse services anytime and anywhere, are being proposed. Various technological configurations are being proposed for these solutions.

[0004] The present disclosure relates to a device and method for managing a beam using an artificial intelligence / machine learning (AI / ML) model in a wireless communication system.

[0005] The present disclosure relates to a device and method for optimizing overhead for beam management using an AI / ML model in a wireless communication system.

[0006] The present disclosure relates to a device and method for determining a format of feedback information for beam management in a wireless communication system.

[0007] The present disclosure relates to a device and method for adaptively adjusting properties for feedback for beam management in a wireless communication system.

[0008] The present disclosure relates to a device and method for adaptively controlling feedback for beam management using reinforcement learning in a wireless communication system.

[0009] The present disclosure relates to a device and method for providing a set of actions for reinforcement learning related to beam management in a wireless communication system.

[0010] The present disclosure relates to a device and method for performing exploitation and exploration using a set of activities for reinforcement learning related to beam management in a wireless communication system.

[0011] The present disclosure relates to a device and method for managing elements within an activity set for reinforcement learning related to beam management in a wireless communication system according to the amount of resources required.

[0012] The present disclosure relates to a device and a method for defining rewards for reinforcement learning related to beam management in a wireless communication system in terms of quality and cost.

[0013] The present disclosure relates to a device and method for combining a plurality of rewards defined for reinforcement learning related to beam management in a wireless communication system into a single reward value.

[0014] The technical objectives to be achieved in the present disclosure are not limited to those mentioned above, and other technical tasks not mentioned can be considered by a person having ordinary skill in the technical field to which the technical configuration of the present disclosure is applied from the embodiments of the present disclosure described below.

[0015] As an example of the present disclosure, a method performed by a terminal in a wireless communication system includes the steps of: receiving a first message requesting capability information from a base station; transmitting a second message including the capability information to the base station; receiving first configuration information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management from the base station; receiving second configuration information related to at least one reference signal from the base station; receiving the at least one reference signal based on the second configuration information; and performing an operation related to the learning based on the at least one reference signal, wherein the first configuration information may include information related to a set including, as elements, formats of a reference signal or report information corresponding to the reference signal as a set of actions for reinforcement learning of the AI / ML model.

[0016] As an example of the present disclosure, a method performed by a base station in a wireless communication system includes the steps of: transmitting a first message requesting capability information to a terminal; receiving a second message including capability information from the terminal; transmitting, to the terminal, first configuration information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management; transmitting, to the terminal, second configuration information related to at least one reference signal; transmitting, based on the second configuration information, the at least one reference signal; and performing an operation related to the learning based on the at least one reference signal, wherein the first configuration information may include information related to a set including, as elements, formats of a reference signal or report information corresponding to the reference signal as a set of actions for reinforcement learning of the AI / ML model.

[0017] As an example of the present disclosure, in a wireless communication system, a terminal includes a transceiver and a processor coupled to the transceiver, wherein the processor is configured to receive a first message requesting capability information from a base station, transmit a second message including the capability information to the base station, receive first configuration information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management from the base station, receive second configuration information related to at least one reference signal from the base station, receive the at least one reference signal based on the second configuration information, and perform an operation related to the learning based on the at least one reference signal, wherein the first configuration information may include information related to a set including, as elements, formats of a reference signal or report information corresponding to the reference signal as a set of actions for reinforcement learning of the AI / ML model.

[0018] As an example of the present disclosure, a communication device includes at least one processor, and at least one computer memory connected to the at least one processor and storing instructions that, when executed by the at least one processor, direct operations, the operations include: receiving a first message requesting capability information from a base station; transmitting a second message including the capability information to the base station; receiving first configuration information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management from the base station; receiving second configuration information related to at least one reference signal from the base station; receiving the at least one reference signal based on the second configuration information; and performing an operation related to the learning based on the at least one reference signal, wherein the first configuration information may include information related to a set of actions for reinforcement learning of the AI / ML model, the set including, as elements, formats of a reference signal or report information corresponding to the reference signal.

[0019] As an example of the present disclosure, a non-transitory computer-readable medium storing at least one instruction includes at least one instruction executable by a processor, wherein the at least one instruction instructs a device to receive a first message requesting capability information from a base station, transmit a second message including the capability information to the base station, receive first configuration information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management from the base station, receive second configuration information related to at least one reference signal from the base station, receive the at least one reference signal based on the second configuration information, and perform an operation related to the learning based on the at least one reference signal, wherein the first configuration information may include information related to a set of actions for reinforcement learning of the AI / ML model, the set including, as elements, formats of a reference signal or report information corresponding to the reference signal.

[0020] The above-described aspects of the present disclosure are only some of the preferred embodiments of the present disclosure, and various embodiments reflecting the technical features of the present disclosure can be derived and understood by a person having ordinary skill in the art based on the detailed description of the present disclosure to be described below.

[0021] The following effects may be achieved by embodiments based on the present disclosure.

[0022] According to the present disclosure, beam management operations such as beam acquisition and beam tracking can be effectively performed.

[0023] The effects that can be obtained from the embodiments of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly derived and understood by those skilled in the art to which the technical configuration of the present disclosure is applied, from the description of the embodiments of the present disclosure below. In other words, unintended effects resulting from implementing the configuration described in the present disclosure can also be derived from the embodiments of the present disclosure by those skilled in the art.

[0024] The accompanying drawings are intended to aid understanding of the present disclosure and, together with detailed descriptions, may provide embodiments of the present disclosure. However, the technical features of the present disclosure are not limited to specific drawings, and the features disclosed in each drawing may be combined with each other to form new embodiments. Reference numerals in each drawing may indicate structural elements.

[0025] Figure 1 illustrates an example of a wireless communication system applicable to the present disclosure.

[0026] FIG. 2 illustrates an example of a wireless device applicable to the present disclosure.

[0027] FIG. 3 illustrates a method for processing a transmission signal applicable to the present disclosure.

[0028] Figure 4 illustrates a communication procedure between a terminal and a base station applicable to the present disclosure.

[0029] FIG. 5 illustrates an example of a communication structure that can be provided in a 6G (6th generation) system applicable to the present disclosure.

[0030] Figure 6 illustrates an electromagnetic spectrum applicable to the present disclosure.

[0031] Figure 7 illustrates a transmitter structure applicable to the present disclosure.

[0032] Figure 8 illustrates an example of a functional framework for application of artificial intelligence technology applicable to the present disclosure.

[0033] Figure 9 illustrates an example of a procedure for utilizing an artificial intelligence model applicable to the present disclosure.

[0034] Figure 10 illustrates a communication procedure based on AI (artificial intelligence) technology applicable to the present disclosure.

[0035] Figure 11 illustrates an example of an AI / ML (machine learning) model for beam prediction.

[0036] FIG. 12a and FIG. 12b illustrate examples of deployment of AI / ML models according to one embodiment of the present disclosure.

[0037] FIG. 13 illustrates the structure of reinforcement learning according to one embodiment of the present disclosure.

[0038] Figure 14 shows the Pareto optimal region of beam performance and feedback cost.

[0039] FIG. 15 illustrates an example of a procedure for providing capability information according to one embodiment of the present disclosure.

[0040] FIG. 16 illustrates an example of a procedure for collecting data using a downlink reference signal in a base station-side model situation according to one embodiment of the present disclosure.

[0041] FIG. 17 illustrates an example of a procedure for collecting data using an uplink reference signal in a base station-side model situation according to one embodiment of the present disclosure.

[0042] FIG. 18 illustrates an example of a procedure for collecting data using a downlink reference signal in a terminal-side model situation according to one embodiment of the present disclosure.

[0043] FIG. 19 illustrates an example of a procedure for collecting data using an uplink reference signal in a terminal-side model situation according to one embodiment of the present disclosure.

[0044] FIG. 20 illustrates an example of a procedure for performing learning for a terminal-side model according to one embodiment of the present disclosure.

[0045] FIG. 21 illustrates an example of a procedure for performing learning for a base station-side model according to one embodiment of the present disclosure.

[0046] Figures 22a and 22b illustrate the Pareto optimal region of beam performance and feedback cost depending on whether the feedback period is fixed.

[0047] Figure 23 illustrates an example of a wireless device applicable to the present disclosure.

[0048] Figure 24 illustrates an example of a portable device applicable to the present disclosure.

[0049] FIG. 25 illustrates an example of a vehicle or autonomous vehicle applicable to the present disclosure.

[0050] Figure 26 illustrates an example of a vehicle applicable to the present disclosure.

[0051] FIG. 27 illustrates an example of an extended reality (XR) device applicable to the present disclosure.

[0052] Figure 28 illustrates an example of a robot applicable to the present disclosure.

[0053] Figure 29 illustrates an example of an AI device applicable to the present disclosure.

[0054] The following embodiments combine the components and features of the present disclosure in a predetermined form. Each component or feature may be considered optional unless explicitly stated otherwise. Each component or feature may be implemented without being combined with other components or features. Furthermore, some components and / or features may be combined to form embodiments of the present disclosure. The order of operations described in the embodiments of the present disclosure may be changed. Some components or features of one embodiment may be included in another embodiment or may be replaced with corresponding components or features of another embodiment.

[0055] In the description of the drawings, procedures or steps that may obscure the gist of the present disclosure are not described, and procedures or steps that can be understood by a person skilled in the art are also not described.

[0056] Throughout the specification, when a part is said to "comprising" or "including" a component, this does not mean that other components may be included, but rather that other components may be excluded, unless otherwise specifically stated. In addition, terms such as "...part," "...unit," and "module" described in the specification mean a unit that processes at least one function or operation, which may be implemented by hardware, software, or a combination of hardware and software. In addition, the words "a" or "an," "one," "the," and similar related words may be used in the context of describing the present disclosure (especially in the context of the claims below) to include both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context.

[0057] Embodiments of the present disclosure described herein focus on the data transmission and reception relationship between a base station and a mobile station. Here, the base station is understood as a terminal node of a network that directly communicates with the mobile station. Certain operations described herein as being performed by the base station may, in some cases, be performed by an upper node of the base station.

[0058] That is, in a network consisting of multiple network nodes including a base station, various operations performed for communication with a mobile station may be performed by the base station or other network nodes other than the base station. In this case, the term 'base station' may be replaced by terms such as fixed station, Node B, eNB (eNode B), gNB (gNode B), ng-eNB, advanced base station (ABS), or access point.

[0059] Additionally, in the embodiments of the present disclosure, the term terminal may be replaced with terms such as user equipment (UE), mobile station (MS), subscriber station (SS), mobile subscriber station (MSS), mobile terminal, or advanced mobile station (AMS).

[0060] Additionally, a transmitter refers to a fixed and / or mobile node that provides data or voice services, and a receiver refers to a fixed and / or mobile node that receives data or voice services. Therefore, for uplink, a mobile station can be the transmitter, and a base station can be the receiver. Similarly, for downlink, a mobile station can be the receiver, and a base station can be the transmitter.

[0061] Embodiments of the present disclosure may be supported by standard documents disclosed in at least one of wireless access systems, such as IEEE 802.xx system, 3rd Generation Partnership Project (3GPP) system, 3GPP Long Term Evolution (LTE) system, 3GPP 5th generation (5G) NR (New Radio) system and 3GPP2 system, and in particular, embodiments of the present disclosure may be supported by 3GPP TS (technical specification) 38.211, 3GPP TS 38.212, 3GPP TS 38.213, 3GPP TS 38.321 and 3GPP TS 38.331 documents.

[0062] Furthermore, the embodiments of the present disclosure can be applied to other wireless access systems and are not limited to the systems described above. For example, they can be applied to systems implemented after the 3GPP 5G NR system and are not limited to a specific system.

[0063] That is, obvious steps or parts not described in the embodiments of the present disclosure can be explained by referring to the above documents. In addition, all terms disclosed in this document can be explained by the above standard documents.

[0064] Hereinafter, preferred embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. The detailed description set forth below, together with the accompanying drawings, is intended to illustrate exemplary embodiments of the present disclosure and is not intended to represent the only embodiments in which the technical configurations of the present disclosure may be implemented.

[0065] Additionally, specific terms used in the embodiments of the present disclosure are provided to aid in understanding of the present disclosure, and the use of such specific terms may be changed to other forms without departing from the technical spirit of the present disclosure.

[0066] The following technology can be applied to various wireless access systems such as CDMA (code division multiple access), FDMA (frequency division multiple access), TDMA (time division multiple access), OFDMA (orthogonal frequency division multiple access), and SC-FDMA (single carrier frequency division multiple access).

[0067]

[0068] For clarity, the following description is based on a 3GPP wireless communication system (e.g., LTE, NR, etc.), but the technical spirit of the present disclosure is not limited thereto. LTE may refer to technology after 3GPP TS 36.xxx Release 8. Specifically, LTE technology after 3GPP TS 36.xxx Release 10 may be referred to as LTE-A, and LTE technology after 3GPP TS 36.xxx Release 13 may be referred to as LTE-A pro. 3GPP NR may refer to technology after TS 38.xxx Release 15. 3GPP 6G may refer to technology after TS Release 17 and / or Release 18. "xxx" refers to a standard document detail number. LTE / NR / 6G may be collectively referred to as a 3GPP system.

[0069] For background information, terms, abbreviations, etc. used in this disclosure, reference may be made to standard documents published prior to this disclosure. For example, reference may be made to standard documents 36.xxx and 38.xxx.

[0070]

[0071] Wireless communication system applicable to the present disclosure

[0072] Although not limited thereto, the various descriptions, functions, procedures, proposals, methods and / or operational flowcharts of the present disclosure disclosed in this document may be applied to various fields requiring wireless communication / connectivity (e.g., 5G) between devices.

[0073] Hereinafter, more specific examples will be provided with reference to the drawings. In the drawings / descriptions below, the same drawing reference numerals may represent identical or corresponding hardware blocks, software blocks, or functional blocks, unless otherwise described.

[0074] Figure 1 illustrates an example of a wireless communication system applied to the present disclosure.

[0075] Referring to FIG. 1, a wireless communication system (100) applied to the present disclosure includes a wireless device, a base station, and a network. Here, the wireless device refers to a device that performs communication using a wireless access technology (e.g., LTE, LTE-A, LTE-A pro, NR, 5G, 5G-A, 6G) and may be referred to as a communication / wireless / 5G device. Although not limited thereto, the wireless device may include a robot (100a), a vehicle (100b-1, 100b-2), an XR (extended reality) device (100c), a hand-held device (100d), a home appliance (100e), an IoT (Internet of Things) device (100f), and an AI (artificial intelligence) device / server (100g). For example, the vehicle may include a vehicle equipped with a wireless communication function, an autonomous vehicle, a vehicle capable of performing vehicle-to-vehicle communication, etc. Here, the vehicles (100b-1, 100b-2) may include unmanned aerial vehicles (UAVs) (e.g., drones). The XR devices (100c) include augmented reality (AR) / virtual reality (VR) / mixed reality (MR) devices, and may be implemented in the form of head-mounted devices (HMDs), head-up displays (HUDs) installed in vehicles, televisions, smartphones, computers, wearable devices, home appliances, digital signage, vehicles, robots, etc. The portable devices (100d) may include smartphones, smart pads, wearable devices (e.g., smartwatches, smart glasses), computers (e.g., laptops, etc.), etc. The home appliances (100e) may include TVs, refrigerators, washing machines, etc. The IoT devices (100f) may include sensors, smart meters, etc.For example, the base station (120) and the network (130) may also be implemented as wireless devices, and a specific wireless device (120a) may act as a base station / network node to other wireless devices.

[0076] Wireless devices (100a to 100f) can be connected to a network (130) via a base station (120). AI technology can be applied to the wireless devices (100a to 100f), and the wireless devices (100a to 100f) can be connected to an AI server (100g) via a network (130). The network (130) can be configured using a 3G network, a 4G (e.g., LTE) network, a 5G (e.g., NR), or a 6G network. The wireless devices (100a to 100f) can communicate with each other via the base station (120) / network (130), but can also communicate directly (e.g., sidelink communication) without going through the base station (120) / network (130). For example, vehicles (100b-1, 100b-2) can communicate directly (e.g., V2V (vehicle to vehicle) / V2X (vehicle to everything) communication). Additionally, an IoT device (100f) (e.g., a sensor) can communicate directly with another IoT device (e.g., a sensor) or another wireless device (100a to 100f).

[0077] Wireless communication / connection (150a, 150b, 150c) can be established between wireless devices (100a to 100f) / base stations (120), and base stations (120) / base stations (120). Here, the wireless communication / connection can be established through various wireless access technologies such as uplink / downlink communication (150a), sidelink communication (150b) (or D2D communication), and base station-to-base station communication (150c) (e.g., relay, IAB (integrated access backhaul)). Through the wireless communication / connection (150a, 150b, 150c), the wireless device and base station / wireless device, and base stations and base stations can transmit / receive wireless signals to / from each other. For example, the wireless communication / connection (150a, 150b, 150c) can transmit / receive signals through various physical channels. To this end, based on various proposals of the present disclosure, at least some of various configuration information setting processes for transmitting / receiving wireless signals, various signal processing processes (e.g., channel encoding / decoding, modulation / demodulation, resource mapping / demapping, etc.), resource allocation processes, etc. may be performed.

[0078]

[0079] Devices applicable to the present disclosure

[0080] FIG. 2 illustrates an example of a wireless device applicable to the present disclosure.

[0081] Referring to FIG. 2, the wireless device (200) can transmit and receive wireless signals via various wireless access technologies (e.g., LTE, LTE-A, LTE-A pro, NR, 5G, 5G-A, 6G). The wireless device (200) includes at least one processor (202) and at least one memory (204), and may additionally include at least one transceiver (206) and / or at least one antenna (208).

[0082] The processor (202) controls the memory (204) and / or the transceiver (206), and may be configured to implement the descriptions, functions, procedures, proposals, methods, and / or operational flowcharts disclosed in this document. For example, the processor (202) may process information in the memory (204) to generate first information / signal, and then transmit a wireless signal including the first information / signal via the transceiver (206). In addition, the processor (202) may receive a wireless signal including second information / signal via the transceiver (206), and then store information obtained from signal processing of the second information / signal in the memory (204). The memory (204) may be connected to the processor (202) and may store various information related to the operation of the processor (202). For example, the memory (204) may store software code including instructions for performing some or all of the processes controlled by the processor (202), or for performing the descriptions, functions, procedures, proposals, methods, and / or operational flowcharts disclosed herein. Here, the processor (202) and the memory (204) may be part of a communication modem / circuit / chip designed to implement wireless communication technology. The transceiver (206) may be connected to the processor (202) and may transmit and / or receive wireless signals via at least one antenna (208). The transceiver (206) may include a transmitter and / or a receiver. The transceiver (206) may be used interchangeably with an RF (radio frequency) unit. In the present disclosure, a wireless device may also mean a communication modem / circuit / chip.

[0083] Hereinafter, the hardware elements of the wireless device (200) will be described in more detail. Although not limited thereto, at least one protocol layer may be implemented by at least one processor (202). For example, at least one processor (202) may implement at least one layer (e.g., a functional layer such as physical (PHY), media access control (MAC), radio link control (RLC), packet data convergence protocol (PDCP), radio resource control (RRC), and service data adaptation protocol (SDAP)). At least one processor (202) may generate at least one Protocol Data Unit (PDU) and / or at least one Service Data Unit (SDU) according to the descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document. At least one processor (202) may generate a message, control information, data, or information according to the descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document. At least one processor (202) can generate a signal (e.g., a baseband signal) including a PDU, an SDU, a message, control information, data or information according to the functions, procedures, proposals and / or methods disclosed in this document, and provide the signal to at least one transceiver (206). At least one processor (202) can receive a signal (e.g., a baseband signal) from at least one transceiver (206) and obtain the PDU, SDU, message, control information, data or information according to the descriptions, functions, procedures, proposals, methods and / or operational flowcharts disclosed in this document.

[0084] At least one processor (202) may be referred to as a controller, a microcontroller, a microprocessor, or a microcomputer. The at least one processor (202) may be implemented by hardware, firmware, software, or a combination thereof. For example, at least one application specific integrated circuit (ASIC), at least one digital signal processor (DSP), at least one digital signal processing device (DSPD), at least one programmable logic device (PLD), or at least one field programmable gate array (FPGA) may be included in the at least one processor (202). The descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document may be implemented using firmware or software, and the firmware or software may be implemented to include modules, procedures, functions, etc. The descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document may be included in the at least one processor (202), or may be stored in at least one memory (204) and executed by the at least one processor (202). The descriptions, functions, procedures, suggestions, methods and / or flowcharts disclosed in this document may be implemented using firmware or software in the form of code, instructions and / or sets of instructions.

[0085] At least one memory (204) can be connected to at least one processor (202) and can store various forms of data, signals, messages, information, programs, codes, instructions and / or commands. The at least one memory (204) can be configured as a read only memory (ROM), a random access memory (RAM), an erasable programmable read only memory (EPROM), a flash memory, a hard drive, a register, a cache memory, a computer readable storage medium and / or a combination thereof. The at least one memory (204) can be located internally and / or externally to the at least one processor (202). In addition, the at least one memory (204) can be connected to the at least one processor (202) via various technologies such as a wired or wireless connection.

[0086] At least one transceiver (206) can transmit user data, control information, wireless signals / channels, etc., mentioned in the methods and / or flowcharts of this document to at least one other device. At least one transceiver (206) can receive user data, control information, wireless signals / channels, etc. mentioned in the descriptions, functions, procedures, proposals, methods and / or flowcharts disclosed in this document from at least one other device. For example, at least one transceiver (206) can be connected to at least one processor (202) and can transmit and receive wireless signals. For example, at least one processor (202) can control at least one transceiver (206) to transmit user data, control information, or wireless signals to at least one other device. Furthermore, at least one processor (202) can control at least one transceiver (206) to receive user data, control information, or wireless signals from at least one other device. In addition, at least one transceiver (206) may be connected to at least one antenna (208), and at least one transceiver (206) may be configured to transmit and receive user data, control information, wireless signals / channels, etc. mentioned in the descriptions, functions, procedures, proposals, methods and / or operation flowcharts disclosed in this document through at least one antenna (208). In this document, at least one antenna may be a plurality of physical antennas or a plurality of logical antennas (e.g., antenna ports). At least one transceiver (206) may convert the received wireless signals / channels, etc. from RF band signals to baseband signals in order to process the received user data, control information, wireless signals / channels, etc. using at least one processor (202). At least one transceiver (206) may convert the processed user data, control information, wireless signals / channels, etc. from baseband signals to RF band signals using at least one processor (202).For this purpose, at least one transceiver (206) may include an (analog) oscillator and / or filter.

[0087] The components of the wireless device described with reference to FIG. 2 may be referred to by different terms in terms of functionality. For example, the processor (202) may be referred to as a control unit, the transceiver (206) as a communication unit, and the memory (204) as a storage unit. In some cases, the communication unit may be used to mean at least a portion of the processor (202) and the transceiver (206).

[0088] The structure of the wireless device described with reference to FIG. 2 can be understood as the structure of at least a portion of various devices. For example, the structure of the wireless device illustrated in FIG. 2 can be at least a portion of various devices described with reference to FIG. 1 (e.g., a robot (100a), a vehicle (100b-1, 100b-2), an XR device (100c), a portable device (100d), a home appliance (100e), an IoT device (100f), an AI device / server (100g)). Furthermore, according to various embodiments, in addition to the components illustrated in FIG. 2, the device may further include other components.

[0089] For example, the device may be a portable device such as a smartphone, a smart pad, a wearable device (e.g., a smart watch, smart glasses), or a portable computer (e.g., a laptop, etc.). In this case, the device may further include at least one of a power supply unit that supplies power and includes a wired / wireless charging circuit, a battery, etc., an interface unit that includes at least one port for connection with another device (e.g., an audio input / output port, a video input / output port), and an input / output unit for inputting and outputting image information / signals, audio information / signals, data, and / or information input from a user.

[0090] For example, the device may be a mobile device such as a mobile robot, a vehicle, a train, an aerial vehicle (AV), a ship, etc. In this case, the device may further include at least one of a driving unit including at least one of an engine, a motor, a power train, wheels, brakes, and a steering unit of the device, a power supply unit including a wired / wireless charging circuit, a battery, etc. that supplies power, a sensor unit that senses status information, environmental information, and user information of the device or its surroundings, an autonomous driving unit that performs functions such as path maintenance, speed control, and destination setting, and a position measurement unit that obtains location information of the mobile device through a global positioning system (GPS) and various sensors.

[0091] For example, the device may be an XR device such as an HMD, a head-up display (HUD) installed in a vehicle, a television, a smartphone, a computer, a wearable device, a home appliance, a digital signage, a vehicle, a robot, etc. In this case, the device may further include at least one of a power supply unit that supplies power and includes a wired / wireless charging circuit, a battery, etc., an input / output unit that obtains control information, data, etc. from the outside and outputs the generated XR object, and a sensor unit that senses status information, environmental information, and user information of the device or the surroundings of the device.

[0092] For example, the device may be a robot that can be classified into industrial, medical, household, military, etc. types depending on the purpose or field of use. In this case, the device may further include at least one of a sensor unit that senses status information, environmental information, and user information of the device or its surroundings, and a driving unit that performs various physical actions, such as moving the robot joints.

[0093] For example, the device may be an AI device such as a TV, a projector, a smartphone, a PC, a laptop, a digital broadcasting terminal, a tablet PC, a wearable device, a set-top box (STB), a radio, a washing machine, a refrigerator, digital signage, a robot, a vehicle, etc. In this case, the device may further include at least one of an input unit that acquires various types of data from the outside, an output unit that generates output related to sight, hearing, or touch, a sensor unit that senses status information, environmental information, and user information of the device or its surroundings, and a training unit that trains a model composed of an artificial neural network using learning data.

[0094] The structure of the wireless device illustrated in FIG. 2 may be understood as a part of a RAN node (e.g., a base station, DU, RU, RRH, etc.). That is, the device illustrated in FIG. 2 may be a RAN node. In this case, the device may further include a wired transceiver for front haul and / or back haul communications. However, if the front haul and / or back haul communications are based on wireless communications, at least one transceiver (206) illustrated in FIG. 2 may be used for front haul and / or back haul communications, and a wired transceiver may not be included.

[0095]

[0096] FIG. 3 illustrates a method for processing a transmission signal applicable to the present disclosure. For example, the transmission signal may be processed by a signal processing circuit. At this time, the signal processing circuit (300) may include scramblers (310), modulators (320), a layer mapper (330), a precoder (340), resource mappers (350), and signal generators (360). At this time, for example, the operation / function of FIG. 3 may be performed in the processor (202) and / or the transceiver (206) of FIG. 2. Furthermore, for example, the hardware elements of FIG. 3 may be implemented in the processor (202) and / or the transceiver (206) of FIG. 2. For example, blocks 310 to 360 may be implemented in the processor (202) of FIG. 2. Additionally, blocks 310 to 350 may be implemented in the processor (202) of FIG. 2, and block 360 may be implemented in the transceiver (206) of FIG. 2, and are not limited to the above-described embodiment.

[0097] The codeword can be converted into a wireless signal through the signal processing circuit (300) of FIG. 3. Here, the codeword is an encoded bit sequence of an information block. The information block may include a transport block (e.g., a UL-SCH transport block, a DL-SCH transport block). Here, the information block may include data related to AI (e.g., training data, AI model data, input data, output data, etc.), and the codeword may be an encoded bit sequence corresponding to the data related to AI. The wireless signal may be transmitted through various physical channels (e.g., a PUSCH, a PDSCH). Specifically, the codeword may be converted into a bit sequence scrambled by scramblers (310). The scramble sequence used for scrambling is generated based on an initialization value, and the initialization value may include ID information of the wireless device, etc. The scrambled bit sequence may be modulated into a modulation symbol sequence by modulators (320). Modulation schemes may include pi / 2-BPSK (pi / 2-binary phase shift keying), m-PSK (m-phase shift keying), m-QAM (m-quadrature amplitude modulation), etc.

[0098] A complex modulation symbol sequence can be mapped to at least one transport layer by a layer mapper (330). Here, a transport layer is a logical resource unit for mapping a signal or data transmitted through spatial resources to antenna ports, and one transport layer can correspond to one stream or one antenna port. Each of the complex modulation symbols included in the complex modulation symbol sequence is mapped to at least one transport layer, thereby determining which antenna port it will be transmitted through. The modulation symbols of each transport layer can be mapped to the corresponding antenna port(s) by a precoder (340). The output z of the precoder (340) can be obtained by multiplying the output y of the layer mapper (330) by a precoding matrix W of NХM. Here, N is the number of antenna ports, and M is the number of transport layers. Here, the precoder (340) may perform precoding after performing transform precoding (e.g., discrete Fourier transform (DFT) transform) on complex modulation symbols. Additionally, the precoder (340) may perform precoding without performing transform precoding.

[0099] Resource mappers (350) can map modulation symbols of each antenna port to time-frequency resources. The time-frequency resources may include a plurality of symbols (e.g., CP-OFDMA symbols, DFT-s-OFDMA symbols) in the time domain and a plurality of subcarriers in the frequency domain. Signal generators (360) generate wireless signals from the mapped modulation symbols, and the generated wireless signals can be transmitted to other devices through each antenna. To this end, each of the signal generators (360) may include an inverse fast Fourier transform (IFFT) module, a cyclic prefix (CP) inserter, a digital-to-analog converter (DAC), a frequency uplink converter, etc.

[0100] The signal processing process for a received signal in a wireless device may be configured in reverse order of the signal processing process (310 to 360) of FIG. 3. For example, a wireless device (e.g., 200 of FIG. 2) may receive a wireless signal from the outside through an antenna port / transceiver. The received wireless signal may be converted into a baseband signal through a signal restorer. For this purpose, the signal restorer may include a frequency downlink converter, an analog-to-digital converter (ADC), a CP remover, and a fast Fourier transform (FFT) module. Thereafter, the baseband signal may be restored to a codeword through a resource demapper process, a postcoding process, a demodulation process, and a descrambling process. The codeword may be restored to the original information block through decoding. Therefore, a signal processing circuit (not shown) for a received signal may include a signal restorer, a resource demapper, a postcoder, a demodulator, a descrambler, and a decoder.

[0101] The signal processing circuit (300) described with reference to FIG. 3 is exemplified as including a plurality of scramblers (310), modulators (320), a plurality of resource mappers (350), and a plurality of signal generators (360). However, at least one of the scramblers, modulators, resource mappers, and signal generators may be implemented as a single integrated structure. That is, the number of at least one of the scramblers, modulators, resource mappers, and signal generators may be smaller than the number of layers. Furthermore, at least one of the components exemplified in FIG. 3 may be omitted.

[0102]

[0103] Figure 4 illustrates a communication procedure between a terminal and a base station applicable to the present disclosure. Figure 4 illustrates operations of a terminal (410) and a base station (420) transmitting and / or receiving data and operations performed prior thereto.

[0104] Referring to FIG. 4, in step 401, the terminal (410) and the base station (420) perform synchronization. For example, the terminal (410) performs an initial cell search operation. Specifically, the terminal (410) can detect at least one synchronization signal transmitted from the base station (420) according to a predefined rule. Here, the synchronization signal can include multiple synchronization signals classified according to structure or purpose (e.g., primary synchronization signal, secondary synchronization signal). Through this, the terminal (410) can check the boundary of the frame, subframe, slot, and / or symbol of the base station (420) and obtain information about the base station (420) (e.g., cell identifier).

[0105] In step 403, the terminal (410) obtains system information transmitted from the base station (420). The system information is information related to the properties, characteristics, and / or capabilities of the base station (420) required to access the base station (420) and use the service, and may be classified by content (e.g., whether it is essential for access), transmission structure (e.g., channel used, whether provided on-demand), etc., and may be classified into, for example, a master information block (MIB) and a system information block (SIB). If necessary, the terminal (410) may transmit a signal requesting system information before receiving the system information. The system information may include information related to an AI function. For example, the system information may include at least one of information related to an AI model, information related to training, and information related to inference / prediction, as information required for operations performed based on AI. However, the request and provision of the system information may be performed after a random access procedure described below.

[0106] In step 405, the terminal (410) and the base station (420) perform a random access procedure. The terminal (410) may transmit and / or receive at least one message (e.g., a random access preamble, a random access response (RAR) message, etc.) for the random access procedure based on information related to the random access channel of the base station (420) obtained through system information (e.g., channel position, channel structure, supported preamble structure, etc.). For example, the terminal (410) may transmit a preamble (e.g., MSG1) through the random access channel, receive an RAR message (e.g., MSG2), transmit a message (e.g., MSG3) including information related to the terminal (410) (e.g., identification information) to the base station (420) using scheduling information included in the RAR message, and receive a message (e.g., MSG4) for contention resolution and / or connection establishment. As another example, MSG1 and MSG3 may be sent and received as one message, or MSG2 and MSG4 may be sent and received as one message.

[0107] In step 407, the terminal (410) and the base station (420) perform signaling of control information. Here, the control information may be defined in various layers, such as a layer that controls a connection (e.g., a radio resource control (RRC) layer), a layer that handles mapping between logical channels and transport channels (e.g., a media access control (MAC) layer), and a layer that handles physical channels (e.g., a physical (PHY) layer). For example, the terminal (410) and the base station (420) may perform at least one of signaling for establishing a connection, signaling for determining settings related to communication, and signaling for indicating allocated resources. In addition, the signaling of the control information may be performed to convey information related to an AI function. For example, the information related to an AI function is information necessary for an operation performed based on AI, and may include at least one of information related to an AI model, information related to training, and information related to inference / prediction. More specifically, information related to the AI ​​function signaled in step 407 may be combined and / or linked with information related to the AI ​​function signaled in step 403, and the two may be defined in a hierarchical, mutually complementary, or substitutive structure.

[0108] In step 409, the terminal (410) and the base station (420) transmit and / or receive data. In other words, the terminal (410) and the base station (420) can process, transmit, and / or receive data based on the signaling of the control information. For example, when transmitting data, the terminal (410) or the base station (420) can perform at least one of channel encoding, rate matching, scrambling, constellation mapping, layer mapping, waveform modulation, antenna mapping, and resource mapping on the information bits. Conversely, when receiving data, the terminal (410) or the base station (420) can perform at least one of signal extraction from resources, waveform demodulation for each antenna, signal arrangement considering layer mapping, constellation demapping, descrambling, and channel decoding. Here, the transmitted data is data related to AI, and may include, for example, data for AI-based operations or data generated by AI-based operations.

[0109] Steps 401 to 409 illustrated with reference to FIG. 4 do not necessarily have to be performed in the order illustrated in FIG. 4, and the order of at least some of the steps may vary. Furthermore, at least some of steps 401 to 409 may be combined into a single step or omitted. That is, the steps illustrated in FIG. 4 may be performed in various modified forms.

[0110]

[0111] 6G wireless communication systems and core implementation technologies of 6G systems

[0112] The 5G system defines various operating bands within FR1 (frequency range 1), which covers 410 MHz to 7125 MHz, and FR2 (frequency range 2), which covers 24,250 MHz to 71,000 MHz. Various frequencies are being discussed as operating bands for the subsequent 6G system, and the use of higher frequencies than 5G systems is also being considered for wider bandwidth and higher transmission speeds. One such band is the THz (terahertz) frequency band, which covers approximately 100 GHz to 10 THz. The THz frequency band is a band that has both the transparency of radio waves and the straightness of light waves, and communications using the THz frequency band are expected to play a transitional role from existing radio-centered communications to lightwave-based communications.

[0113] 6G systems utilizing the THz frequency band are aimed at *?*very high data rates per device, *?*a very large number of connected devices, *?*global connectivity, *?*very low latency, *?*lower energy consumption of battery-free IoT devices, *?*ultra-reliable connectivity, and *?*connected intelligence with machine learning capabilities. The vision of the 6G system can be divided into four aspects: "intelligent connectivity," "deep connectivity," "holographic connectivity," and "ubiquitous connectivity," and the 6G system can be designed to satisfy the requirements as shown in [Table 1] below.

[0114] Per device peak data rate1 TbpsE2E latency1 msMaximum spectral efficiency100 bps / HzMobility supportup to 1000 km / hrSatellite integrationFullyAIFullyAutonomous vehicleFullyXRFullyHaptic CommunicationFully

[0115] At this time, the 6G system may have key factors such as enhanced mobile broadband (eMBB), ultra-reliable low latency communications (URLLC), massive machine type communications (mMTC), AI integrated communication, tactile internet, high throughput, high network capacity, high energy efficiency, low backhaul and access network congestion, and enhanced data security. FIG. 5 illustrates an example of a communication structure that can be provided in a 6G system applicable to the present disclosure. Referring to FIG. 5, the 6G system is expected to have simultaneous wireless communication connectivity that is 50 times higher than that of a 5G wireless communication system. URLLC, a key feature of 5G, is expected to become an even more crucial technology in 6G communications, offering end-to-end latency of less than 1 ms. Furthermore, 6G systems will boast significantly higher volumetric spectral efficiency than the commonly used area spectral efficiency. 6G systems can offer extremely long battery life and advanced battery technologies for energy harvesting, eliminating the need for separate charging for mobile devices in 6G systems. New network characteristics in 6G may include:

[0116] - Satellite integrated network: 6G is expected to integrate with satellites to provide a global mobile network. The integration of terrestrial, satellite, and airborne networks into a single wireless communications system is crucial for 6G.

[0117] Connected Intelligence: Unlike previous generations of wireless communication systems, 6G is revolutionary, upgrading the wireless evolution from "connected objects" to "connected intelligence." AI can be applied at every stage of the communication process (or at every signal processing step, as described below).

[0118] - Seamless integration of wireless information and energy transfer: 6G wireless networks will transfer power to charge the batteries of devices such as smartphones and sensors. Therefore, wireless information and energy transfer (WIET) will be integrated.

[0119] - Ubiquitous super 3D connectivity: Access to networks and core network functions of drones and very low Earth orbit satellites will create super 3D connectivity in 6G ubiquitous.

[0120] Some general requirements for the new network characteristics of 6G, such as the above, may be as follows:

[0121] - Small cell networks: The concept of small cell networks was introduced to improve received signal quality in cellular systems by increasing throughput, energy efficiency, and spectral efficiency. Consequently, small cell networks are essential for 5G and beyond 5G (5GB) wireless communication systems. Accordingly, 6G wireless communication systems also adopt the characteristics of small cell networks.

[0122] Ultra-dense heterogeneous networks: Ultra-dense heterogeneous networks will be another key feature of 6G wireless communication systems. Multi-tier networks comprised of heterogeneous networks improve overall QoS and reduce costs.

[0123] High-capacity backhaul: Backhaul connections are characterized by high-capacity backhaul networks to support high-volume traffic. High-speed fiber optics and free-space optics (FSO) systems may be potential solutions to this problem.

[0124] - Radar technology integrated with mobile technology: High-precision localization (or location-based services) through communications is a key feature of 6G wireless communication systems. Therefore, radar systems will be integrated with 6G networks.

[0125] - Softwarization and virtualization: Softwarization and virtualization are two critical features that form the foundation of the design process for 5GB networks to ensure flexibility, reconfigurability, and programmability. Furthermore, billions of devices can be shared on a shared physical infrastructure.

[0126] To satisfy the above-mentioned characteristics, the core implementation technologies of the 6G system may include artificial intelligence (AI), THz (terahertz) communication, optical wireless technology, FSO backhaul network, massive MIMO technology, blockchain, 3D networking, quantum communication, unmanned aerial vehicles, cell-free communication, wireless information and energy transfer (WIET), integration of sensing and communication, integration of access backhaul networks, holographic beamforming, big data analysis, and large intelligent surface (LIS).

[0127] For example, THz communication can be utilized in 6G systems. THz communication is a communication that utilizes a spectrum in a frequency band between 0.3 THz and 3 THz with a corresponding wavelength in the range of 0.1 mm to 1 mm, as shown in FIG. 6. Referring to FIG. 6, the frequency band of THz waves is located in the middle region between the infrared band and the millimeter wave band, and therefore, THz waves can be understood as radio waves with the shortest wavelength and light waves with the longest wavelength. Therefore, THz waves share some of the characteristics of infrared and microwave waves, and specifically, they can simultaneously have the transparency of electromagnetic waves and the straightness of light waves.

[0128]

[0129] Figure 7 illustrates a transmitter structure applicable to the present disclosure.

[0130] Referring to Figure 7, in order to modulate data into an optical signal, an optical source of a laser can be passed through an optical wave guide to change the phase of the signal, etc. At this time, data is loaded by changing the electrical characteristics through a microwave contact, etc. Therefore, the optical modulator output is formed as a modulated waveform.

[0131] Data may be provided from a data signal generator. Here, the data may include various user data, configuration information, control information, etc. transmitted through a channel. Furthermore, the data may include data related to AI-based operations, such as information for configuring an AI model, input / output data for tasks of the AI ​​model, etc. To this end, components related to AI functions (e.g., an AI processing unit) may be included in the data signal generator or may be linked to the data signal generator.

[0132] An optical / electronic converter (O / E converter) can generate THz pulses by optical rectification using a nonlinear crystal, photoelectric conversion using a photoconductive antenna, or emission from a bunch of relativistic electrons. The THz pulse generated in the above manner can have a length in the range of femtoseconds to picoseconds. The optical / electronic converter (O / E converter) performs down conversion by utilizing the nonlinearity of the device.

[0133] Considering the THz spectrum usage, it is likely that multiple contiguous GHz bands will be used for THz systems, either fixed or for mobile services. For an outdoor scenario, the available bandwidth can be categorized based on an oxygen attenuation of 10^2 dB / km in the spectrum up to 1 THz. Accordingly, a framework in which the available bandwidth is divided into multiple band chunks can be considered. As an example of this framework, if the THz pulse length for a single carrier is set to 50 ps, ​​the bandwidth (BW) becomes approximately 20 GHz.

[0134] Effective down-conversion from the infrared band to the THz band depends on how to utilize the nonlinearity of the optical / electrical converter (O / E converter). In other words, to down-convert to the desired THz band, it is necessary to design an O / E converter with the most ideal non-linearity for transferring to the THz band. If an O / E converter that is not suitable for the target frequency band is used, errors in the amplitude and phase of the pulse are likely to occur.

[0135] A THz transmission and reception system can be implemented using a single optical-to-electrical converter in a single-carrier system. Depending on the channel environment, optical-to-electrical converters may be required as many as the number of carriers in a multi-carrier system. This phenomenon will be particularly noticeable in a multi-carrier system that utilizes multiple broadbands according to the aforementioned spectrum usage plan. In this regard, a frame structure for the multi-carrier system may be considered. A signal down-frequency converted based on an optical-to-electrical converter may be transmitted in a specific resource region (e.g., a specific frame). The frequency region of the specific resource region may include multiple chunks. Each chunk may be composed of at least one component carrier (CC).

[0136]

[0137] 6G systems may introduce AI technology. Efficient resource management and optimization are necessary to maintain connectivity between various services and devices. AI technology may include technologies that perform data analysis, pattern recognition, and predictive modeling using AI / ML (artificial intelligence / machine learning) models. Here, an AI / ML model can be understood as a set of parameter values ​​and / or weight values ​​related to mathematical formulas or algorithms generated through learning to discover patterns in input data or make predictions. To create such an AI / ML model, an AI / ML model learning process is required, which builds an AI / ML model by learning the relationship between inputs and outputs in a data-driven manner. Various learning algorithms, such as supervised learning, unsupervised learning, and reinforcement learning, can be utilized. To generate output, a user can input specific data into a trained AI / ML model, and the process of obtaining output data by inputting input data into an AI / ML model can be referred to as AI / ML "inference" or "prediction."

[0138] Network control parameters can be obtained as output through AI / ML inference using trained AI / ML models. Users can utilize the output parameter values ​​to improve network efficiency. For example, AI technology can be utilized in various fields, such as wireless network resource allocation, traffic management, fault prediction, and quality of service (QoS) management. In particular, machine learning can efficiently allocate resources even in dynamically changing network environments based on real-time data. Therefore, AI technology can be utilized to provide hyper-connectivity and ultra-low latency.

[0139] At this time, AI / ML inference can be performed based on a combination of various devices. For example, the UE and the network can jointly perform AI / ML inference, and such an AI / ML model can be referred to as a two-sided AI / ML model or a two-sided model. In this case, the UE can perform the first part of the inference first, and the base station can perform the remaining inference, or vice versa. Alternatively, inference can be performed entirely on the UE, and such an AI / ML model can be referred to as a UE-side AI / ML model or a UE-side model.

[0140] Additionally, life cycle management (LCM) can be performed for AI / ML models. Life cycle management can include model training, model deployment, model inference, model monitoring, and model updates. This may require support for data collection, model training, functional / model identification, model delivery / transfer, model inference operations, functional / model selection / activation / deactivation / fallback, functional / model monitoring, model updates, and UE capabilities.

[0141] For example, AI / ML technology can be operated based on a functional framework such as FIG. 8. FIG. 8 illustrates an example of a functional framework for application of AI / ML technology applicable to the present disclosure. First, a data collection function (810) performs data preparation on input data collected from objects (e.g., UE, RAN node, network node, etc.) to generate training data (801), monitoring data (803), and / or inference data (805) including processed input data. A model training function (820), which receives training data (801) from the data collection function (810), performs training on an AI / ML model using the training data (801) and provides a trained / updated model (813) to a model repository (840). The model repository (840) can store and retain the received trained / updated model (813).

[0142] A management function (830) may be used to control AI / ML model training. The management function (830) may control the operation of the AI / ML model or AI / ML functions, or supervise their performance. To this end, the management function (830) may receive monitoring data (830) from the data collection function (810) and inference output (809) from the inference function (840). The management function (830) manages the data received from the data collection function (810) and the inference function (840) so that the inference task can be performed efficiently. That is, the management function (830) may transmit performance feedback or a retraining request (807) to the model training function (820) to improve the inference task. Here, the performance feedback may be used to indicate a learning goal or as a reward for reinforcement learning. Additionally, the management function (830) can transmit management instructions (811) that instruct the inference function (840) to select AI / ML models or AL / ML-based functions to be used, activate / deactivate them, or switch to non-AI / ML operation.

[0143] The inference function (840) generates an inference output (809) by performing inference and / or prediction using the inference data (805) received by the data collection function (810). Here, the inference output (809) refers to the inference output of the AI / ML model used by the inference function (840), and the details of the inference output may vary depending on the use case. The AI / ML model used by the inference function (840) can be controlled by the management function (830). That is, the management function (830) can transmit a model transfer / forward request signal (815) to request a necessary AI / ML model to the model repository function (850), and the model repository function (850) can transmit the corresponding AI / ML model to the inference function (840) via a model transfer / forward signal (817). Therefore, the inference function (840) can perform inference using the AI / ML model (817) according to the received management instruction (811).

[0144] Additionally, the management function (830) may trigger or perform a designated task / action based on the inference output (809). Accordingly, the management function (830) may trigger a task / action for another entity (e.g., at least one UE, at least one RAN node, at least one network node, etc.) or for itself. Any one of the functions exemplified in FIG. 8 described above may be performed by two or more entities, including the RAN, the network node, the network operator's OAM, or the UE, in collaboration. This may be referred to as a split AI operation.

[0145] Not all of the functions (810 to 850) illustrated in FIG. 8 need to be used to utilize the AI / ML model, and the method of combining them is not limited to a specific method. Accordingly, the functions (810 to 850) may be operated in an integrated manner, or some functions may be omitted. Furthermore, the functions (810 to 850) illustrated in FIG. 8 are not necessarily limited to being implemented as separate devices or apparatuses. For example, some or all of the functions (810 to 850) may be included in the processor (202) of FIG. 2. Furthermore, the model storage function (850) may be included in the memory (204) of FIG. 2.

[0146]

[0147] FIG. 9 illustrates an example of a procedure for utilizing an AI model applicable to the present disclosure. FIG. 9 illustrates a case where a model training function (820) is included in a network node and a model inference function (840) is included in a RAN node. Referring to FIG. 9, in step 1, RAN node 1 and RAN node 2 transmit input data (e.g., training data) for training an AI model to the network node. Here, RAN node 1 and RAN node 2 may transmit data collected from the UE (e.g., measurements of the UE related to RSRP, RSRQ, SINR of the serving cell and neighboring cells, the UE's position, speed, etc.) together to the network node. In step 2, the network node trains the AI ​​model using the received training data. In step 3, the network node distributes / updates the AI ​​model to RAN node 1 and / or RAN node 2. RAN node 1 and / or RAN node 2 may also continue model training based on the received AI model. In this procedure, it is assumed that the AI ​​model is deployed / updated only to RAN node 1. In step 4, RAN node 1 receives input data (e.g., inference data) for AI model inference from UE and RAN node 2. In step 5, RAN node 1 performs AI model-based inference using the received inference data to generate output data (e.g., prediction or decision). In step 6, if applicable, RAN node 1 may transmit model performance feedback to network nodes. In step 7, RAN node 1, RAN node 2, and UE (or 'RAN node 1 and UE', or 'RAN node 1 and RAN node 2') perform actions based on the output data. For example, in case of load balancing operation, the UE may move from RAN node 1 to RAN node 2. In step 8, RAN node 1 and RAN node 2 transmit feedback information to network nodes.

[0148] Network nodes can manage AI models based on feedback information regarding the AI ​​model's inference results. For example, the network node can perform additional training on the AI ​​model or generate additional information about the AI ​​model (e.g., performance information, accuracy information, etc.). If additional training is performed on the AI ​​model, the network node can distribute the updated AI model to RAN node 1.

[0149] As described with reference to Figure 9, model training can be performed by network nodes, and inference using the model can be performed by RAN node 1. In other words, the model training and inference functions can be distributed. Typically, model training requires significant computational resources because it utilizes large amounts of data and complex algorithms for optimization. In contrast, inference, which uses a trained model to derive conclusions about new data, requires relatively fewer computational resources compared to model training. Therefore, using the procedure of Figure 9, model training can be performed through network nodes when the computational resources of the UE or RAN node are insufficient. Furthermore, security can be ensured for the AI ​​model because the AI ​​model is not disclosed to the UE.

[0150] FIG. 9 illustrates a case where a model training function (820) is included in a network node and a model inference function (840) is included in a RAN node, but the present disclosure is not limited thereto. For example, if the computational resources of the RAN node are sufficient, both the model training function (820) and the model inference function (840) may be included in RAN node 1. In this case, RAN node 1 receives training data for training an AI model from the UE and RAN node 2. RAN node 1 trains the AI ​​model using the received training data. Thereafter, RAN node 1 receives inference data for AI model inference from the UE and RAN node 2. RAN node 1 performs inference based on the AI ​​model using the received inference data to generate output data. Based on the output data, the UE, RAN node 1, and RAN node 2 may perform operations related to communication (e.g., handover, cell change). Thereafter, the UE and RAN node 2 may transmit feedback regarding the operations to RAN node 1. Therefore, RAN node 1 can learn the AI ​​model and update its own AI model using feedback information regarding the AI ​​model's inference results. According to the aforementioned method, signaling with the network is not required for AI model learning and inference, thereby reducing network load and delays until the AI ​​model is trained or inference results are received. Furthermore, since the UE performs inference using the AI ​​model, its personal information is not transmitted to network nodes, etc., thereby enhancing the security of personal information.

[0151] As another example, the model training function (820) may be included in the RAN node, and the model inference function (840) may be included in the UE. The RAN node receives training data for training an AI model from the UE, and trains the AI ​​model using the received training data. The RAN node distributes the trained AI model to the UE. The UE may perform inference based on the received AI model to generate output data. At this time, the data for inference may be received from the RAN node, or the UE may use data acquired on its own. The UE and the RAN node may perform communication-related operations based on the output data generated by inference. Thereafter, the UE may transmit feedback regarding the operation to the RAN node. Therefore, the RAN node may train the AI ​​model and distribute the updated AI model to the UE through feedback information regarding the inference result of the AI ​​model. According to the above-described method, the model training function (820) may be included in the RAN node, and the model inference function (840) may be included in the UE to perform inference, thereby reducing the load on the RAN node. Additionally, if the UE uses data acquired on its own to make inferences, even if the UE loses connection with the RAN node after receiving the AI ​​model, the UE can continue to infer the distributed AI model and perform operations based on the inference results.

[0152] According to the framework and procedures described above, an AI model can be trained and utilized in a wireless communication system. The model training function (820) and the model inference function (840) can be combined in various ways and are not necessarily limited to the case of FIG. 9. In the framework and procedures described above, various types of data, such as input data, training data, and inference data, are introduced, and the specific content of the data described above may vary depending on the task for which the AI ​​model is utilized. For example, information used in various embodiments of the present disclosure described below may be included in the data described above.

[0153]

[0154] Figure 10 illustrates an AI technology-based communication procedure applicable to the present disclosure. The detailed procedures illustrated in Figure 10 can be combined with various embodiments of the present disclosure described below. For example, data generated according to various embodiments of the present disclosure can be used for operations (e.g., configuration, training, inference, and / or data transmission / reception) in at least one of the detailed procedures illustrated in Figure 10. As another example, the results of the inference illustrated in Figure 10 can be used to transmit and / or receive data according to various embodiments of the present disclosure.

[0155] Referring to FIG. 10, in step S1001, at least one of the UE (1010), the RAN node (1020), and the network node (1030) performs an initial access procedure. For example, in this step, at least one of an initial cell search operation, a system information acquisition operation, a random access operation, and a registration operation may be performed. In step S1003, at least one of the UE (1010), the RAN node (1020), and the network node (1030) performs a configuration procedure. Through the configuration procedure, parameters, resources, connections, and / or entities necessary for performing subsequent procedures in layers between the UE (1010) and the RAN node (1020) and / or in at least one layer between the UE (1010) and the network node (1030) may be determined and / or created. In this case, the configuration procedure may be performed based on information, status, and / or characteristics of an AI model used for subsequent training and inference.

[0156] In step S1005, at least one of the UE (1010), the RAN node (1020), and the network node (1030) performs a model training procedure. At least one of the UE (1010), the RAN node (1020), and the network node (1030) may collect training data and perform learning using the training data. For example, the model training procedure may be performed as described with reference to FIG. 15, FIG. 16, or FIG. 17. If an offline-trained model is used, this step may be omitted.

[0157] In step S1007, at least one of the UE (1010), the RAN node (1020), and the network node (1030) performs a task using the trained model. That is, the task may be performed based on the results of inference and / or prediction using the trained model. For example, the task may be a procedure belonging to a communication protocol, and may be a preparatory operation for subsequent data transmission and / or reception, or may be related to data transmission and / or reception, or may be related to data processing (e.g., encoding, decoding, etc.).

[0158] In step S1009, at least one of the UE (1010), the RAN node (1020), and the network node (1030) transmits and / or receives data. At this time, the result of the task performed in step 1007 may be used. In some cases, the task performed in step 1007 may include transmitting and / or receiving data, in which case this step may be omitted as it is part of step 1007.

[0159]

[0160] Specific embodiments of the present disclosure

[0161] The present disclosure relates to a technique for managing beams using an artificial intelligence / machine learning (AI / ML) model in a wireless communication system. In particular, the present disclosure proposes a technique for utilizing an AI / ML model for beam acquisition, beam tracking, etc. at the physical layer. Specifically, the present disclosure proposes various embodiments for adaptively controlling the format of reporting information or reference signals for beam management using an AI / ML model.

[0162]

[0163] The problem of AI / ML models inferring beam information between devices is mainly related to inputting channel information and providing the expected optimal beam information as an output. The input data of the AI / ML model may include information explicitly or implicitly including the channel conditions between the transmitting device and the receiving device. Explicit information may include channel estimation through a beam reference signal. Implicit information may include side information such as device location information that may reflect other channel information. The output of the AI / ML model may include predicted beam candidates. The beam information input may indicate a set of multiple beams that are spatially different and transmitted in a wireless channel, or a set of multiple beams that include multiple continuous temporal information that is observed or estimated temporally.

[0164] For example, in 3GPP standard specifications TS 38.843, an AI / ML model for beam prediction is modeled as in FIG. 11. FIG. 11 illustrates an example of an AI / ML model for beam prediction. Referring to FIG. 11, the input to the AI / ML model (1102) is a set B of beams based on beam measurements. The set B of beams can be classified into a case in which it includes spatial information of the beam and a case in which it includes temporal information. Each case is referred to as 'case 1' and 'case 2'. The output of the AI / ML model (1102) includes the beam probability or the predicted beam included in the set A of beams.

[0165] These AI / ML models may exist only on one side of the base station and the terminal, or may be deployed on both sides. Since the base station's transmission may be performed at multiple transmission / reception points (multi-TRPs), multiple AI / ML models may exist for each multi-TRP. Here, it is preferable that the input of the AI / ML model include channel information between the transmitter and the receiver. For example, when the AI / ML model is used for base station downlink beam management, the input of the AI / ML model may include beam information or channel information measured or estimated by the terminal.

[0166] Therefore, from a system perspective, the receiver must report channel information to the transmitter, or the receiver must transmit a reference signal to the transmitter. In the present disclosure, the operation of transmitting the reporting information or the reference signal may be referred to as 'beam management feedback', 'beam management signaling', 'beam management redundancy', or other terms having equivalent technical meanings thereto. Beam management feedback corresponds to the overhead or cost of the radio resources of the physical layer. Therefore, in the present disclosure, the operation of transmitting the reporting information or the reference signal may be referred to as 'beam management overhead', 'beam management cost', or other terms having equivalent technical meanings thereto. Accordingly, the present disclosure aims to solve the problem of maximizing the beam management performance for beam capturing and tracking of an AI / ML model, and minimizing the physical layer resource cost.

[0167]

[0168] FIG. 12A and FIG. 12B illustrate deployment examples of AI / ML models according to an embodiment of the present disclosure. FIG. 12A illustrates a case where an AI / ML model (1202) is included in a base station, and FIG. 12B illustrates a case where an AI / ML model (1202) is included in a terminal (1210). Referring to FIG. 12A, when an AI / ML model (1202) is included in a base station (1220), the base station (1220) performs beam management using CSI-RS reporting or SRS. Referring to FIG. 12B, when an AI / ML model (1202) is included in a terminal (1210), the terminal (1210) operates the AI / ML model using CSI-RS reception and SRS reporting from the base station (1220). In FIG. 12A and FIG. 12B, a T (φ t ) is the beam φ at the time t of the sender. t Codebook beam vector for a, T (ψ t ) is the beam's ψ at the receiver's time t. tCodebook beam vector for, H t is a channel estimation value at time t, c(t) is a channel information report for time t or a transmission cost for a reference signal (e.g., a function of the number of resource blocks (RBs) containing symbol information and the transmission power of the terminal (1210) f(nRB,TxPower)). Here, it is assumed that the terminal (1210) has a beam correspondence function. The beam correspondence function refers to the ability of the terminal (1210) to select a beam suitable for uplink transmission based on downlink beam measurement without relying on uplink beam sweeping. The beam correspondence characteristic may be referred to as channel reciprocity. The reinforcement learning module may exist at either the base station (1220) or the terminal (1210). The AI / ML model (1202) performs estimation of the downlink beam.

[0169] The terminal is assumed to operate on an infinite time horizon for a time t during system operation. This can be described in terms of reinforcement learning as follows: First, the optimization equation [Mathematical Equation 1] must be satisfied so that the average signal-to-noise ratio (SNR) can be maximized over an infinite time period.

[0170]

[0171] In [Equation 1], a T (φ t ) is the codebook beam vector at time t of the sender, φ t is the beam index of the sender, a T (ψ t ) is the codebook beam vector at the receiver at time t, ψ t is the receiver's beam index, H t is the channel estimate value at time t, and c(t) represents the feedback cost due to channel information reporting or reference signal at time t.

[0172] Additionally, during the operating time, the feedback cost c(t) must be minimized. Optimization for the feedback cost can be expressed as follows [Mathematical Equation 2].

[0173]

[0174] In [Equation 2], c(t) is the feedback cost due to channel information reporting or reference signal at time t, g t is the cost function, nRB is the number of RBs, and TxPower is the transmission power.

[0175]

[0176] For beam management based on AI / ML models according to various embodiments of the present disclosure, reinforcement learning can be utilized. The concept of reinforcement learning is as shown in FIG. 13. FIG. 13 illustrates the structure of reinforcement learning according to one embodiment of the present disclosure. Reinforcement learning is performed based on the interaction between an agent (1302) and an environment (1304). The agent (1302) is in a state S. t In the present disclosure, the agent (1302) is an object including an AI / ML model and may be included in a base station or a terminal. In addition, in the present disclosure, state S t may include a vector or matrix of additional information such as channel status or location information between a base station and a terminal. In addition, in the present disclosure, the environment (1304) includes an opposite base station or terminal, a channel, etc. that does not have an agent (1302). Action A t is information or signal fed back from the agent (1302) to the environment (1304), which may be related to the format of a channel information report or reference signal including a channel measurement value transmitted to a base station or a terminal, for example. Reward R t Silver Action A tis feedback on. For example, in this disclosure, the reward R t includes changes in channel quality (e.g., SNR, RSRP (reference signal received power)).

[0177] If the reinforcement learning model illustrated in Figure 13 is modeled as a Markov decision process, it becomes a problem of finding an optimal state-value function using the Bellman equation. This disclosure defines this problem as model-based reinforcement learning. The optimal policy can be determined using policy iteration or value iteration. Model-free reinforcement learning relies entirely on data from the environment without a model. If the agent's episodes are finite in time, Monte Carlo techniques can be used. S t , R, A, S t-1 Continuous data collection is performed, and state values ​​can be estimated based on the collected data. When episodes are infinite in time, the temporal-difference (TD) technique can be used. Using the TD technique, the state value of the current step is updated based on the difference between the state value of the previous step and the state value of the current step, and the policy can be verified through the state value or action state value. TD techniques are classified into SARSA (state-action-reward-state-action) based on an activation policy and Q-learning based on a deactivation policy.

[0178] When states are infinite, deep Q-learning (DQN) can be used to approximate state-value or action-value functions using a deep neural network (DNN). For example, it is possible to assign states to terminals, base stations, and channels as two-dimensional CNN inputs, unfolding them over a finite time period, and configure action settings as feedback, which is the action space. A value-based method for determining policies from state-value or action-value functions has been described. Another reinforcement learning approach, a policy-based approach, determines decisions directly from the environment.

[0179] Policy-based approaches are model-free techniques based on the policy gradient theorem. Maximizing a value function requires determining its gradient. The policy gradient theorem expresses the operation of determining the gradient as the product of the action-value function and the gradient of the policy function. This allows the gradient of the policy function to be directly determined from the environment, enabling learning. A representative algorithm is the Monte-Carlo policy gradient technique known as "REINFORCE."

[0180] It is also possible to simultaneously and iteratively utilize policy-based and value-based approaches. A representative example is the actor-critic (AC) algorithm, which defines advantage as the difference between the current action value function and the average value function. The AC algorithm uses the average of the slopes of the product of the advantage and the log-policy approximation function as the gradient of the objective function. The AC algorithm's advantage lies in its ability to minimize fluctuations in the value function value as data is collected, enabling stable learning.

[0181]

[0182] For beam management based on AI / ML models according to various embodiments of the present disclosure, multi-objective reinforcement learning can be utilized. An optimization problem for achieving two objectives, such as maximizing beam performance and minimizing feedback cost, can be expressed in the form of multi-objective reinforcement learning. Here, the target of optimization is a beam management feedback resource, specifically, a resource occupied by a reference signal transmitted for beam management or a report of channel information measured based on the reference signal. For example, the beam management feedback resource optimization problem is (S, A, P, K, r) in an infinite time dimension. 1 , … , r K , , ρ0, f) can be treated as solving a Markov decision process problem represented by a tuple. Here, the definition of each component of the tuple is as follows [Table 2].

[0183] Variable DescriptionSa finite set of state spaceAa finite set of actionsP:S×A→Sprobability transition distributionKthe set of K objectivesr k= :S×A→[0,1]reward generated by objective k discount factorρ0the distribution of initial statef:R K →Rthe scalarization function

[0184] The state space S may include channel status, beam identifiers, beam signal quality, location information, etc. P is the action value transition probability. For example, the value of K can be 2. A two-dimensional vector can be defined that aims to maximize beam performance and minimize feedback cost. f is a function that transforms two reward or objective vectors into a single scalar (e.g., f(r1,r2)=(1-α)r1+α2). In this setting, the policy π is defined as follows. * It can be defined as follows [Mathematical Formula 3].

[0185]

[0186]

[0187]

[0188]

[0189]

[0190] In [Equation 3], is the average cumulative reward for objective k and policy π, s t is the state at time t, a t means action at time t.

[0191]

[0192] For AI / ML model-based beam management according to various embodiments of the present disclosure, a Pareto Frontier search technique may be used.

[0193] Policy π * The Pareto optimality of is for all k. ≥ or for some k ≥ . Fig. 14 illustrates the Pareto optimal region of beam performance and feedback cost. Fig. 14 shows the Pareto optimal region of beam performance and feedback cost when K is 2. The maximum region of feedback efficiency is the case where there is no feedback, which corresponds to the right vertical dotted line (1402), and is an unreachable upper limit. The highest region of beam performance corresponds to the upper horizontal dotted line (1403) where the TX-RX beam alignment performance is maximum. The part indicated by the thick line (1406) is the boundary of the optimal region, which is a set of ideal points that the proposed technology aims for. In the upper boundary (1406), the point pursued by the proposed technology is to first satisfy the beam performance and reduce the feedback cost, even if there is a certain level of feedback cost. Therefore, the present disclosure proposes various embodiments that reflect the action space, utilization, and exploration for satisfying these parts.

[0194]

[0195] The present disclosure proposes a technique for optimizing feedback information for beam management using reinforcement learning to solve the aforementioned optimization problem. The optimization of the present disclosure may be referred to as "beam management feedback optimization (BMFO)." A system according to various embodiments includes a data collection module that collects information from the channel environment between a base station and a terminal, a reward module that converts the collected data into rewards, a state module that manages and stores channel states, and a network that determines the next action based on a state or action. The network may be implemented as a neural network and outputs information that directs actions. Here, the network may include a policy network or a value network. The policy network determines the next action based on the probability of performing each action based on a given state, and the value network determines the next action that maximizes the value of the state generated by the next action based on the given state and action. The state module may manage additional information, such as location information, along with the channel information. For example, vector state information may be generated by combining the terminal's location information with a continuous state space for channel H. Location information can take the form of continuous values ​​or index values ​​in a quantized space by region.

[0196] The action space according to various embodiments has the following characteristics.

[0197] - A report message including measurement results for a reference signal (e.g., CSI-RS, SRS) which is feedback information is defined as a set of multiple settings or multiple settings for a reference signal.

[0198] - Elements within a set are sorted in descending or ascending order based on the amount of feedback resource consumed. Alternatively, the set can be structured to select elements in descending or ascending order based on the amount of resource consumed.

[0199] - Elements of the set are indicated by indicators or indices, and are mutually signaled between the base station and the terminal through the indices.

[0200] The action set may be defined to include actions that have a 1:1 relationship with the available elements, or it may include two actions that increase or decrease an indicator indicating the index of one of the available elements (e.g., a first action indicating an increase and a second action indicating a decrease). Here, the increase and decrease are intended to reduce the variance of reinforcement learning.

[0201] Resources are defined as follows:

[0202] - Data transmission rate according to the reporting cycle of reporting messages for channel quality (e.g. RSRP, SINR)

[0203] - Data transmission rate based on the maximum number of reports included in a reporting message about channel quality (e.g. RSRP, SINR).

[0204] - Data transmission rate according to the quantization level of channel quality values ​​for reporting on channel quality (e.g. RSRP, SINR).

[0205] - Data transmission rate according to the maximum number of multi-TRPs involved in reporting on channel quality (e.g. RSRP, SINR)

[0206] - Transmission rate according to density on the RE (resource element) grid according to the downlink reference signal (e.g. CSI-RS) transmission format

[0207] - Transmission rate according to density on the RE grid according to the uplink reference signal (e.g. SRS) transmission format

[0208] Exploitation and exploration for feedback optimization are described below. In reinforcement learning, exploitation refers to the process of maximizing the use of currently known optimal options or knowledge. Exploitation allows for the efficient use of existing knowledge or resources. Exploration in reinforcement learning refers to the process of finding new options or exploring unknown areas. Through exploration, new information or knowledge can be obtained. In this regard, the present disclosure proposes the following embodiments.

[0209] - By action value function or policy function, actions with high feedback resource cost for beam management performance are allocated to utilization.

[0210] - Assign actions with small feedback resource cost for beam management performance to search by action value function or policy function.

[0211] The aforementioned allocation scheme allows reinforcement learning starting with high feedback resources to perform reinforcement learning in low feedback space without degrading beam performance.

[0212]

[0213] In one embodiment, in a deep Q-network, instead of randomly initializing the initial value action function, it is possible to initialize it with an action with a large feedback resource cost and explore actions with a small feedback resource cost using epsilon-greedy. In another embodiment, when performing reinforcement learning based on a policy neural network, a policy π(a) can be used so that feedback resources can be probabilistically selected for actions that require relatively large feedback resources in a state where exploitation-exploration is possible. t |s t) can be initialized. For the remaining probabilities, actions requiring relatively small feedback resources can be placed, and exploration can be performed. When reinforcement learning is performed based on a deterministic technique such as deep deterministic policy gradient (DDPG), initialization can be performed in response to actions requiring relatively large feedback resources, and actions requiring relatively small feedback resources can be placed for exploration noise.

[0214]

[0215] Various embodiments according to the present disclosure can be classified according to the location of the AI / ML model (e.g., base station side or terminal side) and the signal used (e.g., CSI-RS or SRS). Hereinafter, embodiments in which a base station-side model is used based on CSI-RS, embodiments in which a base station-side model is used based on SRS, embodiments in which a terminal-side model is used based on CSI-RS, and embodiments in which a terminal-side model is used based on SRS are described.

[0216]

[0217] [Example #1] CSI-RS-based method using a base station model

[0218] This embodiment can be applied when a base station possesses an AI / ML model. The base station inputs feedback information, including the terminal's CSI report and additional information (e.g., terminal location, terminal speed, etc.), into the model. Accordingly, rewards with the following characteristics can be granted.

[0219] - Quality values ​​(e.g., cri-RSRP, ssb-index-RSRP, cri-SINR, ssb-index-SINR) of a set of serving beams transmitted from one or more CSI-RS resources or SSB resources, or the difference between the values ​​of previous measurement values ​​of the corresponding parameter and the current measurement value.

[0220] - Quality value of CSI-RS estimated at the terminal (e.g. quality value at beam alignment) or difference value between previous measurement value and current measurement value of the corresponding parameter

[0221] - The CQI value of the CSI-RS estimated from the terminal or the difference value between the previous measurement value and the current measurement value of the corresponding parameter.

[0222] If a report is generated due to the aforementioned feedback, compensation for costs may be determined as follows:

[0223] - Size of the report message per unit of time (e.g. total number of bits) or the difference between the size of the previous message and the size of the current message per unit of time.

[0224] - Size and transmission power of RB (resource block) for transmitting reports per unit time

[0225] Since the compensation includes two elements, it can be expressed as a vector value. The vector value is converted to a scalar value by a compensation transformation function. For example, the compensation transformation function can be defined as a weighted sum of two compensation elements. As another example, the compensation transformation function can be defined as an inner product with (a, 1-a). In the present disclosure, the compensation transformation function is shared between the base station and the terminal by being predefined or set, and is not limited to a specific function.

[0226] The action set of this embodiment has the following characteristics.

[0227] - A set containing at least some of all possible combinations of components of a CSI report that a terminal can transmit to a base station.

[0228] - Message size or resource size information corresponding to each element included in the set (e.g., combination of number of RBs and transmission power)

[0229] - In a set, the indices of elements are sorted in descending or ascending order based on message size or resource size information.

[0230] Components of a CSI report may include a set of TRPs, the number of beams being reported (e.g., the M value used when reporting the best-M beams), absolute and / or relative quantization levels of quality information (e.g., RSRP or SINR), a reporting period, and a reporting method (e.g., A(aperiodic) / P(period) / SP(semi-persistent)).

[0231]

[0232] [Example #2] SRS-based method using a base station model

[0233] This embodiment can be applied when a base station possesses an AI / ML model. When performing beam management, the base station may rely on the terminal's SRS transmission and feedback reporting of additional information. If performance is improved due to feedback, compensation with the following characteristics may be provided.

[0234] - The quality (e.g. SINR) of a set of serving beams transmitted from one or more SRS resources, or the difference between the previous and current measurement values ​​of the corresponding parameter.

[0235] - The beam alignment quality of the SRS estimated at the base station (e.g. SINR / RSRP) or the difference value between the previous measurement value and the current measurement value of the corresponding parameter.

[0236] When an SRS is transmitted, compensation for the cost can be given as follows:

[0237] - Number of RBs and transmission power for SRS per unit time, or the difference between the resources of the previous transmission and the current resources per unit time.

[0238] Since the compensation includes two elements, it can be expressed as a vector value. The vector value is converted to a scalar value by a compensation transformation function. For example, the compensation transformation function can be defined as a weighted sum of two compensation elements. As another example, the compensation transformation function can be defined as an inner product with (a, 1-a). In the present disclosure, the compensation transformation function is shared between the base station and the terminal by being predefined or set, and is not limited to a specific function.

[0239] The present disclosure proposes a set of actions. The set of actions has the following characteristics.

[0240] - A set of identifiers, indicators or indices corresponding to multiple candidate settings and settings for transmitting SRS from a terminal to a base station 1:1

[0241] - Message size or resource size information (e.g. number of RBs and transmit power combination) corresponding to each element of the set (e.g. multiple candidate configurations)

[0242] - In a set, the indices of the elements are sorted in descending or ascending order of the corresponding message size or resource size.

[0243]

[0244] [Example #3] CSI-RS-based method using a terminal-side model

[0245] This embodiment can be applied when a terminal possesses an AI / ML model. In beam management, the terminal can manage downlink beams based on CSI-RS transmissions. Compensation with the following characteristics can be provided.

[0246] - Quality values ​​of a set of serving beams transmitted from one or more CSI or SSB resources (e.g., cri-RSRP, ssb-index-RSRP, cri-SINR, ssb-index-SINR values), or the difference between the values ​​of previous measurements of the corresponding parameter and the current measurement value.

[0247] - Quality (e.g. RSRP, SINR) related to the beam alignment estimate value of CSI-RS or SSB estimated at the terminal, or the difference value between the previous measurement value and the current measurement value of the corresponding parameter.

[0248] If measurements are performed, compensation for the costs may be given as follows:

[0249] - The number of RS resources for each CSI-RS that is an element of a set of multiple CSI-RSs or the difference in resource cost from the previous point in time.

[0250] Since the compensation includes two elements, it can be expressed as a vector value. The vector value is converted to a scalar value by a compensation transformation function. For example, the compensation transformation function can be defined as a weighted sum of two compensation elements. As another example, the compensation transformation function can be defined as an inner product with (a, 1-a). In the present disclosure, the compensation transformation function is shared between the base station and the terminal by being predefined or set, and is not limited to a specific function.

[0251] The action set of this embodiment has the following characteristics.

[0252] - A set containing at least some of all possible combinations of components for CSI-RS measurements.

[0253] - In a set, the indices of the elements are sorted in descending or ascending order of the corresponding message size or resource size.

[0254]

[0255] [Example #4] SRS-based method using a terminal-side model

[0256] This embodiment can be applied when a terminal possesses an AI / ML model. For beam management, the terminal can operate the model using measurements of the base station's SRS. If performance is improved through feedback, compensation with the following characteristics can be provided.

[0257] - The quality (e.g. SINR) of a set of beams transmitted from one or more SRS resources or the difference between the previous measurement value and the current measurement value of that parameter.

[0258] - Estimated alignment performance value of the SRS transmission / reception beam of the base station (e.g. SINR) or the difference value between the previous measurement value and the current measurement value of the corresponding parameter.

[0259] When an SRS is transmitted, compensation for the cost may be given as follows:

[0260] - Number of RBs and transmission power resources for SRS per unit time, or the difference between the previous transmission resources and the current transmission resources per unit time.

[0261] Since the compensation includes two elements, it can be expressed as a vector value. The vector value is converted to a scalar value by a compensation transformation function. For example, the compensation transformation function can be defined as a weighted sum of two compensation elements. As another example, the compensation transformation function can be defined as an inner product with (a, 1-a). In the present disclosure, the compensation transformation function is shared between the base station and the terminal by being predefined or set, and is not limited to a specific function.

[0262] The action set of this embodiment has the following characteristics.

[0263] - A set of identifiers, indicators or indices corresponding to multiple candidate settings and settings for transmitting SRS from a terminal to a base station 1:1

[0264] - Message size or resource size information (e.g., combination of number of RBs and transmit power) corresponding to each element of the set (e.g., multiple candidate configurations).

[0265] - In a set, the indices of the elements are sorted in descending or ascending order of the corresponding message size or resource size.

[0266]

[0267] In the above-described embodiments, it has been described that the elements included in the action set are sorted in descending or ascending order based on resource size, etc. This means that the indexing of the elements is based on resource size, etc. For example, when sorted in descending order of resource size, the resource size of the element at index #0 is greater than the resource size of the element at index #1.

[0268] However, according to other embodiments, the elements may not be sorted, and auxiliary information (e.g., a mapping table, etc.) may be defined that allows the elements to be sequentially selected in descending or ascending order of resource size, etc. For example, the auxiliary information may include a list in which the indices of the elements are sorted in descending or ascending order of resource size, etc.

[0269]

[0270] The present disclosure proposes the following procedures for reinforcement learning-based beam management feedback optimization to solve the aforementioned optimization problem.

[0271]

[0272] FIG. 15 illustrates an example of a procedure for providing capability information according to one embodiment of the present disclosure. FIG. 15 illustrates signal exchange for acquiring capability information between a base station (1520) and a terminal (1510).

[0273] Referring to FIG. 15, in step S1501, the base station (1520) transmits a message (e.g., a BMFO_CAPA_REQ message) to the terminal (1510) to inquire whether the terminal supports the beam management feedback optimization (BMFO) function. In step S1503, the terminal (1510) transmits a message (e.g., a BMFO_CAPA_RES message) to indicate whether the terminal supports the beam management feedback optimization function. For example, the message may include an indicator related to whether beam management feedback optimization is supported, information related to the supported optimization method (e.g., model position, signal used, etc.), etc. Through a procedure such as FIG. 15, it can be mutually confirmed whether the terminal (1510) has the capability to support data collection for beam management feedback optimization with the base station (1520).

[0274]

[0275] Once capability information is confirmed, as shown in Figure 15, a data collection procedure can be performed. The data collection procedure for the base station model is required to train the base station's reinforcement learning value neural network or policy neural network. The data collection procedure for the terminal model is required to train the terminal's reinforcement learning value neural network or policy neural network. During the data collection procedure, information related to rewards and actions is exchanged between the base station and terminal, and data is collected.

[0276]

[0277] FIG. 16 illustrates an example of a procedure for collecting data using a downlink reference signal in a base station-side model situation according to one embodiment of the present disclosure. FIG. 16 illustrates signal exchange for a data collection procedure between a base station (1620) and a terminal (1610).

[0278] Referring to FIG. 16, in step S1601, the base station (1620) transmits a message (e.g., a BMFO_DC_CSI_REQ (BMFO data collection request for CSI report) message) to the terminal (1610) to notify the start of a data collection procedure. The message may include information related to an action set of a CSI report. The action set is a set of multiple CSI-RS report configurations, and elements within the set may be sorted in descending or ascending order according to the message size or radio resource cost for each report. Each element of the set has a unique identifier, index, or indicator. In addition, the message may be used to transmit a function that converts a compensation vector into a scalar. Then, in step S1603, the terminal (1610) transmits a message (e.g., a BMFO_DC_CSI_RES (BMFO data collection response for CSI report) message) to the base station (1620) to notify the start of the data collection procedure. Accordingly, the data collection process can then begin.

[0279] In step S1605, the base station (1620) transmits a message (e.g., a BMFO_DC_CSI_IND message) requesting a CSI report to the terminal (1610). That is, the base station (1620) can request a report according to one of the elements included in the behavior set transmitted in step S1601 through signaling of the PHY, MAC, or RRC layer. Here, one of the elements of the behavior set can be indicated by an identifier, an index, or an indicator. Alternatively, if one of the elements of the behavior set is an element k steps above or below the most recently indicated element, the relative index difference k can be signaled. Then, in step S1607, the base station (1620) transmits a CSI-RS to the terminal (1610).

[0280] In step S1609, the terminal (1610) transmits a message including a CSI report (e.g., a BMFO_DC_CSI_RPT message) to the base station (1620). That is, the terminal (1610) can transmit a compensation. Here, the compensation can include the following information. For example, the compensation can include at least one of an RSRP, SINR, or equivalent link quality value related to the transmission / reception beam alignment quality of the CSI-RS estimated by the terminal (1610), a CQI value of the CSI-RS estimated by the terminal (1610), or an RSRP value of the CSI-RS estimated by the terminal (1610). Here, the RSRP, SINR, or link quality value can be replaced with a difference value between a previous specific value of the corresponding parameter and a current measurement value. Accordingly, the base station (1620) can determine the compensation based on the CSI report including the RSRP value of the CSI-RS, etc. Here, each reward value is normalized to a range from 0 to 1, forming a feedback reward and a reward vector. A reward determined by a scalar transformation function can be applied.

[0281] Although not shown in FIG. 16, the terminal (1610) or the base station (1620) can terminate the data collection procedure by transmitting a message requesting termination of the currently ongoing data collection procedure (e.g., a BMFO_DC_REL (BMFO data release request) message).

[0282]

[0283] FIG. 17 illustrates an example of a procedure for collecting data using an uplink reference signal in a base station-side model situation according to one embodiment of the present disclosure. FIG. 17 illustrates signal exchange for a data collection procedure between a base station (1720) and a terminal (1710).

[0284] Referring to FIG. 17, in step S1701, the base station (1720) transmits a message (e.g., a BMFO_DC_SRS_REQ (BMFO data collection request for SRS) message) to the terminal (1710) to notify the start of a data collection procedure. The message may include information related to a behavior set that includes settings for multiple SRS transmissions as elements. In the behavior set, the elements may be sorted in descending or ascending order of the cost of the radio resource. The elements have a unique identifier, index, or indicator. In addition, the message may be used to transmit a function that converts a reward vector into a scalar. Then, in step S1703, the terminal (1710) transmits a message (e.g., a BMFO_DC_SRS_RES (BMFO data collection response for SRS) message) to the base station (1720) to notify the start of the data collection procedure. Accordingly, the data collection procedure may then begin.

[0285] In step S1705, the base station (1720) transmits a message (e.g., a BMFO_DC_SRS_IND message) requesting transmission of an SRS to the terminal (1710). That is, the base station (1720) can request SRS transmission according to one of the elements included in the behavior set transmitted in step S1701 through signaling of the PHY, MAC, or RRC layer. Here, one of the elements of the behavior set can be indicated by an identifier, an index, or an indicator. Alternatively, if one of the elements of the behavior set is an element k steps above or below the most recently indicated element, the relative index difference k can be signaled.

[0286] In step S1707, the base station (1720) transmits an SRS to the terminal (1710). Accordingly, the base station (1720) can perform measurements on the SRS and determine compensation based on the SRS. For example, the compensation may include RSRP, SINR, or an equivalent link quality value related to the transmission and reception beam alignment quality of the SRS estimated by the base station (1720). Here, the RSRP, SINR, or link quality value may be replaced with a difference value between a previous specific value of the corresponding parameter and a current measurement value.

[0287] Although not shown in FIG. 17, the terminal (1710) or the base station (1720) can terminate the data collection procedure by transmitting a message (e.g., a BMFO_DC_REL message) requesting termination of the currently ongoing data collection procedure.

[0288]

[0289] FIG. 18 illustrates an example of a procedure for collecting data using a downlink reference signal in a terminal-side model situation according to one embodiment of the present disclosure. FIG. 18 illustrates signal exchange for the data collection procedure between a base station (1820) and a terminal (1810).

[0290] Referring to FIG. 18, in step S1801, the terminal (1810) transmits a message (e.g., a BMFO_DC_CSI_REQ (BMFO data collection request for CSI-RS) message) to the base station (1820) to notify the start of a data collection procedure. The message may include information related to a behavior set that includes configurations for multiple CSI-RS transmissions as elements. In the behavior set, the elements may be sorted in descending or ascending order of the cost of the radio resource. The elements have a unique identifier, index, or indicator. In addition, the message may be used to transmit a function that converts a reward vector into a scalar. Then, in step S1803, the base station (1820) transmits a message (e.g., a BMFO_DC_CSI_RES (BMFO data collection response for CSI-RS) message) to the terminal (1810) to notify confirmation of the start of the data collection procedure. Accordingly, the data collection procedure may then begin.

[0291] In step S1805, the terminal (1810) transmits a message (e.g., a BMFO_DC_CSI-IND message) requesting CSI-RS transmission to the base station (1820). That is, the terminal (1810) can request CSI-RS transmission according to one of the elements included in the behavior set transmitted in step S1801 through signaling of the PHY, MAC, or RRC layer. Here, one of the elements of the behavior set can be indicated by an identifier, an index, or an indicator. Alternatively, if one of the elements of the behavior set is an element k steps above or below the most recently indicated element, the relative index difference k can be signaled.

[0292] In step S1807, the base station (1820) transmits a CSI-RS to the terminal (1810). Accordingly, the terminal (1810) may perform measurement on the CSI-RS and determine compensation based on the CSI-RS. For example, the compensation may include RSRP, SINR, or an equivalent link quality value related to the transmission and reception beam alignment quality of the CSI-RS estimated by the terminal (1810). Here, the RSRP, SINR, or link quality value may be replaced with a difference value between a previous specific value of the corresponding parameter and a current measurement value. Here, each compensation value is a value normalized to a range of 0 to 1, and forms a feedback compensation and a compensation vector. A compensation determined by a scalar transformation function may be applied.

[0293] Although not shown in FIG. 18, the terminal (1810) or the base station (1820) can terminate the data collection procedure by transmitting a message (e.g., a BMFO_DC_REL message) requesting termination of the currently ongoing data collection procedure.

[0294]

[0295]

[0296] FIG. 19 illustrates an example of a procedure for collecting data using an uplink reference signal in a terminal-side model situation according to one embodiment of the present disclosure. FIG. 19 illustrates signal exchange for the data collection procedure between a base station (1920) and a terminal (1910).

[0297] Referring to FIG. 19, in step S1901, the terminal (1910) transmits a message (e.g., a BMFO_DC_SRS_REQ (BMFO data collection request for SRS report) message) to the base station (1920) to notify the start of a data collection procedure. The message may include information related to a behavior set that includes settings for multiple SRS transmissions as elements. In the behavior set, the elements may be sorted in descending or ascending order of the cost of the radio resource. The elements have a unique identifier, index, or indicator. In addition, the message may be used to transmit a function that converts a reward vector into a scalar. Then, in step S1903, the base station (1920) transmits a message (e.g., a BMFO_DC_SRS_RES (BMFO data collection response for SRS report) message) to the terminal (1910) to notify confirmation of the start of the data collection procedure. Accordingly, the data collection procedure may then begin.

[0298] In step S1905, the terminal (1910) transmits a message (e.g., a BMFO_DC_SRS_IND message) requesting a measurement report for SRS to the base station (1920). That is, the terminal (1910) can request a report according to one of the elements included in the behavior set transmitted in step S1901 through signaling of the PHY, MAC, or RRC layer. Here, one of the elements of the behavior set can be indicated by an identifier, an index, or an indicator. Alternatively, if one of the elements of the behavior set is an element k steps above or below the most recently indicated element, the relative index difference k can be signaled. Then, in step S1907, the terminal (1910) transmits an SRS to the base station (1920).

[0299] In step S1909, the base station (1920) transmits a message (e.g., a BMFO_DC_SRS_RPT message) including a measurement report for the SRS to the terminal (1910). That is, the base station (1920) can transmit a compensation. Here, the compensation can include the following information. For example, the compensation can include RSRP, SINR, or an equivalent link quality value related to the transmission and reception beam alignment quality of the SRS estimated by the base station (1920). Here, the RSRP, SINR, or link quality value can be replaced with a difference value between a previous specific value of the corresponding parameter and a current measurement value. Accordingly, the terminal (1910) can determine a compensation. Here, each compensation value is a value normalized to a range of 0 to 1, and forms a feedback compensation and a compensation vector. A compensation determined by a scalar transformation function can be applied.

[0300] Although not shown in FIG. 19, the terminal (1910) or the base station (1920) can terminate the data collection procedure by transmitting a message (e.g., a BMFO_DC_REL message) requesting termination of the currently ongoing data collection procedure.

[0301]

[0302] As described above, training of AI / ML models for beam management can be performed using reinforcement learning. Application examples of the various embodiments described above are as follows.

[0303]

[0304] Here's an example of how feedback frequency works. Assume N beams are transmitted in free space, free of reflections. The cell planner can observe a maximum UE movement speed of 20 km / h within the cell. In this case, the simplest rule is to select the shortest feedback frequency, which is optimal for 20 km / h, and apply it to all UEs, which is every 10 ms. However, not all UEs need to provide feedback at the same frequency. This is because a UE moving directly toward a specific beam may not change beams. A UE traversing N beams orthogonally (e.g., moving between beam coverage areas) should transmit feedback at the fastest frequency, which is every 10 ms. A UE moving directly toward a specific beam (e.g., moving toward the base station within the coverage area of ​​a specific beam) will not experience any degradation in SNR even if it transmits feedback at a frequency of 100 ms. Applying the proposed technique, these aspects can be adaptively controlled by reinforcement learning's "state," "value (e.g., action and state)" and "reinforcement through learning." Frequent feedback actions while looking straight ahead will not provide benefits from beam changes, so they will be negatively compensated by feedback. In other words, a 100ms feedback action will provide more benefits than a 10ms one. For a terminal crossing beams, it would be preferable to use a 10ms cycle because the negative compensation from feedback will be amortized only if the compensation from beam changes is provided at a 10ms cycle. The above method works effectively because it distinguishes the states of the terminal in detail and reinforces the so-called necessary "value" by determining which action maximizes the cumulative final average gain in each state. For a policy p, a value function q(s, a) = E[Return|S=s,A=a] of state s for action a is learned.At this time, the final policy p* is p(s=crossing state)=arg max q(s,{10ms,100ms})=10ms, p(s=front facing state)=arg max q(s,{10ms,100ms})=100ms.

[0305]

[0306] FIG. 20 illustrates an example of a procedure for performing learning for a terminal-side model according to one embodiment of the present disclosure. FIG. 20 illustrates a method performed by a terminal.

[0307] Referring to FIG. 20, in step S2001, the terminal transmits capability information. For this purpose, although not illustrated in FIG. 20, the terminal may receive a message requesting capability information. The capability information may include information related to hardware or software capabilities related to communication performance of the terminal, information related to procedures, functions, or features supported by the terminal. At this time, according to various embodiments of the present disclosure, the capability information may include information related to an optimization procedure for feedback for beam management based on reinforcement learning (hereinafter, “beam management feedback optimization procedure”). Specifically, the capability information may include at least one of information related to whether the beam management feedback optimization procedure is supported, information related to a signal used for the beam management feedback optimization procedure, and information related to the position of a model for the beam management feedback optimization procedure.

[0308] In step S2003, the terminal performs signaling of configuration information for learning. The signaling of the configuration information includes at least one of transmitting or receiving a request message for data collection required for learning, transmitting or receiving a response message to the request message, and transmitting or receiving an instruction message for transmitting a reference signal. At this time, the request message includes information on a set of actions for reinforcement learning, and the set of actions may include settings for terminal operations related to the reference signal. In addition, the request message may include information related to a function for converting a reward vector into a scalar value.

[0309] In step S2005, the terminal receives configuration information related to the reference signal. For example, the configuration information related to the reference signal may include at least one of information related to resources allocated for the reference signal and information related to the sequence of the reference signal. For example, the configuration information may include information related to a downlink reference signal (e.g., CSI-RS) or an uplink reference signal (e.g., SRS). In addition, the configuration information related to the reference signal may include information related to reporting of measurements for the reference signal. For example, the information related to reporting of measurements may include at least one of information related to resources for measurement, information related to reported items, and information related to a reporting method (e.g., periodic, aperiodic, quasi-static). According to one embodiment, the information related to reporting of measurements may include information related to one of the elements included in the behavior set. According to another embodiment, the instruction message for transmitting the reference signal described in step S2003 may be replaced with the configuration information of this step.

[0310] In step S2007, the terminal transmits or receives a reference signal. The terminal may transmit or receive the reference signal based on the configuration information received in step S2005. That is, the terminal may transmit an uplink reference signal or receive a downlink reference signal through a resource indicated by the configuration information. Here, the reference signal may be repeatedly transmitted through multiple beams. For example, the downlink reference signal may include a CSI-RS, and the uplink reference signal may include an SRS.

[0311] In step S2009, the terminal performs an operation for learning based on the reference signal. According to various embodiments, the operation for learning may include at least one of measuring the reference signal, reporting the measurement for the reference signal, determining a reward for reinforcement learning, and feedback of the reward. For example, the terminal may receive a downlink reference signal, measure the downlink reference signal, and report the measurement result to the base station or perform model learning using the measurement result. As another example, the terminal may transmit an uplink reference signal, receive the measurement result for the uplink reference signal from the base station, and perform model learning using the measurement result. However, if the AI / ML model exists on the base station side and the uplink reference signal is used, this step may be omitted.

[0312]

[0313] FIG. 21 illustrates an example of a procedure for performing learning for a base station-side model according to one embodiment of the present disclosure. FIG. 21 illustrates a method performed by a base station.

[0314] Referring to FIG. 21, in step S2101, the base station transmits capability information. For this purpose, although not illustrated in FIG. 21, the base station may transmit a message requesting capability information to the terminal. The capability information may include information related to hardware or software capabilities related to communication performance of the terminal, information related to procedures, functions, or characteristics supported by the terminal. At this time, according to various embodiments of the present disclosure, the capability information may include information related to a feedback optimization procedure for beam management based on reinforcement learning (hereinafter, “beam management feedback optimization procedure”). Specifically, the capability information may include at least one of information related to whether the beam management feedback optimization procedure is supported, information related to a signal used for the beam management feedback optimization procedure, and information related to the position of a model for the beam management feedback optimization procedure. Through this, the base station can confirm that the terminal supports the beam management feedback optimization procedure and determine whether and how to perform the beam management feedback optimization procedure.

[0315] In step S2103, the base station performs signaling of configuration information for learning. The signaling of the configuration information includes at least one of transmitting or receiving a request message for data collection required for learning, transmitting or receiving a response message to the request message, and transmitting or receiving an instruction message for transmission of a reference signal. At this time, the request message includes information on a set of actions for reinforcement learning, and the set of actions may include settings for terminal operations related to the reference signal. In addition, the request message may include information related to a function for converting a reward vector into a scalar value.

[0316] In step S2105, the base station transmits configuration information related to the reference signal. For example, the configuration information related to the reference signal may include at least one of information related to resources allocated for the reference signal and information related to the sequence of the reference signal. For example, the configuration information may include information related to a downlink reference signal (e.g., CSI-RS) or an uplink reference signal (e.g., SRS). In addition, the configuration information related to the reference signal may include information related to reporting of measurements for the reference signal. For example, the information related to reporting of measurements may include at least one of information related to resources for measurement, information related to reported items, and information related to a reporting method (e.g., periodic, aperiodic, quasi-static). According to one embodiment, the information related to reporting of measurements may include information related to one of the elements included in the behavior set. According to another embodiment, the instruction message for transmitting the reference signal described in step S2103 may be replaced with the configuration information of this step.

[0317] In step S2107, the base station transmits or receives a reference signal. The base station can transmit or receive the reference signal based on the configuration information transmitted in step S2105. That is, the terminal can transmit an uplink reference signal or receive a downlink reference signal through a resource indicated by the configuration information. Here, the reference signal can be repeatedly transmitted through multiple beams. For example, the downlink reference signal can include a CSI-RS, and the uplink reference signal can include an SRS.

[0318] In step S2109, the base station performs an operation for learning based on the reference signal. According to various embodiments, the operation for learning may include at least one of measuring the reference signal, reporting the measurement for the reference signal, determining a reward for reinforcement learning, and feedback of the reward. For example, the base station may receive an uplink reference signal, measure the uplink reference signal, and report the measurement result to the terminal or perform model learning using the measurement result. As another example, the base station may transmit a downlink reference signal, receive the measurement result for the downlink reference signal from the terminal, and perform model learning using the measurement result. However, if the AI / ML model exists on the terminal side and the downlink reference signal is used, this step may be omitted.

[0319]

[0320] According to the various embodiments described above, an AI / ML model can be trained to determine an appropriate feedback format. After the training of the AI / ML model according to the embodiments described above is completed, a beam management procedure using the trained AI / ML model can be performed. At this time, in the beam management procedure, signaling may be performed to notify available formats for measurement reports or reference signals. In other words, after the reinforcement learning of the AI / ML model is completed, signaling of information related to a set of actions for beam management operations may be performed. That is, after the reinforcement learning of the AI / ML model is completed, the base station can transmit information related to a set of actions for beam management operations to the terminal. However, if the set of actions used for training is the same as the set of actions used for beam management operations, the terminal participating in the training may not receive information related to the set of actions for beam management operations.

[0321]

[0322] Figures 22a and 22b illustrate the Pareto optimal regions of beam performance and feedback cost depending on whether the feedback period is fixed. As shown in Figure 22a, when the feedback period is fixed, feedback efficiency is limited. On the other hand, as shown in Figure 22b, when the proposed technique is applied, feedback efficiency increases.

[0323]

[0324] Here's an example of using CSI-RS: An AI / ML model is deployed at a base station, and a terminal transmits a CSI report. Here, the action set includes a set of multiple CSI reporting configurations that the terminal can use for reporting, and within the set, the CSI reporting configurations are sorted in descending or ascending order according to message size per unit time. Each element of the set has a unique indicator, and the base station can indicate at least one element to the terminal at the RRC, MAC, or PHY layer. For example, if CSI reporting is performed periodically, let's assume that there are eight elements with slot periods of 4, 5, 8, 10, 16, 20, 40, and 80. If the CSI report is configured to report measurements for N beams, the feedback transmission rate will be highest for period 4 and lowest for period 80 among the eight reporting configurations. In other words, the elements are sorted in descending or ascending order of feedback transmission rate. The set can be expressed as {k1, k2, …, k8}, the action set corresponding to each element can be expressed as {a1, a2, …, a8}, and the negative reward can be expressed as {r1, r2, …, r8}. Additionally, in this application example, in order to reduce the variance of the cumulative reward, the action set can be defined as a relative concept. For example, an action of increasing the feedback transmission rate by one step can be defined as a1, and an action of decreasing it by one step can be defined as a2. Here, the reward can include a combination of the number of RBs required to transmit a message and the transmission power, or the message size.

[0325] An application example of SRS is as follows. An AI / ML model is deployed at a base station, and a terminal transmits an SRS. Here, the action set includes a set of multiple SRS configurations that the terminal can transmit, and the SRS configurations in the set are sorted in descending or ascending order according to the cost of frequency and time-space resources. The base station can instruct the terminal to transmit the SRS for each configuration. The arrangement of resource elements for the SRS can take various forms. For example, the SRS can occupy the 1st, 2nd, and 4th symbols in the time domain, and a certain number of RBs within the bandwidth part (BWP) in the frequency domain. The multiple SRS configurations can be sorted according to the size of the feedback resource. Alternatively, the increase and decrease of the feedback resource can be defined as elements of the action set.

[0326] An application example of using multi-TRP is as follows. A terminal includes a multi-panel with N panels, and there can be M TRPs. In the process of selecting N beams among M, RSRP reports for the M beams can be transmitted. Depending on the status of the terminal located at a specific point in a specific cell, the optimal M can be selected according to the proposed technique. Among the M candidates that do not contribute to improving SNR or RSRP, the candidate with the lowest transmission rate can be selected from the action set.

[0327] An application example related to the quantization level is as follows. When a terminal reports multiple RSRP values, one reference beam can be quantized into x bits, and the difference from the reference beam can be quantized into y bits for at least one remaining beam. In this case, the behavior set includes multiple transmission formats. For example, within the behavior set, elements can be sorted in descending or ascending order, such as {(x,y)|(7,4),(8,4),(4,3),(3,1)}.

[0328] When performing beam management between a base station and a terminal, the exchange of channel information between the transmitter and receiver is essential. This exchange of channel information involves transmitting and receiving reference signals and beam measurement reports. According to various embodiments of the present disclosure, learning can be performed to minimize feedback depending on channel conditions and transitions. For example, the measurement and feedback cycles can be determined through reinforcement learning. When a terminal traverses multiple beams transmitted from a single TRP, a short feedback cycle is desirable because the beams frequently change. In this case, a longer feedback cycle deteriorates beam alignment performance, so a short-cycle feedback action can be suggested by an action value function or policy function. Even with high mobility, when moving toward or away from a TRP within the coverage area of ​​a beam transmitted from the TRP, this state space transition does not result in frequent beam changes. Therefore, reinforcement learning assigns a higher value to the action so that the feedback cycle of the terminal traversing the beams is longer. This feedback method not only conserves radio resources between the base station and the terminal, but also benefits terminal power savings. The power of the terminal can be maintained optimally by setting the optimal cycle.

[0329] As another example, feedback resource costs can be optimized by defining a set of beams transmitted from multiple TRPs in the action space and filtering out TRPs that are unnecessary for performance improvement. Specifically, by initially setting actions with high feedback resource costs to exploitation in reinforcement learning, and setting remaining actions with lower costs to exploration, reinforcement learning can be performed to maximize the system's beam alignment performance while minimizing feedback resources. This allows the system to operate stably.

[0330]

[0331] Below, examples of wireless device utilization to which various embodiments of the present disclosure are applied are described.

[0332] Figure 23 illustrates an example of a wireless device applicable to the present disclosure. The wireless device may be implemented in various forms depending on the use case / service (see Figure 1).

[0333] Referring to FIG. 23, the wireless device (200) corresponds to the wireless device (200) of FIG. 2 and may be composed of various elements, components, units / units, and / or modules. For example, the wireless device (200) may include a communication unit (210), a control unit (220), a memory unit (230), and additional elements (240). The communication unit may include a communication circuit (212) and a transceiver(s) (214). For example, the communication circuit (212) may include one or more processors (202) and / or one or more memories (204) of FIG. 2. For example, the transceiver(s) (214) may include one or more transceivers (206) and / or one or more antennas (208) of FIG. 2. The control unit (220) is electrically connected to the communication unit (210), the memory unit (230), and the additional elements (240) and controls the overall operations of the wireless device. For example, the control unit (220) can control the electrical / mechanical operations of the wireless device based on the program / code / command / information stored in the memory unit (230). In addition, the control unit (220) can transmit information stored in the memory unit (230) to an external device (e.g., another communication device) via a wireless / wired interface through the communication unit (210), or store information received from an external device (e.g., another communication device) via a wireless / wired interface in the memory unit (230).

[0334] The additional element (240) may be configured in various ways depending on the type of the wireless device. For example, the additional element (240) may include at least one of a power unit / battery, an input / output unit (I / O unit), a driving unit, and a computing unit. Although not limited thereto, the wireless device may be implemented in the form of a robot (Fig. 1, 100a), a vehicle (Fig. 1, 100b-1, 100b-2), an XR device (Fig. 1, 100c), a portable device (Fig. 1, 100d), a home appliance (Fig. 1, 100e), an IoT device (Fig. 1, 100f), a digital broadcasting terminal, a hologram device, a public safety device, an MTC device, a medical device, a fintech device (or a financial device), a security device, a climate / environmental device, an AI server / device (Fig. 1, 400), a base station (Fig. 1, 200), a network node, etc. Wireless devices may be mobile or stationary depending on the use / service.

[0335] In FIG. 23, various elements, components, units / parts, and / or modules within the wireless device (200) may be entirely interconnected via a wired interface, or at least some may be wirelessly connected via a communication unit (210). For example, within the wireless device (200), the control unit (220) and the communication unit (210) may be wired, and the control unit (220) and a first unit (e.g., 230, 240) may be wirelessly connected via the communication unit (210). In addition, each element, component, unit / part, and / or module within the wireless device (200) may further include one or more elements. For example, the control unit (220) may be composed of a set of one or more processors. For example, the control unit (220) may be composed of a set of a communication control processor, an application processor, an electronic control unit (ECU), a graphics processing processor, a memory control processor, etc. As another example, the memory unit (130) may be composed of RAM (Random Access Memory), DRAM (Dynamic RAM), ROM (Read Only Memory), flash memory, volatile memory, non-volatile memory, and / or a combination thereof.

[0336] Below, the implementation example of Fig. 23 is described in more detail with reference to the drawings.

[0337] Figure 24 illustrates examples of portable devices applicable to the present disclosure. Portable devices may include smartphones, smart pads, wearable devices (e.g., smartwatches, smartglasses), and portable computers (e.g., laptops, etc.). Portable devices may also be referred to as mobile stations (MS), user terminals (UT), mobile subscriber stations (MSS), subscriber stations (SS), advanced mobile stations (AMS), or wireless terminals (WT).

[0338] Referring to FIG. 24, the portable device (200) may include an antenna unit (208), a communication unit (210), a control unit (220), a memory unit (230), a power supply unit (240a), an interface unit (240b), and an input / output unit (240c). The antenna unit (208) may be configured as a part of the communication unit (210). Blocks 210 to 230 / 240a to 240c of FIG. 24 correspond to blocks 210 to 230 / 240 of FIG. 23, respectively.

[0339] The communication unit (210) can transmit and receive signals (e.g., data, control signals, etc.) with other wireless devices and base stations. The control unit (220) can control components of the mobile device (200) to perform various operations. The control unit (220) can include an AP (Application Processor). The memory unit (230) can store data / parameters / programs / codes / commands required for operating the mobile device (200). In addition, the memory unit (230) can store input / output data / information, etc. The power supply unit (240a) supplies power to the mobile device (200) and can include a wired / wireless charging circuit, a battery, etc. The interface unit (240b) can support connection between the mobile device (200) and other external devices. The interface unit (240b) can include various ports (e.g., audio input / output ports, video input / output ports) for connection with external devices. The input / output unit (240c) can input or output video information / signals, audio information / signals, data, and / or information input from a user. The input / output unit (240c) may include a camera, a microphone, a user input unit, a display unit (240d), a speaker, and / or a haptic module.

[0340] For example, in the case of data communication, the input / output unit (240c) obtains information / signals (e.g., touch, text, voice, image, video) input by the user, and the obtained information / signals can be stored in the memory unit (230). The communication unit (210) converts the information / signals stored in the memory into wireless signals, and can directly transmit the converted wireless signals to other wireless devices or to a base station. In addition, the communication unit (210) can receive wireless signals from other wireless devices or base stations, and then restore the received wireless signals to the original information / signals. The restored information / signals can be stored in the memory unit (230) and then output in various forms (e.g., text, voice, image, video, haptic) through the input / output unit (240c).

[0341] Figure 25 illustrates examples of vehicles or autonomous vehicles applicable to the present disclosure. The vehicles or autonomous vehicles may be implemented as mobile robots, cars, trains, manned / unmanned aerial vehicles (AVs), ships, etc.

[0342] Referring to FIG. 25, a vehicle or autonomous vehicle (200-1) may include an antenna unit (208-1), a communication unit (210-1), a control unit (220-1), a driving unit (240a-1), a power supply unit (240b-1), a sensor unit (240c-1), and an autonomous driving unit (240d-1). The antenna unit (208-1) may be configured as a part of the communication unit (210-1). Blocks 210-1 / 230-1 / 240a-1 to 240d-1 of FIG. 25 correspond to blocks 210 / 230 / 240 of FIG. 23, respectively.

[0343] The communication unit (210-1) can transmit and receive signals (e.g., data, control signals, etc.) with external devices such as other vehicles, base stations (e.g., base stations, roadside base stations (ROS), etc.), and servers. The control unit (220-1) can control elements of the vehicle or autonomous vehicle (200-1) to perform various operations. The control unit (220-1) may include an ECU (Electronic Control Unit). The drive unit (240a-1) can drive the vehicle or autonomous vehicle (200-1) on the ground. The drive unit (240a-1) may include an engine, a motor, a power train, wheels, brakes, a steering device, etc. The power supply unit (240b-1) supplies power to the vehicle or autonomous vehicle (200-1) and may include a wired / wireless charging circuit, a battery, etc. The sensor unit (240c-1) can obtain vehicle status, surrounding environment information, user information, etc. The sensor unit (240c-1) may include an IMU (inertial measurement unit) sensor, a collision sensor, a wheel sensor, a speed sensor, an incline sensor, a weight detection sensor, a heading sensor, a position module, a vehicle forward / backward sensor, a battery sensor, a fuel sensor, a tire sensor, a steering sensor, a temperature sensor, a humidity sensor, an ultrasonic sensor, an illuminance sensor, a pedal position sensor, etc. The autonomous driving unit (240d-1) may implement a technology for maintaining a driving lane, a technology for automatically controlling speed such as adaptive cruise control, a technology for automatically driving along a set path, a technology for automatically setting a path and driving when a destination is set, etc.

[0344] For example, the communication unit (210-1) can receive map data, traffic information data, etc. from an external server. The autonomous driving unit (240d-1) can generate an autonomous driving route and driving plan based on the acquired data. The control unit (220-1) can control the drive unit (240a-1) so that the vehicle or autonomous vehicle (200-1) moves along the autonomous driving route according to the driving plan (e.g., speed / direction control). During autonomous driving, the communication unit (210-1) can irregularly / periodically acquire the latest traffic information data from an external server and can acquire surrounding traffic information data from surrounding vehicles. In addition, during autonomous driving, the sensor unit (240c-1) can acquire vehicle status and surrounding environment information. The autonomous driving unit (240d-1) can update the autonomous driving route and driving plan based on newly acquired data / information. The communication unit (210-1) can transmit information regarding the vehicle location, autonomous driving route, driving plan, etc. to an external server. The external server can predict traffic information data in advance using AI technology, etc. based on information collected from the vehicle or autonomous vehicles, and provide the predicted traffic information data to the vehicle or autonomous vehicles. If the device (220-2) is an autonomous vehicle, it can perform the same procedure as the vehicle or autonomous vehicle (200-1). In addition, if the device (220-2) is a base station or a roadside base station, the device (220-2) can transmit data, control signals, etc. to the vehicle or autonomous vehicle (200-1) through the communication unit (210-2).

[0345] Figure 26 illustrates an example of a vehicle applicable to the present disclosure. The vehicle may also be implemented as a means of transportation, a train, an aircraft, a ship, etc. Referring to Figure 26, the vehicle (200) may include a communication unit (210), a control unit (220), a memory unit (230), an input / output unit (240a), and a position measurement unit (240b). Here, blocks 210 to 230 / 240a to 240b correspond to blocks 210 to 230 / 240 of Figure 23, respectively.

[0346] The communication unit (210) can transmit and receive signals (e.g., data, control signals, etc.) with other vehicles or external devices such as base stations. The control unit (220) can control components of the vehicle (200) to perform various operations. The memory unit (230) can store data / parameters / programs / codes / commands that support various functions of the vehicle (100). The input / output unit (240a) can output AR / VR objects based on information in the memory unit (230). The input / output unit (240a) can include a HUD. The position measurement unit (240b) can obtain position information of the vehicle (200). The position information can include absolute position information of the vehicle (200), position information within a driving line, acceleration information, position information with respect to surrounding vehicles, etc. The position measurement unit (240b) can include GPS and various sensors.

[0347] For example, the communication unit (210) of the vehicle (200) can receive map information, traffic information, etc. from an external server and store them in the memory unit (230). The location measurement unit (240b) can obtain vehicle location information through GPS and various sensors and store the information in the memory unit (230). The control unit (220) can create a virtual object based on the map information, traffic information, and vehicle location information, and the input / output unit (240a) can display the created virtual object on the vehicle window (240a-1, 240a-2). In addition, the control unit (220) can determine whether the vehicle (200) is being driven normally within the driving line based on the vehicle location information. If the vehicle (200) abnormally deviates from the driving line, the control unit (220) can display a warning on the vehicle window through the input / output unit (240a). Additionally, the control unit (220) can broadcast a warning message regarding driving abnormalities to surrounding vehicles through the communication unit (210). Depending on the situation, the control unit (220) can transmit vehicle location information and information regarding driving / vehicle abnormalities to relevant authorities through the communication unit (210).

[0348] Figure 27 illustrates examples of XR devices applicable to the present disclosure. The XR devices may be implemented as HMDs, head-up displays (HUDs) installed in vehicles, televisions, smartphones, computers, wearable devices, home appliances, digital signage, vehicles, robots, and the like.

[0349] Referring to FIG. 27, the XR device (200a) may include a communication unit (210), a control unit (220), a memory unit (230), an input / output unit (240a), a sensor unit (240b), and a power supply unit (240c). Here, blocks 210 to 230 / 240a to 240c of FIG. 27 correspond to blocks 210 to 230 / 240 of FIG. 23, respectively.

[0350] The communication unit (210) can transmit and receive signals (e.g., media data, control signals, etc.) with external devices such as other wireless devices, portable devices, or media servers. The media data can include videos, images, sounds, etc. The control unit (220) can control components of the XR device (200a) to perform various operations. For example, the control unit (220) can be configured to control and / or perform procedures such as video / image acquisition, (video / image) encoding, metadata generation and processing, etc. The memory unit (230) can store data / parameters / programs / codes / commands required for driving the XR device (200a) / generating XR objects. The input / output unit (240a) can obtain control information, data, etc. from the outside, and output the generated XR object. The input / output unit (240a) can include a camera, a microphone, a user input unit, a display unit, a speaker, and / or a haptic module. The sensor unit (240b) can obtain the XR device status, surrounding environment information, user information, etc. The sensor unit (240b) may include a proximity sensor, an illuminance sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, and / or a radar. The power supply unit (240c) supplies power to the XR device (200a) and may include a wired / wireless charging circuit, a battery, etc.

[0351] For example, the memory unit (230) of the XR device (200a) may include information (e.g., data, etc.) required for creating an XR object (e.g., AR / VR / MR object). The input / output unit (240a) may obtain a command to operate the XR device (200a) from the user, and the control unit (220) may operate the XR device (200a) according to the user's operating command. For example, when the user attempts to watch a movie, news, etc. through the XR device (200a), the control unit (220) may transmit content request information to another device (e.g., a mobile device (200b)) or a media server through the communication unit (230). The communication unit (230) may download / stream content such as movies and news from another device (e.g., a mobile device (200b)) or a media server to the memory unit (230). The control unit (220) controls and / or performs procedures such as video / image acquisition, (video / image) encoding, and metadata generation / processing for content, and can generate / output an XR object based on information about surrounding space or real objects acquired through the input / output unit (240a) / sensor unit (240b).

[0352] In addition, the XR device (200a) is wirelessly connected to the mobile device (200b) through the communication unit (210), and the operation of the XR device (200a) can be controlled by the mobile device (200b). For example, the mobile device (200b) can act as a controller for the XR device (200a). To this end, the XR device (200a) can obtain 3D location information of the mobile device (200b), and then generate and output an XR object corresponding to the mobile device (200b).

[0353] Figure 28 illustrates examples of robots applicable to the present disclosure. Robots can be classified into industrial, medical, household, and military types, depending on their intended use or field.

[0354] Referring to FIG. 28, the robot (200) may include a communication unit (210), a control unit (220), a memory unit (230), an input / output unit (240a), a sensor unit (240b), and a driving unit (240c). Here, blocks 210 to 230 / 240a to 240c of FIG. 28 correspond to blocks 210 to 230 / 240 of FIG. 23, respectively.

[0355] The communication unit (210) can transmit and receive signals (e.g., driving information, control signals, etc.) with external devices such as other wireless devices, other robots, or control servers. The control unit (220) can control components of the robot (200) to perform various operations. The memory unit (230) can store data / parameters / programs / codes / commands that support various functions of the robot (200). The input / output unit (240a) can obtain information from the outside of the robot (200) and output information to the outside of the robot (200). The input / output unit (240a) can include a camera, a microphone, a user input unit, a display unit, a speaker, and / or a haptic module. The sensor unit (240b) can obtain internal information of the robot (200), surrounding environment information, user information, etc. The sensor unit (240b) may include a proximity sensor, an illuminance sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, a radar, etc. The driving unit (240c) may perform various physical operations, such as moving the robot joints. In addition, the driving unit (240c) may enable the robot (200) to drive on the ground or fly in the air. The driving unit (240c) may include an actuator, a motor, wheels, brakes, propellers, etc.

[0356] Figure 29 illustrates an example of an AI device applicable to the present disclosure.

[0357] AI devices can be implemented as fixed or mobile devices, such as TVs, projectors, smartphones, PCs, laptops, digital broadcasting terminals, tablet PCs, wearable devices, set-top boxes (STBs), radios, washing machines, refrigerators, digital signage, robots, and vehicles.

[0358] Referring to FIG. 29, the AI ​​device (200) may include a communication unit (210), a control unit (220), a memory unit (230), an input / output unit (240a / 240b), a learning processor unit (240c), and a sensor unit (240d). Blocks 210 to 230 / 240a to 240d of FIG. 29 correspond to blocks 210 to 230 / 140 of FIG. 23, respectively.

[0359] The communication unit (210) can transmit and receive wired and wireless signals (e.g., sensor information, user input, learning models, control signals, etc.) with external devices such as other AI devices (e.g., 100a to 100f, 120 of FIG. 1) or AI servers (e.g., 100g of FIG. 1) using wired and wireless communication technology. To this end, the communication unit (210) can transmit information within the memory unit (230) to the external device or transfer a signal received from the external device to the memory unit (230).

[0360] The control unit (220) may determine at least one executable operation of the AI ​​device (200) based on information determined or generated using a data analysis algorithm or a machine learning algorithm. In addition, the control unit (220) may control components of the AI ​​device (200) to perform the determined operation. For example, the control unit (220) may request, search, receive, or utilize data from the learning processor unit (240c) or the memory unit (230), and may control components of the AI ​​device (200) to perform at least one executable operation, a predicted operation, or an operation determined to be desirable. In addition, the control unit (220) may collect history information including the operation contents of the AI ​​device (200) or user feedback on the operation, and store the collected history information in the memory unit (230) or the learning processor unit (240c), or transmit the collected history information to an external device such as an AI server (FIG. 1, 100g). The collected history information may be used to update a learning model.

[0361] The memory unit (230) can store data that supports various functions of the AI ​​device (200). For example, the memory unit (230) can store data obtained from the input unit (240a), data obtained from the communication unit (210), output data of the learning processor unit (240c), and data obtained from the sensing unit (140). In addition, the memory unit (230) can store control information and / or software codes necessary for the operation / execution of the control unit (220).

[0362] The input unit (240a) can obtain various types of data from the outside of the AI ​​device (200). For example, the input unit (220) can obtain learning data for model learning, input data to which the learning model will be applied, etc. The input unit (240a) may include a camera, a microphone, and / or a user input unit. The output unit (240b) may generate output related to vision, hearing, or touch. The output unit (240b) may include a display unit, a speaker, and / or a haptic module, etc. The sensing unit (140d) can obtain at least one of internal information of the AI ​​device (200), information about the surrounding environment of the AI ​​device (200), and user information using various sensors. The sensing unit (140d) may include a proximity sensor, an illuminance sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, and / or a radar, etc.

[0363] The learning processor unit (240c) can train a model composed of an artificial neural network using learning data. The learning processor unit (240c) can perform AI processing together with the learning processor unit of the AI ​​server (Fig. 1, 100g). The learning processor unit (240c) can process information received from an external device via the communication unit (210) and / or information stored in the memory unit (230). In addition, the output value of the learning processor unit (240c) can be transmitted to an external device via the communication unit (210) and / or stored in the memory unit (230).

[0364]

[0365] The proposed methods described above can be implemented independently, but they can also be implemented as a combination (or merge) of some of the proposed methods. Rules can be defined so that the base station notifies the terminal of the applicability of the proposed methods (or information about the rules of the proposed methods) through a predefined signal (e.g., a physical layer signal or a higher layer signal).

[0366] The present disclosure may be embodied in other specific forms without departing from the technical ideas and essential features described herein. Therefore, the above detailed description should not be construed as limiting in all respects but rather as illustrative. The scope of the present disclosure should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present disclosure are intended to be included within the scope of the present disclosure. Furthermore, claims that are not explicitly cited in the claims may be combined to form an embodiment or incorporated into a new claim through a post-filing amendment.

[0367] The embodiments described herein may be applied to various wireless access systems. Examples of such wireless access systems include the 3rd Generation Partnership Project (3GPP) or 3GPP2 systems.

[0368] The embodiments described herein can be applied not only to the various wireless access systems described above, but also to all technical fields utilizing such systems. Furthermore, the proposed method can be applied to mmWave and THz wireless communication systems utilizing ultra-high frequency bands.

[0369] Additionally, the embodiments can be applied to various applications such as autonomous vehicles and drones.

Claims

1. In a method performed by a terminal in a wireless communication system, A step of receiving a first message requesting capability information from a base station; A step of transmitting a second message including capability information to the base station; A step of receiving first configuration information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management from the base station; A step of receiving second setting information related to at least one reference signal from the base station; A step of receiving at least one reference signal based on the second setting information; A step of performing an operation related to the learning based on at least one reference signal, A method in which the first setting information includes information related to a set of actions for reinforcement learning of the AI / ML model, the set including formats of a reference signal or report information corresponding to the reference signal as elements.

2. In claim 1, A method in which the above elements are sorted in descending or ascending order of resource consumption required by each of the above formats within the above action set.

3. In claim 1, The steps for performing the above learning-related actions are: A method comprising the step of determining a reward for the reinforcement learning based on the results of the measurement for the at least one reference signal.

4. In claim 3, A method wherein the compensation comprises a first compensation related to a quality of a channel measured based on the at least one reference signal, and a second compensation related to a format applied to the at least one reference signal or a report corresponding to the at least one reference signal among the formats.

5. In claim 4, The above first reward and the above second reward are converted into one scalar value, A method in which the above scalar value is used as a reward value for the above reinforcement learning.

6. In claim 4, A method in which the above scalar value is determined using the function determined based on information related to the function received from the base station through the first setting information.

7. In claim 1, The formats of the above reporting information are classified by one of the following: a set of transmission / reception points (TRPs), the number of beams being reported, absolute and / or relative quantization levels of quality information, a reporting cycle, and a reporting method.

8. In claim 1, A step of transmitting information related to one of the elements included in the above action set to the base station, A method in which the format indicated by the above element is applied to at least one reference signal or report information corresponding to the at least one reference signal.

9. In claim 8, The above elements are selected based on policies related to exploitation and exploration for the reinforcement learning, The above utilization includes actions that require relatively large resource consumption for each of the formats of the above elements, The above search method includes an action that requires relatively small resource consumption for each of the formats of the above elements.

10. In claim 1, The state for the above reinforcement learning includes the channel state between the base station and the terminal and additional information of the terminal, A method wherein the above additional information includes at least one of the position or speed of the terminal.

11. In claim 1, A method further comprising the step of receiving information related to a set of actions for beam management operations after reinforcement learning of the AI / ML model is completed.

12. In a method performed by a base station in a wireless communication system, A step of transmitting a first message requesting capability information to a terminal; A step of receiving a second message including capability information from the terminal; A step of transmitting first setting information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management to the terminal; A step of transmitting second setting information related to at least one reference signal to the terminal; A step of transmitting at least one reference signal based on the second setting information; A step of performing an operation related to the learning based on at least one reference signal, A method in which the first setting information includes information related to a set of actions for reinforcement learning of the AI / ML model, the set including formats of a reference signal or report information corresponding to the reference signal as elements.

13. In a wireless communication system, at a terminal, Transmitter and receiver; and comprising a processor coupled to the above transceiver, The above processor, Receive a first message requesting capability information from a base station, Transmitting a second message including capability information to the base station, Receive first configuration information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management from the above base station, Receive second setting information related to at least one reference signal from the base station, Receive at least one reference signal based on the second setting information, configured to perform an operation related to the learning based on at least one reference signal; The above first setting information is a terminal including information related to a set of actions for reinforcement learning of the AI / ML model, the set including formats of a reference signal or report information corresponding to the reference signal as elements.

14. In communication devices, At least one processor; At least one computer memory connected to said at least one processor and storing instructions that direct operations when executed by said at least one processor, The above actions are, A step of receiving a first message requesting capability information from a base station; A step of transmitting a second message including capability information to the base station; A step of receiving first configuration information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management from the base station; A step of receiving second setting information related to at least one reference signal from the base station; A step of receiving at least one reference signal based on the second setting information; A step of performing an operation related to the learning based on at least one reference signal, A communication device in which the above first setting information includes information related to a set of actions for reinforcement learning of the AI / ML model, the set including formats of a reference signal or report information corresponding to the reference signal as elements.

15. In a non-transitory computer-readable medium storing at least one instruction, comprising at least one instruction executable by the processor, At least one of the above commands causes the device to: Receive a first message requesting capability information from a base station, Transmitting a second message including capability information to the base station, Receive first configuration information related to learning for an AI / ML (artificial intelligence / machine learning) model for determining feedback information for beam management from the above base station, Receive second setting information related to at least one reference signal from the base station, Receive at least one reference signal based on the second setting information, Instructing the user to perform an action related to the learning based on at least one reference signal; The first setting information is a computer-readable medium including information related to a set of actions for reinforcement learning of the AI / ML model, the set including formats of a reference signal or report information corresponding to the reference signal as elements.

Citation Information

Patent Citations

  • Method and device for wireless communication

    US20230421221A1

  • Time gaps for artificial intelligence and machine learning models in wireless communication

    US20240008067A1

  • Method and device of communication in a communication system using an open radio access network

    WO2021187954A1

  • Method and apparatus for monitoring model in beam management by using artificial intelligence and machine learning

    WO2024080740A1

  • Method and device for managing model in beam management using artificial intelligence and machine learning

    WO2024091047A1