Method and apparatus for SWIN transformer-based wireless channel estimation

The Swin transformer-based channel estimation model addresses the challenge of high-dimensional channel estimation by flexibly capturing antenna and subcarrier dependencies, enhancing accuracy and reducing computational complexity in 5G and beyond systems.

WO2026005540A1PCT designated stage Publication Date: 2026-01-02SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/009141
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-03
Filing Date
2025-06-27
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing channel estimation methods struggle to accurately estimate high-dimensional wireless channels, particularly in environments with multiple antennas and subcarriers, leading to inefficiencies and suboptimal performance in 5G and beyond systems.

Method used

Employing a Swin transformer-based channel estimation model that incorporates a shifted window attention mechanism to flexibly capture dependencies in antenna and subcarrier dimensions, reducing unnecessary calculations and improving estimation accuracy.

Benefits of technology

The Swin transformer-based model enhances channel estimation performance by adapting to asymmetric dependencies, reducing computational overhead, and improving accuracy compared to conventional CNN-based models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025009141_02012026_PF_FP_ABST
    Figure KR2025009141_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure relates to a 5G or 6G communication system for supporting higher data transmission rates and satisfying various service requirements. A method performed by a base station in a wireless communication system is provided. The method includes generating channel data associated with a channel between the base station and a terminal; preprocessing the channel data for a training of a channel estimation model, the channel estimation model being based on a an Swin transformer; training the channel estimation model based on preprocessed channel data; and performing a channel estimation for the channel by using the channel estimation model, wherein an SRS received from the terminal on the channel is an input of the channel estimation model.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR SWIN TRANSFORMER-BASED WIRELESS CHANNEL ESTIMATION

[0001] The present disclosure relates to a user equipment (UE) and a base station (BS) in a wireless communication system. In particular, the disclosure related to a method and an apparatus for a shifted window (Swin) transformer-based wireless channel estimation (CE).

[0002] 5th generation (5G) mobile communication technology defines a wide frequency band to enable fast transmission speeds and new services, and can be implemented not only in the sub-6GHz frequency band such as 3.5 gigahertz (3.5GHz), but also in the ultra-high frequency band called millimeter wave (㎜Wave) such as 28GHz and 39GHz ('Above 6GHz'). In addition, in the case of 6G mobile communication technology, which is called the system after 5G communication (Beyond 5G), it is expected that it will be very important to secure new frequency resources such as the sub-6GHz band, ultra-high frequency bands, and upper mid band (7-24GHz) in order to handle the rapidly increased data traffic due to the spread of AI (artificial intelligence) technology and the increase in streaming services and to improve user perceived performance, and to efficiently utilize all available frequency resources as needed. To this end, reallocation, reuse, or sharing of existing frequency bands from 2G to 5G for 6G can be considered. Separately, since the introduction of 5G, the communications market has been increasingly interested in improving system operation efficiency, sustainability, and user experience. Accordingly, in addition to improving traditional communications performance such as data transmission speed and delay time, the introduction of new innovative technologies such as AI, reducing operating costs, improving energy efficiency, expanding service coverage, and introducing new services are becoming increasingly important.

[0003] In the early stages of 5G mobile communication technology, the goal is to support services and satisfy performance requirements for enhanced mobile broadband (eMBB), ultra-reliable low-latency communications (URLLC), and massive machine-type communications (mMTC), including beamforming and massive multiple-input multiple-output (MIMO) to mitigate path loss of radio waves in ultra-high frequency bands and increase the range of radio transmission, support for various numerologies (such as operation of multiple subcarrier intervals) and dynamic operation of slot formats for efficient use of ultra-high frequency resources, initial access technology to support multi-beam transmission and wideband, definition and operation of band-width part (BWP), new channel coding methods such as low density parity check (LDPC) codes for large-capacity data transmission and polar codes for reliable transmission of control information, L2 pre-processing, and networks that provide dedicated networks specialized for specific services. Standardization of slicing (network slicing) etc. has been progressing.

[0004] Since the early days of 5G mobile communication technology, discussions have been held on improving and enhancing the initial 5G mobile communication technology in consideration of the services that 5G mobile communication technology was intended to support, including vehicle-to-everything (V2X) to help autonomous vehicles make decisions based on their own location and status information transmitted by the vehicle and to increase user convenience, new radio unlicensed (NR-U) for system operation that meets various regulatory requirements in unlicensed bands, NR terminal low power consumption technology (i.e., UE power saving), non-terrestrial network (NTN), which is direct terminal-satellite communication to secure coverage in areas where communication with terrestrial networks is impossible, positioning, NR support up to 71GHz, support of reduced capability NR devices for lower cost and complexity compared to general terminals, UE power saving enhancement for improved power management in preparation for the use of various terminal types, and sidelink. Standardization of the physical layer has been carried out for technologies such as sidelink enhancement, duplex enhancements which study a new form of duplexing called subband non-overlapping full duplex (SBFD), network energy saving which secures the idle period in which the base station operates in maximum power saving mode to the maximum extent and reduces power consumption, and network controlled repeaters which have improved performance compared to existing repeaters by having the function of receiving and processing side control information from the network.

[0005] In addition, standardization of the wireless interface architecture / protocol layer for technologies such as the industrial internet of things (IIoT) for supporting new services through linkage and convergence with other industries, integrated access and backhaul (IAB) that provides nodes for expanding network service areas by integrating wireless backhaul links and access links, mobility enhancement technology including conditional handover (CHO) and dual active protocol stack (DAPS) handover, 2-step random access channel (RACH) for NR that simplifies random access procedures, multicast and broadcast, standardization of support for multi universal subscriber identity module (USIM) devices that provide services to users using information of two or more subscriber identity modules (SIMs), sidelink relay that provides relay-related functions to support connections between terminals in long distances and between terminals and networks, small data transfer (SDT) which transmits small data or signaling in an inactive state without transitioning to a connected state, mobility enhancements including L1 / L2 triggered mobility (LTM) / subsequent conditional PSCell addition / change (SCPAC) / and conditional handover (CHO) with candidate SCGs, extended reality support (XR enhancement) to support XR services in NR systems, etc. has also been carried out, and standardization of system architecture / services such as 5G baseline architecture (e.g., service-based architecture, service-based interface) for grafting network functions virtualization (NFV) and software-defined networking (SDN) technologies, mobile edge computing (MEC) that provides services based on the location of the terminal, non-public networks (NPN) that can be used only by some permitted terminals for non-public purposes, disaster roaming that supports the use of communication services through other carriers' networks in the event of a communication disaster, proximity-based service via 5GS, and unmanned Standardization has also been made in the system architecture / service areas, including support of UAS to support remote identification, tracking, and authorization of uncrewed aerial vehicles (UAVs), structural enhancements to support XR and interactive media services, 5GS to support AI / ML (artificial intelligence / machine learning) services, and advanced mobile edge computing to provide edge computing services in roaming networks, etc. has also been carried out.

[0006] Currently, standardization is in progress for technologies such as beam prediction using AI / ML technology, CSI (channel state information) prediction to improve positioning accuracy, ultra-low-power terminal technology using low-power wake-up receivers, technology for transmitting long term evolution (LTE) broadcasts to 5G networks, MIMO transmission technology using multiple base stations, and ultra-low-power terminals (ambient IoT) that transmit data by obtaining power from an external source without a battery. At the radio interface architecture / protocol layer, standardization is in progress for technologies such as LTM scenario support and conditional LTM support between Central Units (CUs), simultaneous support for the same XR service between multiple devices, NTN coverage enhancement and evolution, AI / ML-based mobility support, and terminal-to-terminal connection relay across multiple hops between terminals and networks. In addition, standardization of system architecture / service fields for satellite communication optimization methods, 5G system energy usage management and efficiency, SBI-based user plane evolution, ambient IoT technology, data service provision methods in IMS (IP multimedia subsystem), and avatar communication service is in progress. When such 5G mobile communication systems are commercialized, an explosive increase in connected devices will be connected to the communication network, and accordingly, it is expected that the functions and performance of 5G mobile communication systems will be strengthened and integrated operation of connected devices will be required. To this end, new research will be additionally conducted on extended reality (XR) to efficiently support augmented reality (AR), virtual reality (VR), and mixed reality (MR), 5G performance improvement and complexity reduction using AI / ML, AI service support, metaverse service support, and drone communication.

[0007] In addition, the development of these 5G mobile communication systems is expected to serve as the basis for enhancing 5G performance and ultimately evolving into 6G. In the 6G era, the three major 5G services mentioned above, eMBB, URLLC, and mMTC services, are expected to evolve into immersive communication (IC), hyper-reliable and low-latency communication (HRLLC), and massive communication (MC) services, respectively. In addition, new services such as artificial intelligence (AI) and communication, integrated sensing and communication, and ubiquitous connectivity are expected to be additionally supported. For these various 6G services, improved performance requirements compared to 5G are also essential, and standardization to define these is also in progress.

[0008] In this way, in order to satisfy the expanded services and improved performance requirements of 6G, it is expected that it will be essential to optimize and improve system operation, such as introducing AI technology, improving energy efficiency, expanding coverage, and applying next-generation security technology, as well as developing sustainable communication technology, in addition to simply improving existing communication performance.

[0009] To this end, the latest AI technology is applied to all areas from the communication system design stage to development, management, and operation to improve communication performance and realize AI internalization technology that realizes network automation and efficiency; technology that improves user-perceived performance and network operation efficiency by improving power consumption of networks and terminals; technology that reduces power consumption in core base station components such as radio frequency (RF) and modems and in the channel coding and signal modulation and transmission / reception processes; multi-antenna transmission technology (eXtreme MIMO, X-MIMO) that utilizes large antennas to overcome propagation path loss due to high frequency compared to the 3.5 GHz band of 5G communication and provide equivalent coverage; transmission / reception technology based on multiple base stations (distributed MIMO, D-MIMO) to improve quality in cell edge areas; full-duplex communication (sub-band non-overlapping full duplex, SBFD) technology to improve frequency efficiency and system network; next-generation encryption technology (post quantum cryptography, PQC) and zero trust architecture (ZTA) technology to strengthen 6G communication security; initial access delay and mobility Research will be focused on technologies to minimize delay, design a hardware-friendly protocol structure for ultra-high-speed data processing, and expand the application of integrity protection technologies.

[0010] In addition, research will be conducted on the structure of mobile communication systems (prevention of redundant functions, simplification of functions, etc.), introduction of new planes for providing service providers, user privacy protection measures, realistic services, enhancement of network resiliency, network sharing technologies, improved security technologies (false base stations, lower layer protection, etc.), and intent-based network operation and management.

[0011] In wireless communication, channel estimation processes are utilized to provide reliable transmission of data between transmitters and receivers. Wireless communication systems are inherently susceptible to various impairments and variations in the radio propagation environment, leading to fluctuations in the channel characteristics. Channel estimation seeks to mitigate the adverse effects of these variations by providing accurate information about the current state of the communication channel.

[0012] The disclosure relates to operations of a user equipment (UE), a base station (BS), and / or a network entity in a wireless communication system. More particularly, the disclosure relates to a method and an apparatus for Swin transformer-based channel estimation.

[0013] Accordingly, an aspect of the disclosure is to provide a method and apparatus for a channel estimation model based on the Swin transformer for improving channel estimation accuracy.

[0014] Furthermore, an aspect of the disclosure is to provide an enhanced attention window in the Swin transformer used for the channel estimation to capture dependencies in antenna and subcarrier dimensions more flexibly.

[0015] Additional aspects will be set forth in part in the description which follows and, in part, will be apparent from the description, or may be learned by practice of the presented embodiments.

[0016] Aspects of the disclosure are to address at least the above-mentioned problems and / or disadvantages and to provide at least the advantages described below.

[0017] In accordance with an aspect of the disclosure, a method performed by a base station in a wireless communication system is provided. The method includes generating channel data associated with a channel between the base station and a terminal; preprocessing the channel data for a training of a channel estimation model, the channel estimation model being based on an Swin transformer; training the channel estimation model based on preprocessed channel data; and performing a channel estimation for the channel by using the channel estimation model, wherein an SRS received from the terminal on the channel is an input of the channel estimation model.

[0018] In accordance with another aspect of the disclosure, a base station in a wireless communication system is provided. The base station includes a transceiver; a processor coupled to the at least one transceiver; and a memory coupled to the processor, storing instructions executable by the the processor to cause the base station to: generate channel data associated with a channel between the base station and a terminal, preprocess the channel data for a training of a channel estimation model, the channel estimation model being based on an Swin transformer, train the channel estimation model based on preprocessed channel data, and perform a channel estimation for the channel by using the channel estimation model, wherein a sounding reference signal (SRS) received from the terminal on the channel is an input of the channel estimation model.

[0019] In one embodiment, a base station (BS) is provided. The BS includes a processor configured to generate training data for a shifted window (Swin) transformer-based channel estimation (CE) model, preprocess the training data, and train the Swin transformer-based CE model with the preprocessed training data. The BS also includes a transceiver operably coupled to the transceiver. The transceiver is configured to receive, over a wireless communication channel, a sounding reference signal (SRS). The processor is also configured to provide the SRS as an input image to the trained Swin transformer-based CE model, and receive as output from the trained Swin transformer-based CE model, a CE for the wireless communication channel.

[0020] In another embodiment, a method of operating a BS is provided. The method includes generating training data for a Swin transformer-based CE model, preprocessing the training data, and receiving, over a wireless communication channel, a SRS. The method also includes providing the SRS as an input image to the trained Swin transformer-based CE model, and receiving as output from the trained Swin transformer-based CE model, a CE for the wireless communication channel.

[0021] In yet another embodiment, a non-transitory computer readable medium embodying a computer program is provided. The computer program includes program code that, when executed by a processor of a device, causes the device to generate training data for a Swin transformer-based CE model, preprocess the training data, and train the Swin transformer-based CE model with the preprocessed training data. The computer program also includes program code that, when executed by the processor of a device, causes the device to receive, over a wireless communication channel, a SRS, provide the SRS as an input image to the trained Swin transformer-based CE model, and receive as output from the trained Swin transformer-based CE model, a CE for the wireless communication channel.

[0022] Other technical features may be readily apparent to one skilled in the art from the following figures, descriptions, and claims.

[0023] According to an embodiment of the disclosure, by applying the Swin transformer to channel estimation, channel estimation performance can be improved compared to existing CNN-based models or simple transformer models.

[0024] Furthermore, according to an embodiment of the disclosure, by flexibly designing the size of the attention window, unnecessary calculations can be reduced while learning considering the asymmetric dependency of antenna dimensions and subcarrier dimensions can be performed.

[0025] The effects obtainable in the disclosure are not limited to the above-mentioned effects, and other effects not mentioned herein will be clearly understood from the following description by those skilled in the art to which the disclosure belongs.

[0026] For a more complete understanding of this disclosure and its advantages, reference is now made to the following description, taken in conjunction with the accompanying drawings, in which:

[0027] Figure 1 illustrates an example wireless network according to embodiments of the present disclosure;

[0028] Figure 2a illustrates example wireless transmit and receive paths according to embodiments of the present disclosure;

[0029] Figure 2b illustrates example wireless transmit and receive paths according to embodiments of the present disclosure;

[0030] Figure 3a illustrates an example UE according to embodiments of the present disclosure;

[0031] Figure 3b illustrates an example gNB according to embodiments of the present disclosure;

[0032] Figure 4 illustrates an example procedure for Swin transformer-based channel estimation according to embodiments of the present disclosure;

[0033] Figure 5 illustrates an example method for preprocessing channel data according to embodiments of the present disclosure;

[0034] Figure 6 illustrates an example architecture of a Swin transformer-based channel estimation model according to embodiments of the present disclosure;

[0035] Figure 7 illustrates an example Swin transformer layer (STL) structure according to embodiments of the present disclosure;

[0036] Figure 8 illustrates an example of shifted window partitioning according to embodiments of the present disclosure;

[0037] Figure 9 illustrates an example limitation of square attention windows according to embodiments of the present disclosure; and

[0038] Figure 10 illustrates an example method for Swin transformer-based wireless CE according to embodiments of the present disclosure.

[0039] Before undertaking the Mode for Invention, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The term "couple" and its derivatives refer to any direct or indirect communication between two or more elements, whether or not those elements are in physical contact with one another. The terms "transmit," "receive," and "communicate," as well as derivatives thereof, encompass both direct and indirect communication. The terms "include" and "comprise," as well as derivatives thereof, mean inclusion without limitation. The term "or" is inclusive, meaning and / or. The phrase "associated with," as well as derivatives thereof, means to include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, juxtapose, be proximate to, be bound to or with, have, have a property of, have a relationship to or with, or the like. The term "controller" means any device, system or part thereof that controls at least one operation. Such a controller may be implemented in hardware or a combination of hardware and software and / or firmware. The functionality associated with any particular controller may be centralized or distributed, whether locally or remotely. The phrase "at least one of," when used with a list of items, means that different combinations of one or more of the listed items may be used, and only one item in the list may be needed. For example, "at least one of: A, B, and C" includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A and B and C.

[0040] Moreover, various functions described below can be implemented or supported by one or more computer programs, each of which is formed from computer readable program code and embodied in a computer readable medium. The terms "application" and "program" refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or a portion thereof adapted for implementation in a suitable computer readable program code. The phrase "computer readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer readable medium" includes any type of medium capable of being accessed by a computer, such as read only memory (ROM), random access memory (RAM), a hard disk drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory. A "non-transitory" computer readable medium excludes wired, wireless, optical, or other communication links that transport transitory electrical or other signals. A non-transitory computer readable medium includes media where data can be permanently stored and media where data can be stored and later overwritten, such as a rewritable optical disc or an erasable memory device.

[0041] Definitions for other certain words and phrases are provided throughout this patent document. Those of ordinary skill in the art should understand that in many if not most instances, such definitions apply to prior as well as future uses of such defined words and phrases.

[0042] Figures 1 through 10, discussed below, and the various embodiments used to describe the principles of this disclosure in this patent document are by way of illustration only and should not be construed in any way to limit the scope of the disclosure. Those skilled in the art will understand that the principles of this disclosure may be implemented in any suitably arranged wireless communication system.

[0043] To meet the demand for wireless data traffic having increased since deployment of 4G communication systems and to enable various vertical applications, 5G / NR communication systems have been developed and are currently being deployed. The 5G / NR communication system is considered to be implemented in higher frequency (mmWave) bands, e.g., 28 GHz or 60GHz bands, so as to accomplish higher data rates or in lower frequency bands, such as 6 GHz, to enable robust coverage and mobility support. To decrease propagation loss of the radio waves and increase the transmission distance, the beamforming, massive multiple-input multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, an analog beam forming, large scale antenna techniques are discussed in 5G / NR communication systems.

[0044] In addition, in 5G / NR communication systems, development for system network improvement is under way based on advanced small cells, cloud radio access networks (RANs), ultra-dense networks, device-to-device (D2D) communication, wireless backhaul, moving network, cooperative communication, coordinated multi-points (CoMP), reception-end interference cancelation and the like.

[0045] The discussion of 5G systems and frequency bands associated therewith is for reference as certain embodiments of the present disclosure may be implemented in 5G systems. However, the present disclosure is not limited to 5G systems or the frequency bands associated therewith, and embodiments of the present disclosure may be utilized in connection with any frequency band. For example, aspects of the present disclosure may also be applied to deployment of 5G communication systems, 6G or even later releases which may use terahertz (THz) bands.

[0046] Figures 1-3b below describe various embodiments implemented in wireless communications systems and with the use of orthogonal frequency division multiplexing (OFDM) or orthogonal frequency division multiple access (OFDMA) communication techniques. The descriptions of Figures 1-3b are not meant to imply physical or architectural limitations to the manner in which different embodiments may be implemented. Different embodiments of the present disclosure may be implemented in any suitably arranged communications system.

[0047] Figure 1illustrates an example wireless network 100 according to embodiments of the present disclosure. The embodiment of the wireless network shown in Figure 1 is for illustration only. Other embodiments of the wireless network 100 could be used without departing from the scope of this disclosure.

[0048] As shown in Figure 1, the wireless network includes a gNB 101 (e.g., base station, BS), a gNB 102, and a gNB 103. The gNB 101 communicates with the gNB 102 and the gNB 103. The gNB 101 also communicates with at least one network 130, such as the internet, a proprietary internet protocol (IP) network, or other data network.

[0049] The gNB 102 provides wireless broadband access to the network 130 for a first plurality of user equipments (UEs) within a coverage area 120 of the gNB 102. The first plurality of UEs includes a UE 111, which may be located in a small business; a UE 112, which may be located in an enterprise; a UE 113, which may be a Wi-Fi hotspot; a UE 114, which may be located in a first residence; a UE 115, which may be located in a second residence; and a UE 116, which may be a mobile device, such as a cell phone, a wireless laptop, a wireless PDA, or the like. The gNB 103 provides wireless broadband access to the network 130 for a second plurality of UEs within a coverage area 125 of the gNB 103. The second plurality of UEs includes the UE 115 and the UE 116. In some embodiments, one or more of the gNBs 101-103 may communicate with each other and with the UEs 111-116 using 5G / NR, long term evolution (LTE), long term evolution-advanced (LTE-A), WiMAX, Wi-Fi, or other wireless communication techniques.

[0050] Depending on the network type, the term "base station" or "BS" can refer to any component (or collection of components) configured to provide wireless access to a network, such as transmit point (TP), transmit-receive point (TRP), an enhanced base station (eNodeB or eNB), a 5G / NR base station (gNB), a macrocell, a femtocell, a Wi-Fi access point (AP), or other wirelessly enabled devices. Base stations may provide wireless access in accordance with one or more wireless communication protocols, e.g., 5G / NR 3rdgeneration partnership project (3GPP) NR, long term evolution (LTE), LTE advanced (LTE-A), high speed packet access (HSPA), Wi-Fi 802.11a / b / g / n / ac, etc. For the sake of convenience, the terms "BS" and "TRP" are used interchangeably in this patent document to refer to network infrastructure components that provide wireless access to remote terminals. Also, depending on the network type, the term "user equipment" or "UE" can refer to any component such as "mobile station," "subscriber station," "remote terminal," "wireless terminal," "receive point," or "user device." For the sake of convenience, the terms "user equipment" and "UE" are used in this patent document to refer to remote wireless equipment that wirelessly accesses a BS, whether the UE is a mobile device (such as a mobile telephone or smartphone) or is normally considered a stationary device (such as a desktop computer or vending machine).

[0051] Dotted lines show the approximate extents of the coverage areas 120 and 125, which are shown as approximately circular for the purposes of illustration and explanation only. It should be clearly understood that the coverage areas associated with gNBs, such as the coverage areas 120 and 125, may have other shapes, including irregular shapes, depending upon the configuration of the gNBs and variations in the radio environment associated with natural and man-made obstructions.

[0052] As described in more detail below, one or more of the UEs 111-116 include circuitry, programing, or a combination thereof, for shifted window (Swin) transformer-based wireless channel estimation (CE). In certain embodiments, one or more of the gNBs 101-103 includes circuitry, programing, or a combination thereof, to support Swin transformer-based wireless CE in a wireless communication system.

[0053] Although Figure 1 illustrates one example of a wireless network, various changes may be made to Figure 1. For example, the wireless network could include any number of gNBs and any number of UEs in any suitable arrangement. Also, the gNB 101 could communicate directly with any number of UEs and provide those UEs with wireless broadband access to the network 130. Similarly, each gNB 102-103 could communicate directly with the network 130 and provide UEs with direct wireless broadband access to the network 130. Further, the gNBs 101, 102, and / or 103 could provide access to other or additional external networks, such as external telephone networks or other types of data networks.

[0054] Figures 2a and 2billustrate example wireless transmit and receive paths according to embodiments of the present disclosure.

[0055] In the following description, a transmit path 200 may be described as being implemented in a gNB (such as gNB 102), while a receive path 250 may be described as being implemented in a UE (such as UE 116). However, it will be understood that the receive path 250 can be implemented in a gNB and that the transmit path 200 can be implemented in a UE. In some embodiments, the transmit path 200 and / or the receive path 250 is configured to implement and / or support Swin transformer-based wireless CE as described in embodiments of the present disclosure.

[0056] The transmit path 200 includes a channel coding and modulation block 205, a serial-to-parallel (S-to-P) block 210, a size N inverse fast fourier transform (IFFT) block 215, a parallel-to-serial (P-to-S) block 220, an add cyclic prefix block 225, and an up-converter (UC) 230. The receive path 250 includes a down-converter (DC) 255, a remove cyclic prefix block 260, a serial-to-parallel (S-to-P) block 265, a size N fast fourier transform (FFT) block 270, a parallel-to-serial (P-to-S) block 275, and a channel decoding and demodulation block 280.

[0057] In the transmit path 200, the channel coding and modulation block 205 receives a set of information bits, applies coding (such as a low-density parity check (LDPC) coding), and modulates the input bits (such as with quadrature phase shift keying (QPSK) or auadrature amplitude modulation (QAM)) to generate a sequence of frequency-domain modulation symbols. The serial-to-parallel block 210 converts (such as de-multiplexes) the serial modulated symbols to parallel data in order to generate N parallel symbol streams, where N is the IFFT / FFT size used in the gNB 102 and the UE 116. The size N IFFT block 215 performs an IFFT operation on the N parallel symbol streams to generate time-domain output signals. The parallel-to-serial block 220 converts (such as multiplexes) the parallel time-domain output symbols from the size N IFFT block 215 in order to generate a serial time-domain signal. The add cyclic prefix block 225 inserts a cyclic prefix to the time-domain signal. The up-converter 230 modulates (such as up-converts) the output of the add cyclic prefix block 225 to an RF frequency for transmission via a wireless channel. The signal may also be filtered at baseband before conversion to the RF frequency.

[0058] A transmitted RF signal from the gNB 102 arrives at the UE 116 after passing through the wireless channel, and reverse operations to those at the gNB 102 are performed at the UE 116. The down-converter 255 down-converts the received signal to a baseband frequency, and the remove cyclic prefix block 260 removes the cyclic prefix to generate a serial time-domain baseband signal. The serial-to-parallel block 265 converts the time-domain baseband signal to parallel time domain signals. The size N FFT block 270 performs an FFT algorithm to generate N parallel frequency-domain signals. The parallel-to-serial block 275 converts the parallel frequency-domain signals to a sequence of modulated data symbols. The channel decoding and demodulation block 280 demodulates and decodes the modulated symbols to recover the original input data stream.

[0059] Each of the gNBs 101-103 may implement a transmit path 200 that is analogous to transmitting in the downlink to UEs 111-116 and may implement a receive path 250 that is analogous to receiving in the uplink from UEs 111-116. Similarly, each of UEs 111-116 may implement a transmit path 200 for transmitting in the uplink to gNBs 101-103 and may implement a receive path 250 for receiving in the downlink from gNBs 101-103.

[0060] Each of the components in Figures 2a and 2b can be implemented using only hardware or using a combination of hardware and software / firmware. As a particular example, at least some of the components in Figures 2a and 2b may be implemented in software, while other components may be implemented by configurable hardware or a mixture of software and configurable hardware. For instance, the FFT block 270 and the IFFT block 215 may be implemented as configurable software algorithms, where the value of size N may be modified according to the implementation.

[0061] Furthermore, although described as using FFT and IFFT, this is by way of illustration only and should not be construed to limit the scope of this disclosure. Other types of transforms, such as discrete fourier transform (DFT) and inverse discrete fourier transform (IDFT) functions, can be used. It will be appreciated that the value of the variable N may be any integer number (such as 1, 2, 3, 4, or the like) for DFT and IDFT functions, while the value of the variable N may be any integer number that is a power of two (such as 1, 2, 4, 8, 16, or the like) for FFT and IFFT functions.

[0062] Although Figures 2a and 2b illustrate examples of wireless transmit and receive paths, various changes may be made to Figures 2a and 2b. For example, various components in Figures 2a and 2b can be combined, further subdivided, or omitted and additional components can be added according to particular needs. Also, Figures 2a and 2b are meant to illustrate examples of the types of transmit and receive paths that can be used in a wireless network. Any other suitable architectures can be used to support wireless communications in a wireless network.

[0063] Figure 3aillustrates an example UE 116 according to embodiments of the present disclosure. The embodiment of the UE 116 illustrated in Figure 3a is for illustration only, and the UEs 111-115 of Figure 1 could have the same or similar configuration. However, UEs come in a wide variety of configurations, and Figure 3a does not limit the scope of this disclosure to any particular implementation of a UE.

[0064] As shown in Figure 3a, the UE 116 includes antenna(s) 305, a transceiver(s) 310, and a microphone 320. The UE 116 also includes a speaker 330, a processor 340, an input / output (I / O) interface (IF) 345, an input 350, a display 355, and a memory 360. The memory 360 includes an operating system (OS) 361 and one or more applications 362.

[0065] The transceiver(s) 310 receives, from the antenna 305, an incoming RF signal transmitted by a gNB of the network 100. The transceiver(s) 310 down-converts the incoming RF signal to generate an intermediate frequency (IF) or baseband signal. The IF or baseband signal is processed by RX processing circuitry in the transceiver(s) 310 and / or processor 340, which generates a processed baseband signal by filtering, decoding, and / or digitizing the baseband or IF signal. The RX processing circuitry sends the processed baseband signal to the speaker 330 (such as for voice data) or is processed by the processor 340 (such as for web browsing data).

[0066] TX processing circuitry in the transceiver(s) 310 and / or processor 340 receives analog or digital voice data from the microphone 320 or other outgoing baseband data (such as web data, e-mail, or interactive video game data) from the processor 340. The TX processing circuitry encodes, multiplexes, and / or digitizes the outgoing baseband data to generate a processed baseband or IF signal. The transceiver(s) 310 up-converts the baseband or IF signal to an RF signal that is transmitted via the antenna(s) 305.

[0067] The processor 340 can include one or more processors or other processing devices and execute the OS 361 stored in the memory 360 in order to control the overall operation of the UE 116. For example, the processor 340 could control the reception of DL channel signals and the transmission of UL channel signals by the transceiver(s) 310 in accordance with well-known principles. In some embodiments, the processor 340 includes at least one microprocessor or microcontroller.

[0068] The processor 340 is also capable of executing other processes and programs resident in the memory 360, for example, processes for Swin transformer-based wireless CE as discussed in greater detail below. The processor 340 can move data into or out of the memory 360 as required by an executing process. In some embodiments, the processor 340 is configured to execute the applications 362 based on the OS 361 or in response to signals received from gNBs or an operator. The processor 340 is also coupled to the I / O interface 345, which provides the UE 116 with the ability to connect to other devices, such as laptop computers and handheld computers. The I / O interface 345 is the communication path between these accessories and the processor 340.

[0069] The processor 340 is also coupled to the input 350, which includes for example, a touchscreen, keypad, etc., and the display 355. The operator of the UE 116 can use the input 350 to enter data into the UE 116. The display 355 may be a liquid crystal display, light emitting diode display, or other display capable of rendering text and / or at least limited graphics, such as from web sites.

[0070] The memory 360 is coupled to the processor 340. Part of the memory 360 could include a random-access memory (RAM), and another part of the memory 360 could include a Flash memory or other read-only memory (ROM).

[0071] Although Figure 3a illustrates one example of UE 116, various changes may be made to Figure 3a. For example, various components in Figure 3a could be combined, further subdivided, or omitted and additional components could be added according to particular needs. As a particular example, the processor 340 could be divided into multiple processors, such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs). In another example, the transceiver(s) 310 may include any number of transceivers and signal processing chains and may be connected to any number of antennas. Also, while Figure 3a illustrates the UE 116 configured as a mobile telephone or smartphone, UEs could be configured to operate as other types of mobile or stationary devices.

[0072] Figure 3billustrates an example gNB 102 according to embodiments of the present disclosure. The embodiment of the gNB 102 illustrated in Figure 3b is for illustration only, and the gNBs 101 and 103 of Figure 1 could have the same or similar configuration. However, gNBs come in a wide variety of configurations, and Figure 3B does not limit the scope of this disclosure to any particular implementation of a gNB.

[0073] As shown in Figure 3b, the gNB 102 includes multiple antennas 370a-370n, multiple transceivers 372a-372n, a controller / processor 378, a memory 380, and a backhaul or network interface 382.

[0074] The transceivers 372a-372n receive, from the antennas 370a-370n, incoming RF signals, such as signals transmitted by UEs in the network 100. The transceivers 372a-372n down-convert the incoming RF signals to generate IF or baseband signals. The IF or baseband signals are processed by receive (RX) processing circuitry in the transceivers 372a-372n and / or controller / processor 378, which generates processed baseband signals by filtering, decoding, and / or digitizing the baseband or IF signals. The controller / processor 378 may further process the baseband signals.

[0075] Transmit (TX) processing circuitry in the transceivers 372a-372n and / or controller / processor 378 receives analog or digital data (such as voice data, web data, e-mail, or interactive video game data) from the controller / processor 378. The TX processing circuitry encodes, multiplexes, and / or digitizes the outgoing baseband data to generate processed baseband or IF signals. The transceivers 372a-372n up-converts the baseband or IF signals to RF signals that are transmitted via the antennas 370a-370n.

[0076] The controller / processor 378 can include one or more processors or other processing devices that control the overall operation of the gNB 102. For example, the controller / processor 378 could control the reception of uplink (UL) channel signals and the transmission of downlink (DL) channel signals by the transceivers 372a-372n in accordance with well-known principles. The controller / processor 378 could support additional functions as well, such as more advanced wireless communication functions. For instance, the controller / processor 378 could support beam forming or directional routing operations in which outgoing / incoming signals from / to multiple antennas 370a-370n are weighted differently to effectively steer the outgoing signals in a desired direction. Any of a wide variety of other functions could be supported in the gNB 102 by the controller / processor 378.

[0077] The controller / processor 378 is also capable of executing programs and other processes resident in the memory 380, such as an OS and, for example, processes to support Swin transformer-based wireless CE as discussed in greater detail below. The controller / processor 378 can move data into or out of the memory 380 as required by an executing process.

[0078] The controller / processor 378 is also coupled to the backhaul or network interface 382. The backhaul or network interface 382 allows the gNB 102 to communicate with other devices or systems over a backhaul connection or over a network. The interface 382 could support communications over any suitable wired or wireless connection(s). For example, when the gNB 102 is implemented as part of a cellular communication system (such as one supporting 5G / NR, LTE, or LTE-A), the interface 382 could allow the gNB 102 to communicate with other gNBs over a wired or wireless backhaul connection. When the gNB 102 is implemented as an access point, the interface 382 could allow the gNB 102 to communicate over a wired or wireless local area network or over a wired or wireless connection to a larger network (such as the Internet). The interface 382 includes any suitable structure supporting communications over a wired or wireless connection, such as an Ethernet or transceiver.

[0079] The memory 380 is coupled to the controller / processor 378. Part of the memory 380 could include a RAM, and another part of the memory 380 could include a Flash memory or other ROM.

[0080] Although Figure 3b illustrates one example of gNB 102, various changes may be made to Figure 3b. For example, the gNB 102 could include any number of each component shown in Figure 3b. Also, various components in Figure 3b could be combined, further subdivided, or omitted and additional components could be added according to particular needs.

[0081] A wireless channel is a dynamic medium through which signals are transmitted. Wireless channels can be affected by factors such as multi-path fading, interference, noise, and mobility. Channel estimation serves as a mechanism to track and adapt to these dynamic changes, allowing the communication system to optimize its performance. Channel estimation is a process of estimating the characteristics of a wireless communication channel, such as its frequency response, delay spread, and fading coefficients, which are used by the receiver to demodulate and decode transmitted signals accurately. Channel estimation is important for coherent detection and decoding of transmitted signals, as well as for optimization of transmission parameters, such as power allocation, modulation scheme, and coding rate. Channel estimation can improve the accuracy and reliability of received signals, as well as increase the capacity and performance of wireless communication systems.

[0082] Various methods for channel estimation often utilize pilot signals, which are known symbols inserted into the transmitted signal, allowing the receiver to measure the channel response at specific points in time. These measurements are then used to interpolate the channel characteristics between pilot symbols, thus providing an estimate of the channel conditions. Various methods and techniques have been proposed to tackle this problem, such as linear interpolation, least squares, minimum mean square error, maximum likelihood, Bayesian interference, and deep learning.

[0083] However, channel estimation is also challenging, especially for high-dimensional signals that involve multiple antennas, multiple subcarriers and multiple users. The channel estimation problem can be formulated as finding the optimal solution of a system of equations that relate the transmitted and received signals with the channel coefficients and the noise. The complexity and difficulty of this problem depends on the number and arrangement of the channel coefficients, the availability and quality of the pilot signals, the noise level and distribution, and the channel dynamics and variations. Some channel estimation solutions such as least square (LS) and linear minimum mean square error (LMMSE) fail to achieve desirable estimation accuracy with reasonable complexity, particularly in the low signal-to-noise ratio (SNR) regime.

[0084] Recently, the integration of machine learning techniques into channel estimation processes has gained substantial attention and shown great promise in improving the accuracy and efficiency of channel estimation. Machine learning-based channel estimation leverages the power of artificial intelligence and data-driven approaches to adapt and learn from the wireless channel's behavior, making channel estimation more robust to varying conditions and potentially reducing the need for explicit pilot signals.

[0085] Some common machine learning-based channel estimation methods include deep learning approaches, reinforcement learning, autoencoders, transfer learning, and diffusion models.

[0086] - Deep learning approaches utilizing deep neural networks (DNNs) (including convolutional neural networks [CNNs] and recurrent neural networks [RNNs]) have been applied to channel estimation tasks. These networks can learn complex relationships between received signals and the channel characteristics, allowing for accurate and efficient estimation.

[0087] - Reinforcement learning techniques can be used to optimize the transmission and reception strategies in response to changing channel conditions, effectively improving channel estimation and overall system performance.

[0088] - Autoencoders are neural network architectures that can be used for unsupervised learning of channel representations. Autoencoders can capture channel characteristics and reduce the reliance on pilot signals.

[0089] - Transfer learning techniques enable the adaptation of pre-trained models to specific channel environments, enhancing the generalization of channel estimation algorithms across different scenarios.

[0090] - Diffusion models include denoising diffusion probabilistic models (DDPM) and score matching with Langevin dynamics (SMLD). Based on the SMLD algorithm, some channel estimation solutions first learn a score function of the channel data using denoising score matching, obtain a close-form score function of the likelihood, and finally complete a posterior sampling process following annealed Langevin dynamics.

[0091] Machine learning-based channel estimation methods hold the potential to make wireless communication systems more adaptive, efficient, and robust, particularly in challenging environments. As the field of machine learning continues to advance, these methods are expected to play an increasingly important role in optimizing wireless communication systems for a wide range of applications, including 5G, IoT, and beyond. Shifted window (Swin) transformer approaches have shown great promise in machine learning-based channel estimation, as Swin transformer approaches integrate the advantages of both CNN and Transformer based approaches. Swin transformer approaches have the advantage of CNN based approaches to process images of large size due to the local attention mechanism. Swin transformer approaches also have the advantage of Transformer based approaches to model long-range dependency with the shifted window scheme.

[0092] Various embodiments of the present disclosure apply a Swin transformer-based image restoration algorithm to the channel estimation task.

[0093] In the present disclosure, the channel estimation problem may be solved as follows:

[0094] In the frequency domain, the input-output relationship at pilot tones (subcarriers) between transmitted and received signals can be expressed as

[0095]

[0096] In particular, the mathematical model described in equation (1) is applicable to different types of signal models (including but not limited to single-input single-output (SISO), single-input multiple-output (SIMO), multiple-input multiple-output (MIMO) cases etc.). For example, in a SIMO signal model, and can be used to represent the number of the pilot tones (subcarriers) in the frequency domain over one OFDM symbol and the number of the receive antennas, respectively. On the other hand, in a SISO case, and can be used to represent the number of the pilot tones (subcarriers) in the frequency domain over one OFDM symbol and the number of the OFDM symbols containing pilot tones, respectively. Note that MIMO signal models can be readily converted to a SIMO case where pilot signals from different transmitted antennas are separated in time, frequency, or code domains.

[0097] The goal of the channel estimation task is to estimate the channel matrixHbased on pilot signalsXand received signalsY. Without loss of generality, the present disclosure assumes pilot signalsXto be an identity matrix, and thus the signal model in equation (1) can be rewritten as

[0098] .   (2)

[0099] Note that the various embodiments of the present disclosure can be readily applied to the case where pilot signalsXare not an identity matrix.

[0100] Figure 4illustrates an example procedure 400 for Swin transformer-based channel estimation according to embodiments of the present disclosure. An embodiment of the procedure illustrated in Figure 4 is for illustration only. One or more of the components illustrated in Figure 4 may be implemented in specialized circuitry configured to perform the noted functions or one or more of the components may be implemented by one or more processors executing instructions to perform the noted functions. Other embodiments of a procedure for Swin transformer-based channel estimation could be used without departing from the scope of this disclosure.

[0101] In the example of Figure 4, the procedure 400 for Swin transformer-based channel estimation includes four steps. In some embodiments, the four steps of procedure 400 may be performed by a network entity. For example, procedure 400 may be performed by a BS (such as gNB 102 of Figure 1). However, it should be understood that in some embodiments, each step of procedure 400 can be performed by different apparatuses, such as other network entities, simulators, user equipment, etc.

[0102] Procedure 400 begins at step 401.At step 401, wireless channel data is generated to be used for model training and testing of a Swin transformer-based CE model. As described herein, channel data generation refers to a process to obtain channel response data or received signal data. During the channel data generation process, additional information about the channel or received signals, e.g., signal to noise ratio (SNR) or transmission power, should also be estimated and stored. In some embodiments, the channel data may be generated based on a simulated channel. For example, channel data generation may be performed by channel simulation based on wireless channel models. In some embodiments, channel data generation may be performed in the field. For example, the channel data may be generated by a UE (such as UE 116 of Figure 1) being served by a network entity (such as gNB 102 of Figure 1).

[0103]

[0104] By transforming the channel response to other domains, such as a delay-antenna domain or a delay-angular domain, the channel data could become sparser. Increased sparsity of data can bring some benefits to the Swin transformer-based CE model, such as:

[0105] - Regularization effect: Sparse data acts as a natural form of regularization. When the available data is limited, models tend to generalize better because they focus on essential patterns rather than memorizing noise.

[0106] - Feature importance: Sparse data highlights the importance of features. Rare but informative features receive more attention from the model.

[0107] - Efficient storage and processing: Sparse representations require less memory and computational resources, making sparse representations efficient for large-scale applications.

[0108] At step 402, the channel data generated at step 401 is preprocessed so that the channel data has a sparser structure. In some embodiments, the channel data may be preprocessed as shown in Figure 5.

[0109] Figure 5illustrates an example method 500 for preprocessing channel data according to embodiments of the present disclosure. An embodiment of the method illustrated in Figure 5 is for illustration only. One or more of the components illustrated in Figure 5 may be implemented in specialized circuitry configured to perform the noted functions or one or more of the components may be implemented by one or more processors executing instructions to perform the noted functions. Other embodiments of a method for preprocessing channel data could be used without departing from the scope of this disclosure.

[0110] In the example of Figure 5, method 500 begins at step 421. At step 421, the channel data generated at step 401 is transformed from the frequency-antenna domain to a delay-antenna domain using an Inverse Fast Fourier Transform (IFFT).

[0111] At step 422, the data transformed at step 421 is transformed from the delay-antenna domain to a delay-angular domain using a 2-dimensional Fast Fourier Transform (2D FFT) based on the structure of the antenna. The channel response on the transformed domain can then be used as the input and output of the Swin transformer-based CE model.

[0112] Although Figure 5 illustrates one example method 500 for preprocessing channel data, various changes may be made to Figure 5. For example, while shown as a series of steps, various steps in Figure 5 could overlap, occur in parallel, occur in a different order, occur any number of times, be omitted, or replaced by other steps.

[0113] At step 403, the Swin transformer-based CE model is trained with the preprocessed data set (images) resulting from the preprocessing of the channel data at step 402. The Swin transformer-based CE model takes a noisy channel response image as input and outputs an estimated channel response image. In some embodiments, in correspondence with the data preprocessing at step 402, the estimation result (i.e., the output of the Swin transformer-based CE model) may be converted back to the frequency-antenna domain. In some embodiments, the architecture of the Swin transformer-based CE model may be as shown in Figure 6.

[0114] Figure 6illustrates an example architecture 600 of a Swin transformer-based channel estimation model according to embodiments of the present disclosure. The embodiment of an architecture of a Swin transformer-based CE model of Figure 6 is for illustration only. Different embodiments of an architecture of a Swin transformer-based CE could be used without departing from the scope of this disclosure.

[0115] In the example of Figure 6, architecture 600 includes three components:

[0116] - Shallow Feature Extraction (431): This initial step extracts low-level features from the input noisy channel response image. In the example of Figure 6, a convolutional layer is used as the shallow feature extraction module. However, convolutional layers of other sizes may be used for shallow feature extraction, and embodiments of the present disclosure are not limited to using a convolutional layer for shallow feature extraction.

[0117] - Deep Feature Extraction (432): This module comprises several residual Swin Transformer blocks (RSTBs). Each RSTB combines multiple Swin Transformer layers (STLs) with a residual connection. The STLs capture long-range dependencies across antennas and subcarriers, while the residual connection is for stable training. In the example of Figure 6, a single RSTB including 7 STLs is used as the deep feature extraction module. However, additional RSTBs may be used for deep feature extraction, and any number of STLs may be used within an RSTB. Embodiments of the present disclosure are not limited to using a single RSTB or RSTBs with 7 STLs for deep feature extraction.

[0118] - High-Quality Image Reconstruction (433): This step reconstructs the high-quality image (i.e., the true channel response) using the extracted features. In the example of Figure 6, a convolutional layer is used as the high-quality image reconstruction module. However, convolutional layers of other sizes may be used for high-quality image reconstruction, and embodiments of the present disclosure are not limited to using a convolutional layer for high-quality image reconstruction.

[0119] The hyper parameters for the Swin transformer-based CE model of Figure 6 include the number of STLs in one RSTB, the number of RSTBs in deep feature extraction 432, and the network structure used for shallow feature extraction 431 and high-quality image reconstruction 433. In some embodiments, the hyper parameters may be obtained by experiments on the dataset generated at step 402. The hyper parameters can be tunable based on the size of the dataset, computation complexity constraints, memory constraints, etc. However, these hyper parameters are only illustrative, and do not restrict the various embodiments of the present disclosure.

[0120] Although Figure 6 illustrates one example architecture 600 of a Swin transformer-based channel estimation model, various changes may be made to Figure 6. For example, various changes to the number of STLs could be made, etc. according to particular needs.

[0121] Figure 7illustrates an example STL structure 700 according to embodiments of the present disclosure. The embodiment of an STL structure of Figure 7 is for illustration only. Different embodiments of an STL structure could be used without departing from the scope of this disclosure.

[0122] In the example of Figure 7, a structure of two successive STLs is shown. For example, the two successive STLs of Figure 7 may represent any two of the STLs shown in the RTSB for deep feature extraction 432 in Figure 6.

[0123] STLs are the core of Swin transformer-based CE models as described herein. STLs are based on the standard multi-head self-attention (MSA) mechanism (steps 442 and 446 in Figure 7) of the original transformer layer. The primary distinctions are the incorporation of the local attention and the shift-invariant mechanism. Assuming input of dimensions , an STL as shown in Figure 7 initially partitions the input image into non-overlapping local windows of shape , reshaping the input image into a feature. Here is the number of local windows. Subsequently, the STL computes the standard self-attention independently within each window, referred to as local attention.

[0124]

[0125] One of the advantages of the local window attention mechanism is the local window attention mechanism can enable efficient attention calculation. For example, for a feature with shape , where and are the height and width of the feature and is the embedding dimension, the computational complexity of a global MSA module and a window based MSA are

[0126]

[0127] The complexity of the global MSA is quadratic to the feature size , while the windowed complexity is linear when is fixed. Global self-attention computation is generally unaffordable for a large . On the other hand, the window based self-attention is scalable. As long as an appropriate window size is found, the complexity of the windowed MSA can be much lower than the global MSA.

[0128] As shown in Figure 7, first a LayerNorm (LN) layer is applied (steps 441 and 445) to the input data. Then, in steps 442 or 446, the attention function is performed for times in parallel and the results are concatenated for multihead self-attention (MSA), where is the number of attention heads. Then, a multi-layered perceptron (MLP), comprising two fully connected layers with Gaussian error linear unit (GELU) non-linearities in between, is used to transform features further in steps 444 and 448. Before both MSA and MLP, a LN layer (443 and 447) is applied, and residual connections are utilized for both components. However, the self-attention module using window lacks inter-window connections, limiting the self-attention module's ability to model the global dependency. For introducing connections across windows while retaining efficient computations in non-overlapping windows, a shifted window partitioning method is used that alternates between two configuration settings in successive Swin Transformer layers. For example, a shifted window partitioning is applied in the second STL, as shown in step 446 in Figure 7.

[0129] Although Figure 7 illustrates one example STL structure 700, various changes may be made to Figure 7. For example, various changes to number of successive STLs could be made, etc. according to particular needs.

[0130] As described herein, shifted window partitioning means shifting the feature by a stride of pixels before partitioning. In some embodiments, there are two steps of shifted window partitioning, as shown in Figure 8.

[0131] Figure 8illustrates an example of shifted window partitioning 800 according to embodiments of the present disclosure. The embodiment of shifted window partitioning of Figure 8 is for illustration only. Different embodiments of shifted window partitioning could be used without departing from the scope of this disclosure.

[0132] In the example of Figure 8, the shifted window partitioning 800 begins at step 451. At step 451, a feature with a feature size , where is the window size, is cyclically shifted by pixels to the right bottom. Then, in step 452, the cyclically shifted feature is partitioned to 4 non-overlapping windows with shape . After MSA is performed (e.g., at step 446), the feature is shifted back using a cyclic shift.

[0133] While the example of Figure 8 uses a feature size , embodiments of the present disclosure are not limited to features of feature size , and the steps of shifted window partitioning 800 can be applied to other feature sizes.

[0134] Although Figure 8 illustrates one example of shifted window partitioning 800, various changes may be made to Figure 8. For example, various changes to window size could be made, etc. according to particular needs.

[0135] With the shifted window partitioning approach, the consecutive Swin Transformer blocks of Figure 7 are blocks are computed as

[0136]

[0137] At step 404, the trained Swin transformer-based CE model is applied to estimate a channel. For example, an SRS may be used as an input image for the trained Swin transformer-based CE model, and the trained Swin transformer-based CE model may output a channel estimation based on the input SRS.

[0138] In some embodiments, the trained Swin transformer-based CE model may be updated. For example, channel data such an SRS received during step 404 may be used to refine the trained Swin transformer-based CE model over time.

[0139] Although Figure 4 illustrates one example procedure 400 for Swin transformer-based channel estimation, various changes may be made to Figure 4. For example, while shown as a series of steps, various steps in Figure 4 could overlap, occur in parallel, occur in a different order, occur any number of times, be omitted, or replaced by other steps.

[0140] To develop a Swin transformer-based CE model capable of functioning effectively across a range of SNR cases, the Swin transformer-based CE model should be trained using diverse training data encompassing various SNRs. However, employing a uniform loss function, such as the mean squared error (MSE), for samples with different SNRs often results in a skewed performance profile. In particular, the Swin transformer-based CE model may demonstrate a propensity for excelling at low SNR cases while simultaneously underperforming at high SNR cases. This phenomenon can be attributed to the fact that the losses incurred by low-SNR data samples are generally more substantial compared to those experienced by high SNR samples. As a consequence, during the training process, the Swin transformer-based CE model naturally gravitates towards focusing on learning from low-SNR cases, as doing so contributes significantly to the reduction of the overall loss function. This bias towards low-SNR cases ultimately hinders the Swin transformer-based CE model's ability to perform well across the entire spectrum of SNR conditions.

[0141] To overcome this issue, some embodiments may apply an SNR weighted MSE as the function as follows:

[0142] ,    (8)

[0143] where Loss is the MSE loss, is number of different SNRs in the training set and the weights

[0144] .    (9)

[0145] In Swin transformer approaches, the input image is divided into non-overlapping windows. Self-attention is done among patches inside each window. As a result, the complexity of an STL increases quadratically with the window size. In the example of Figure 8, the attention window is fixed as a square shape. In other words, each attention window can only include the same number of rows and columns of patches. In the context of channel estimation, the two dimensions of the channel matrix correspond to subcarriers and antennas, respectively. In many cases, the correlation across these two dimensions can be very different. Therefore, a square attention window may not be able to handle the dependency on the two dimensions efficiently. For example, in the frequency domain, a continuous forty subcarriers located within the coherent bandwidth have relatively strong correlation. To better catch the dependency and improve the performance of the CE, an attention window with size forty on the frequency dimension can be used. However, it is possible that the correlation on the antenna domain is relatively weak and a window size four is sufficient to catch the dependency on the antenna domain. In this case, using a square attention window with size forty on the antenna dimension wastes considerable computational resources without improving the performance. Furthermore, in cases with a small number of antennas, using a large attention window size becomes even impossible, as shown in the FIG 9.

[0146] Figure 9illustrates an example limitation of square attention windows 900 according to embodiments of the present disclosure. The embodiment of a limitation of square attention windows of Figure 9 is for illustration only. Different embodiments of a limitation of square attention windows could be used without departing from the scope of this disclosure.

[0147] In the example of Figure 9, a attention window is applied to a input image 902, and a attention window is applied to a input image 904. In the example of input image 902, the window could be expanded to a square window, but this may be inefficient, and the window is unable to capture the entire first dimension of the input image. In the example of input image 904, the attention window cannot be expanded while maintaining the square shape.

[0148] Although Figure 9 illustrates one example limitation of square attention windows 900, various changes may be made to Figure 9. For example, various changes to attention window sizes, the input window sizes, etc., could be made according to particular needs.

[0149] To overcome the limitations of a square attention window, a rectangular attention window may be applied to a Swin transformer-based CE model to catch the dependencies in the antenna and subcarrier dimensions more flexibly. A rectangular attention window can achieve a better balance of performance and model complexity. For example, in some embodiments a rectangular attention window with size may be used, where and are different integers, resulting in a different attention window width for the frequency dimension and antenna dimension. In these embodiments, assuming input of dimensions , an STL as shown in Figure 7 initially partitions the input image into non-overlapping local windows of shape , reshaping the input image into a feature. Here is the number of local windows. Correspondingly, in the SW-MSA step (step 446 in Figure 7) of the Swin transformer-based CE model, the stride of the cyclic shift will be . A Swin transformer-based CE model using a rectangular attention window can catch long dependencies on the frequency domain even if the number of antennas in the channel is small.

[0150] Figure 10illustrates an example method 1000 for Swin transformer-based wireless CE according to embodiments of the present disclosure. An embodiment of the method illustrated in Figure 10 is for illustration only. One or more of the components illustrated in Figure 10 may be implemented in specialized circuitry configured to perform the noted functions or one or more of the components may be implemented by one or more processors executing instructions to perform the noted functions. Other embodiments of a method for Swin transformer-based wireless CE could be used without departing from the scope of this disclosure.

[0151] In the example of Figure 10, method 1000 begins at step 1010. At step 1010, a BS (such as gNB 102 of Figure 1) generates training data for a Swin transformer-based CE model. For example, the BS may generate the training data for the Swin transformer-based CE model similar as described regarding step 401 of Figure 4. In some embodiments, to generate at least some of the training data, the BS may store a plurality of SRSs received by the BS over a period of time. For example, the SRSs may be received by one or more of UEs 111-116 of Figure 1. In some embodiments, to generate at least some of the training data, the BS may perform a channel simulation based on at least one wireless channel model.

[0152] In some embodiments, the Swin transformer-based CE model may include (i) a shallow feature extraction module (such as shallow feature extraction 431 of Figure 6), (ii) a deep feature extraction module comprising at least one residual Swin transformer block (RSTB) that includes a plurality of Swin transformer layers (STLs) (such as deep feature extraction 432 of Figure 6), and (iii) a high-quality image construction module configured to reconstruct a high quality image from features extracted by the shallow feature extraction module and the deep feature extraction module (such as high quality image reconstruction 433 of Figure 6). In some embodiments, each STL of the plurality of STLs may include (i) a first layer norm (LN) layer (such as LN 441 or LN 445 of Figure 7), (ii) an attention function (such as W-MSA 442 or SW-MSA 446 of Figure 7), (iii) a second LN layer (such as LN 443 or LN 447 of FIG 7), and (iv) a multi-layered perceptron function (such as MLP 444 or MLP 448 of Figure 7). In some embodiments, successive STLs of the plurality of STLs may alternate between two configuration settings to apply shifted window partitioning in every other STL of the plurality of STLs (such as STL structure 700 of Figure 7).

[0153] At step 1020, the BS preprocesses the training data. For example, the BS may preprocess the training data similar as described regarding step 402 of Figure 4. In some embodiments, to preprocess the training data, the BS may transform the training data, with an inverse fast Fourier transform (IFTT), from a frequency domain to a delay domain (such as in step 421 of Figure 5), and transform the transformed training data, with a 2-dimensional fast Fourier transform (2D FFT), from the delay domain to an angular domain (such as in step 422 of Figure 5).

[0154] At step 1030, the BS trains the Swin transformer-based CE model with the preprocessed training data. For example, the BS may train the Swin transformer-based CE model with the preprocessed training data similar as described regarding step 403 of Figure 4.

[0155] At step 1040, the BS receives, over a wireless communication channel, an SRS. For example, the BS may receive an SRS from one of UEs 111-116 of Figure 1.

[0156] At step 1050, the BS provides the SRS as an input image to the trained Swin transformer-based CE model. Input image is used by the trained Swin transformer-based CE model to estimate the channel. For example, the trained Swin transformer-based CE model may estimate the channel similar as described regarding step 404 of Figure 4.

[0157] At step 1060, the BS receives as output from the trained Swin transformer-based CE model, a CE for the wireless communication channel.

[0158] In some embodiments, the BS may update the Swin transformer-based CE model based on the SRS received over the wireless communication channel.

[0159] In some embodiments, the Swin transformer-based CE model may be configured to perform CEs based on a rectangular attention window of size [M,N], whereMandNare different integers,Mcorresponds with subcarrier features, andNcorresponds with antenna features.

[0160] Although Figure 10 illustrates one example method 1000 for Swin transformer-based wireless CE, various changes may be made to Figure 10. For example, while shown as a series of steps, various steps in Figure 10 could overlap, occur in parallel, occur in a different order, occur any number of times, be omitted, or replaced by other steps.

[0161] Any of the above variation embodiments can be utilized independently or in combination with at least one other variation embodiment. The above flowcharts illustrate example methods that can be implemented in accordance with the principles of the present disclosure and various changes could be made to the methods illustrated in the flowcharts herein. For example, while shown as a series of steps, various steps in each figure could overlap, occur in parallel, occur in a different order, or occur multiple times. In another example, steps may be omitted or replaced by other steps.

[0162] Although the present disclosure has been described with exemplary embodiments, various changes and modifications may be suggested to one skilled in the art. It is intended that the present disclosure encompass such changes and modifications as fall within the scope of the appended claims. None of the description in this application should be read as implying that any particular element, step, or function is an essential element that must be included in the claim scope. The scope of patented subject matter is defined by the claims.

[0163] Meanwhile, although specific embodiments of the present disclosure have been described in detail, various modifications may be made without departing from the scope of the present disclosure. Therefore, the scope of the present disclosure should not be limited to the described embodiments, but should be defined by the claims and equivalents thereof.

Claims

1.A method performed by a base station in a wireless communication system, the method comprising:generating channel data associated with a channel between the base station and a terminal;preprocessing the channel data for a training of a channel estimation model, the channel estimation model being based on a shifted-window (Swin) transformer;training the channel estimation model based on preprocessed channel data; andperforming a channel estimation for the channel by using the channel estimation model,wherein a sounding reference signal (SRS) received from the terminal on the channel is an input of the channel estimation model.2.The method of claim 1,wherein preprocessing the channel data includes:transforming the channel data with frequency-antenna domain to a first training data with delay-antenna domain by using an inverse fast Fourier transform (IFFT); andtransforming the first training data to a second training data with a delay-angular domain by using a two-dimensional fast Fourier transform (2D FFT), the 2D FFT being based on an antenna structure of the base station,wherein the second training data is used for the training of the channel estimation model.3.The method of claim 1,wherein information on a signal-to-noise ratio (SNR) is generated based on the channel data, the SNR being associated with a signal received on the channel,wherein a loss function for the channel estimation model is configured as a mean squared error (MSE) loss divided by a weight, the MSE loss and the weight being associated with the SNR, andwherein the weight increases as the SNR decreases.4.The method of claim 1,wherein the channel data includes at least one of a plurality of SRSs received on the channel for a period of time or channel response data based on a simulation for the channel, andwherein the SRS is used for an update of the channel estimation model.5.The method of claim 1,wherein the channel estimation model includes:a shallow feature extraction module extracting a low-level feature from the preprocessed channel data;a deep feature extraction module comprising at least one residual Swin transformer block (RSTB) that includes a plurality of Swin transformer layers (STLs); anda high-quality image construction module configured to reconstruct a high-quality image from features extracted by the shallow feature extraction module and the deep feature extraction module,wherein an STL of the plurality of STLs is configured to partition the low-level feature into a window set, the window set comprising a plurality of non-overlapping windows.6.The method of claim 5,wherein a first STL and a second STL successive to the first STL are included in the plurality of STLs,wherein the first STL partitions a first low-level feature into a first window set,wherein the first low-level feature is cyclic-shifted to a second low-level feature,wherein the second STL partitions the second low-level feature into a second window set, andwherein the second low-level feature is shifted back to the first low-level feature after an attention algorithm is performed for the second window.7.The method of claim 5,wherein a window included in the window set is configured to have a size of [M,N],wherein theMand theNare different integers,wherein theMcorresponds to a subcarrier feature on the channel, andwherein theNcorresponds to an antenna feature of the base station.8.A base station in a wireless communication system, the base station comprising:a transceiver;a processor coupled to the transceiver; anda memory coupled to the processor, storing instructions executable by the processor to cause the base station to:generate channel data associated with a channel between the base station and a terminal,preprocess the channel data for a training of a channel estimation model, the channel estimation model being based on a shifted-window (Swin) transformer,train the channel estimation model based on preprocessed channel data, andperform a channel estimation for the channel by using the channel estimation model,wherein a sounding reference signal (SRS) received from the terminal on the channel is an input of the channel estimation model.9.The base station of claim 8,wherein to preprocess the channel data, the instructions executable by the processor further cause the base station to:transform the channel data with frequency-antenna domain to a first training data with delay-antenna domain by using an inverse fast Fourier transform (IFFT), andtransform the first training data to a second training data with a delay-angular domain by using a two-dimensional fast Fourier transform (2D FFT), the 2D FFT being based on an antenna structure of the base station,wherein the second training data is used for the training of the channel estimation model.10.The base station of claim 8,wherein information on a signal-to-noise ratio (SNR) is generated based on the channel data, the SNR being associated with a signal received on the channel,wherein a loss function for the channel estimation model is configured as a mean squared error (MSE) loss divided by a weight, the MSE loss and the weight being associated with the SNR, andwherein the weight increases as the SNR decreases.11.The base station of claim 8,wherein the channel data includes at least one of a plurality of SRSs received on the channel for a period of time or channel response data based on a simulation for the channel, andwherein the SRS is used for an update of the channel estimation model.12.The base station of claim 8,wherein the channel estimation model includes:a shallow feature extraction module extracting a low-level feature from the preprocessed channel data;a deep feature extraction module comprising at least one residual Swin transformer block (RSTB) that includes a plurality of Swin transformer layers (STLs); anda high-quality image construction module configured to reconstruct a high-quality image from features extracted by the shallow feature extraction module and the deep feature extraction module,wherein an STL of the plurality of STLs is configured to partition the low-level feature into a window set, the window set comprising a plurality of non-overlapping windows.13.The base station of claim 12,wherein a first STL and a second STL successive to the first STL are included in the plurality of STLs,wherein the first STL partitions a first low-level feature into a first window set,wherein the first low-level feature is cyclic-shifted to a second low-level feature,wherein the second STL partitions the second low-level feature into a second window set, andwherein the second low-level feature is shifted back to the first low-level feature after an attention algorithm is performed for the second window.14.The base station of claim 12,wherein a window included in the window set is configured to have a size of [M,N],wherein theMand theNare different integers,wherein theMcorresponds to a subcarrier feature on the channel, andwherein theNcorresponds to an antenna feature of the base station.

Citation Information

Patent Citations

  • Machine Learning-Based Channel Estimation

    US20220376957A1

  • Systems and methods for radio frequency calibration exploiting channel reciprocity in distributed input distributed output wireless communications

    WO2014151150A1

  • Data collection method and apparatus for ai / ML model

    WO2023245498A1