Hybrid sequential training for encoder and decoder models

By adopting a hybrid sequential training method in wireless communication technology, the problem of inefficient training of encoder and decoder models is solved, and a more efficient and flexible model training process is achieved.

CN120051941APending Publication Date: 2025-05-27QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072510.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-04
Filing Date
2023-07-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing wireless communication technologies have problems of inefficiency and complex training processes in the training of encoder and decoder models.

Method used

Using a mixed sequential training method, the second model is trained by a first device to receive a function associated with the trained model and select a weight associated with the second model based on the function, parallel training of the encoder and decoder models is realized.

Benefits of technology

It improves the training efficiency of encoder and decoder models, simplifies the training process, and enhances the accuracy and flexibility of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051941A_ABST
    Figure CN120051941A_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure generally relate to wireless communications. In some aspects, a first device may receive, from a second device, a function associated with a trained first model, the function configured to output one or more gradients associated with the trained first model. The first device may train the second model based on selecting one or more weights associated with the second model using the one or more gradients obtained based on inputting the one or more activations and the one or more inputs into the function. Numerous other aspects are described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This patent application claims priority to PCT Patent Application No. PCT / CN2022 / 129967, titled "HYBRID SEQUENTIAL TRAINING FOR ENCODER AND DECODER MODELS", filed on November 4, 2022, and assigned to the assignee of this application. The disclosure of the prior application is considered to be a part of this patent application and is incorporated herein by reference.

[0003] Field of Disclosure

[0004] Aspects of the present disclosure generally relate to wireless communication and relate to techniques and apparatus for hybrid sequential training for encoder and decoder models. Background Art

[0005] Wireless communication systems are widely deployed to provide various telecommunication services such as telephony, video, data, messaging, and broadcasting. A typical wireless communication system may employ multiple access techniques capable of supporting communication with multiple users by sharing available system resources (e.g., bandwidth, transmit power, etc.). Examples of such multiple access techniques include Code Division Multiple Access (CDMA) systems, Time Division Multiple Access (TDMA) systems, Frequency Division Multiple Access (FDMA) systems, Orthogonal Frequency Division Multiple Access (OFDMA) systems, Single - Carrier Frequency Division Multiple Access (SC - FDMA) systems, Time Division Synchronous Code Division Multiple Access (TD - SCDMA) systems, and Long Term Evolution (LTE). LTE / Advanced LTE is an enhanced set of the Universal Mobile Telecommunications System (UMTS) mobile standard promulgated by the Third Generation Partnership Project (3GPP).

[0006] A wireless network may include one or more network nodes that support communication for wireless communication devices such as user equipment (UE) or multiple UEs. A UE may communicate with a network node via downlink communication and uplink communication. "Downlink" (or "DL") refers to the communication link from a network node to a UE, while "uplink" (or "UL") refers to the communication link from a UE to a network node. Some wireless networks may support device - to - device communication (such as via a local link, a Wireless Local Area Network (WLAN) link, and / or a Wireless Personal Area Network (WPAN) link, etc.).

[0007] The above multiple access techniques have been adopted in various telecommunication standards to provide a common protocol that enables different UEs to communicate at the urban, national, regional, and / or global levels. New Radio (NR), which may be referred to as 5G, is an enhanced set of the LTE mobile standard promulgated by 3GPP. NR is designed to better support mobile broadband Internet access by using Orthogonal Frequency Division Multiplexing (OFDM) with Cyclic Prefix (CP) (CP-OFDM) on the downlink, CP-OFDM and / or Single Carrier Frequency Division Multiplexing (SC-FDM) (also known as Discrete Fourier Transform Spread OFDM (DFT-s-OFDM)) on the uplink, and supporting beamforming, Multiple-Input Multiple-Output (MIMO) antenna technology, and carrier aggregation to improve spectral efficiency, reduce costs, improve services, utilize new spectra, and better integrate with other open standards. As the demand for mobile broadband access continues to grow, further improvements to LTE, NR, and other radio access technologies are still useful.

[0008] Overview

[0009] Some aspects described herein relate to a first device for wireless communication. The first device may include one or more memories and one or more processors coupled to the one or more memories. The one or more processors may be configured to receive, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model. The one or more processors may be configured to train a second model based on using the one or more gradients to select one or more weights associated with the second model, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function.

[0010] Some aspects described herein relate to a first device for wireless communication. The first device may include one or more memories and one or more processors coupled to the one or more memories. The one or more processors may be configured to train a first model based on one or more inputs to obtain a trained first model, the trained first model being associated with one or more activations associated with an output of the trained first model. The one or more processors may be configured to transmit, to a second device, a function associated with the trained first model, the function being configured to output one or more activations based on ground truth inputs.

[0011] Some aspects described herein relate to a wireless communication method performed by a first device. The method may include receiving, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model. The method may include training a second model based on selecting one or more weights associated with the second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function.

[0012] Some aspects described herein relate to a wireless communication method performed by a first device. The method may include training a first model based on one or more inputs to obtain a trained first model, the trained first model being associated with one or more activations associated with an output of the trained first model. The method may include transmitting, to a second device, a function associated with the trained first model, the function being configured to output one or more activations based on ground truth inputs.

[0013] Some aspects described herein relate to a non-transitory computer-readable medium storing an instruction set for wireless communication by a first device. The instruction set, when executed by one or more processors of the first device, may cause the first device to receive, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model. The instruction set, when executed by one or more processors of the first device, may cause the first device to train a second model based on selecting one or more weights associated with the second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function.

[0014] Some aspects described herein relate to a non-transitory computer-readable medium storing an instruction set for wireless communication by a first device. The instruction set, when executed by one or more processors of the first device, may cause the first device to train a first model based on one or more inputs to obtain a trained first model, the trained first model being associated with one or more activations associated with an output of the trained first model. The instruction set, when executed by one or more processors of the first device, may cause the first device to transmit, to a second device, a function associated with the trained first model, the function being configured to output one or more activations based on ground truth inputs.

[0015] Some aspects described herein relate to a device for wireless communication. The device may include means for receiving, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model. The device may include means for training a second model based on selecting one or more weights associated with a second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function.

[0016] Some aspects described herein relate to a device for wireless communication. The device may include means for training a first model based on one or more inputs to obtain a trained first model, the trained first model being associated with one or more activations associated with an output of the trained first model. The device may include means for transmitting, to a second device, a function associated with the trained first model, the function being configured to output one or more activations based on ground truth inputs.

[0017] Aspects generally include methods, apparatus, systems, computer program products, non-transitory computer-readable media, user equipment, base stations, network entities, network nodes, wireless communication devices, and / or processing systems substantially as described herein with reference to the figures and the description and as illustrated in the figures and the description.

[0018] The foregoing has outlined rather broadly the features and technical advantages of examples according to the present disclosure in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The disclosed concepts and specific examples may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes as the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein, in terms of both their organization and method of operation, as well as associated advantages, will be better understood when considered in conjunction with the following description taken in connection with the accompanying figures. Each of the figures is provided for the purpose of illustration and description and is not to be construed as limiting the scope of the claims.

[0019] While aspects are described herein by way of illustration of some examples, those skilled in the art will understand that such aspects can be implemented in many different arrangements and scenarios. The techniques described herein can be implemented using different platform types, devices, systems, shapes, sizes, and / or packaging arrangements. For example, some aspects can be implemented via integrated chip embodiments or other devices based on non-module components (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / shopping devices, medical devices, and / or artificial intelligence devices). Aspects can be implemented in chip-level components, module components, non-module components, non-chip-level components, device-level components, and / or system-level components. Devices incorporating the described aspects and features can include additional components and features for implementing and practicing the claimed and described aspects. For example, the transmission and reception of wireless signals can include one or more components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders, and / or summers). The aspects described herein are intended to be practiced in a wide variety of devices, components, systems, distributed arrangements, and / or end-user devices of various sizes, shapes, and configurations. Brief Description of the Drawings

[0021] To understand the above-described features of the present disclosure in detail, a more specific description of the above briefly summarized content can be made with reference to the aspects, some of which are illustrated in the drawings. However, it should be noted that the drawings only illustrate some typical aspects of the present disclosure and should not be considered as limiting its scope, as the description may allow other equally effective aspects. The same reference numerals in different drawings can identify the same or similar elements.

[0022] Figure 1 is a diagram illustrating an example of a wireless network according to the present disclosure.

[0023] Figure 2 is a diagram illustrating an example of a network node and a user equipment in communication in a wireless network according to the present disclosure.

[0024] Figure 3 is a diagram illustrating an example of a decomposed base station architecture according to the present disclosure.

[0025] Figure 4 is a diagram illustrating an example architecture of a functional framework for radio access network intelligence achieved through data collection according to the present disclosure.

[0026] Figure 5 is a diagram illustrating an example architecture associated with artificial intelligence / machine learning (AI / ML)-based channel state feedback compression according to the present disclosure.

[0027] Figure 6 is a diagram illustrating an example associated with multi - vendor AI / ML training according to the present disclosure.

[0028] Figure 7A and Figure 7B is a diagram illustrating an example associated with concurrent training for encoder and decoder models according to the present disclosure.

[0029] Figure 8 is a diagram illustrating an example associated with sequential training for encoder and decoder models according to the present disclosure.

[0030] Figure 9A and Figure 9B is a diagram illustrating an example associated with vector quantization according to the present disclosure.

[0031] Figure 10 is a diagram illustrating an example associated with hybrid sequential training for encoder and decoder models according to the present disclosure.

[0032] Figure 11 is a diagram illustrating an example associated with hybrid sequential training for encoder and decoder models according to the present disclosure.

[0033] Figure 12 is a diagram illustrating an example process, such as that performed by a first device, according to the present disclosure.

[0034] Figure 13 is a diagram illustrating an example process, such as that performed by a first device, according to the present disclosure.

[0035] Figure 14 is a diagram of an example apparatus for wireless communication according to the present disclosure.

[0036] Figure 15 is a diagram of an example apparatus for wireless communication according to the present disclosure.

[0037] Detailed description

[0038] Aspects of the present disclosure are described more fully hereinafter with reference to the accompanying drawings. However, the present disclosure may be embodied in many different forms and should not be construed as limited to any specific structure or function presented throughout this disclosure. Rather, these aspects are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art. Those skilled in the art will appreciate that the scope of the present disclosure is intended to cover any aspect of the present disclosure disclosed herein, whether implemented independently or combined with any other aspect of the present disclosure. For example, any number of the aspects set forth herein may be used to implement an apparatus or practice a method. Additionally, the scope of the present disclosure is intended to cover such apparatus or methods practiced using other structures, functionality, or a combination of structures and functionality that supplement or are additional to the various aspects of the present disclosure set forth herein. It should be understood that any aspect of the present disclosure disclosed herein may be implemented by one or more elements of a claim.

[0039] Certain aspects of a telecommunications system will now be presented with reference to various apparatuses and techniques. These apparatuses and techniques will be described in detail hereinafter and illustrated in the drawings by various blocks, modules, components, circuits, steps, processes, algorithms, etc. (collectively referred to as "elements"). These elements may be implemented using hardware, software, or a combination thereof. Whether such elements are implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system.

[0040] Although aspects may be described herein using terminology commonly associated with 5G or New Radio (NR) radio access technology (RAT), aspects of the present disclosure may be applied to other RATs, such as 3G RAT, 4G RAT, and / or a RAT after 5G (e.g., 6G).

[0041] Figure 1FIG. is a diagram illustrating an example of a wireless network 100 in accordance with the present disclosure. The wireless network 100 can be a 5G (e.g., NR) network and / or a 4G (e.g., Long Term Evolution (LTE)) network, etc. or can include elements thereof. The wireless network 100 can include one or more network nodes 110 (shown as network nodes 110a, network nodes 110b, network nodes 110c, and network nodes 110d), user equipment (UE) 120 or multiple UEs 120 (shown as UEs 120a, UEs 120b, UEs 120c, UEs 120d, and UEs 120e), and / or other entities. The network nodes 110 are network nodes that communicate with the UEs 120. As shown, the network nodes 110 can include one or more network nodes. For example, the network nodes 110 can be a centralized network node, which means that the centralized network node is configured to utilize a radio protocol stack physically or logically integrated within a single radio access network (RAN) node (e.g., integrated within a single device or unit). As another example, the network nodes 110 can be a decomposed network node (sometimes referred to as a decomposed base station), which means that the network nodes 110 are configured to utilize a protocol stack physically or logically distributed among two or more nodes (such as one or more central units (CUs), one or more distributed units (DUs), or one or more radio units (RUs)).

[0042] In some examples, the network nodes 110 are or include network nodes that communicate with the UEs 120 via a radio access link, such as an RU. In some examples, the network nodes 110 are or include network nodes that communicate with other network nodes 110 via a fronthaul link or a midhaul link, such as a DU. In some examples, the network nodes 110 are or include network nodes that communicate with other network nodes 110 via a midhaul link or communicate with a core network via a backhaul link, such as a CU. In some examples, the network nodes 110 (such as a centralized network node 110 or a decomposed network node 110) can include multiple network nodes, such as one or more RUs, one or more CUs, and / or one or more DUs. The network nodes 110 can include, for example, an NR base station, an LTE base station, a Node B, an eNB (e.g., in 4G), a gNB (e.g., in 5G), an access point, a transmission reception point (TRP), a DU, an RU, a CU, a mobility element of the network, a core network node, a network element, network equipment, a RAN node, or a combination thereof. In some examples, the network nodes 110 can be interconnected with each other or with one or more other network nodes 110 in the wireless network 100 using any suitable transport network via various types of fronthaul, midhaul, and / or backhaul interfaces (such as a direct physical connection, an air interface, or a virtual network).

[0043] In some examples, network node 110 may provide communication coverage for a specific geographical area. In the 3rd Generation Partnership Project (3GPP), the term "cell" may refer to the coverage area of radio node 110 and / or the network node subsystem serving that coverage area, depending on the context in which the term is used. Network node 110 may provide communication coverage for macro cells, pico cells, femto cells, and / or another type of cell. A macro cell may cover a relatively large geographical area (e.g., with a radius of several kilometers) and may allow unconstrained access by UEs 120 with a service subscription. A pico cell may cover a relatively small geographical area and may allow unconstrained access by UEs 120 with a service subscription. A femto cell may cover a relatively small geographical area (e.g., a residence) and may allow constrained access by UEs 120 associated with that femto cell (e.g., UEs 120 in a Closed Subscriber Group (CSG)). The network node 110 for a macro cell may be referred to as a macro network node. The network node 110 for a pico cell may be referred to as a pico network node. The network node 110 for a femto cell may be referred to as a femto network node or a home network node. In Figure 1 the example shown, network node 110a may be a macro network node for macro cell 102a, network node 110b may be a pico network node for pico cell 102b, and network node 110c may be a femto network node for femto cell 102c. A network node may support one or more (e.g., three) cells. In some examples, a cell may not necessarily be stationary, and the geographical area of the cell may move according to the location of the moving network node 110 (e.g., a mobile network node).

[0044] In some aspects, the term "base station" or "network node" may refer to a centralized base station, a decomposed base station, an integrated access and backhaul (IAB) node, a relay node, or one or more of its components. For example, in some aspects, a "base station" or "network node" may refer to a CU, a DU, an RU, a near real-time (near-RT) RAN intelligent controller (RIC), or a non-real-time (non-RT) RIC, or a combination thereof. In some aspects, the term "base station" or "network node" may refer to a single device configured to perform one or more functions, such as those described herein in connection with network node 110. In some aspects, the term "base station" or "network node" may refer to multiple devices configured to perform one or more functions. For example, in some distributed systems, each of several different devices (which may be located in the same geographical location or different geographical locations) may be configured to perform at least a portion of a function, or to repeat the performance of at least a portion of the function, and the term "base station" or "network node" may refer to any one or more of these different devices. In some aspects, the term "base station" or "network node" may refer to one or more virtual base stations or one or more virtual base station functions. For example, in some aspects, two or more base station functions may be instantiated on a single device. In some aspects, the term "base station" or "network node" may refer to one of the base station functions and not the other. In this way, a single device may include more than one base station.

[0045] Wireless network 100 may include one or more relay stations. A relay station is a network entity that can receive a transmission of data from an upstream node (e.g., network node 110 or UE 120) and send the transmission of the data to a downstream node (e.g., UE 120 or network node 110). A relay station may be a UE 120 that is capable of relaying transmissions for other UEs 120. In Figure 1 the example shown, network node 110d (e.g., a relay network node) may communicate with network node 110a (e.g., a macro network node) and UE120d to facilitate communication between network node 110a and UE 120d. A network node 110 that relays communications may be referred to as a relay station, a relay base station, a relay network node, a relay node, a relay, etc.

[0046] Wireless network 100 may be a heterogeneous network that includes different types of network nodes 110, such as macro network nodes, pico network nodes, femto network nodes, or relay network nodes, etc. These different types of network nodes 110 may have different transmit power levels, different coverage areas, and / or different impacts on interference in wireless network 100. For example, a macro network node may have a high transmit power level (e.g., 5 to 40 watts), while pico network nodes, femto network nodes, and relay network nodes may have lower transmit power levels (e.g., 0.1 to 2 watts).

[0047] The network controller 130 may be coupled to or communicate with a set of network nodes 110 and may provide coordination and control of these network nodes 110. The network controller 130 may communicate with the network nodes 110 via a backhaul communication link or a midhaul communication link. The network nodes 110 may communicate with each other directly or indirectly via a wireless or wired backhaul communication link. In some aspects, the network controller 130 may be a CU or a core network device, or may include a CU or a core network device.

[0048] Each UE 120 may be dispersed throughout the wireless network 100, and each UE 120 may be stationary or mobile. The UE 120 may include, for example, an access terminal, a terminal, a mobile station, and / or a subscriber unit. The UE 120 may be a cellular phone (e.g., a smart phone), a personal digital assistant (PDA), a wireless modem, a wireless communication device, a handheld device, a laptop computer, a cordless phone, a wireless local loop (WLL) station, a tablet device, a camera, a gaming device, a netbook, a smartbook, a superbook, a medical device, a biometric device, a wearable device (e.g., a smart watch, smart clothing, smart glasses, smart wristbands, smart jewelry (e.g., a smart ring or a smart bracelet)), an entertainment device (e.g., a music device, a video device, and / or a satellite radio), a vehicle-mounted component or sensor, a smart meter / sensor, industrial manufacturing equipment, a global positioning system device, UE functionality of a network node, and / or any other suitable device configured to communicate via a wireless or wired medium.

[0049] Some UEs 120 may be considered machine type communication (MTC) UEs, or evolved or enhanced machine type communication (eMTC) UEs. The MTC UE and / or eMTC UE may include, for example, robots, drones, remote devices, sensors, meters, monitors, and / or location tags, which may communicate with network nodes, another device (e.g., a remote device), or some other entity. Some UEs 120 may be considered Internet of Things (IoT) devices, and / or may be implemented as narrowband IoT (NB-IoT) devices. Some UEs 120 may be considered client equipment. The UE 120 may be included inside a housing that houses components of the UE 120, such as a processor component and / or a memory component. In some examples, the processor component and the memory component may be coupled together. For example, the processor component (e.g., one or more processors) and the memory component (e.g., a memory) may be operatively coupled, communicatively coupled, electronically coupled, and / or electrically coupled.

[0050] Generally, any number of wireless networks 100 can be deployed in a given geographical area. Each wireless network 100 can support a specific RAT and can operate on one or more frequencies. The RAT can be referred to as radio technology, air interface, etc. The frequency can be referred to as carrier, frequency channel, etc. Each frequency can support a single RAT in a given geographical area to avoid interference between wireless networks of different RATs. In some cases, an NR or 5G RAT network can be deployed.

[0051] In some examples, two or more UEs 120 (e.g., shown as UEs 120a and 120e) can communicate directly using one or more sidelink channels (e.g., without using network node 110 as an intermediary to communicate with each other). For example, UEs 120 can communicate using peer-to-peer (P2P) communication, device-to-device (D2D) communication, vehicle-to-everything (V2X) protocols (e.g., which can include vehicle-to-vehicle (V2V) protocols, vehicle-to-infrastructure (V2I) protocols, or vehicle-to-pedestrian (V2P) protocols), and / or mesh networking. In such examples, UEs 120 can perform scheduling operations, resource selection operations, and / or other operations described elsewhere herein as being performed by network node 110.

[0052] In some examples, the wireless network 100 can include one or more servers, such as servers 135a, 135b, and 135c. In some examples, servers 135a, 135b, and 135c can be connected wirelessly or otherwise, such as via a wired connection. Servers 135a, 135b, and 135c can be UE-side servers and can communicate with one or more UEs (such as UEs 120a, 120b, and / or 120c). For example, server 135a can be a UE-side server associated with a first UE vendor and can communicate with UE 120a (e.g., a UE associated with the first UE vendor). Server 135b can be a second UE-side server associated with a second UE vendor different from the first UE vendor, and can communicate with UE120b (e.g., UE 120b can be associated with the second UE vendor). A vendor can be a manufacturer or entity that designs, markets, maintains, and / or sells a device (such as a UE or a network node) or one or more components of the device. Server 135b can have similar functionality to server 135a, as described in more detail elsewhere herein, such as in connection with Figures 10 - 15The server 135c can be a network - side server and can communicate with one or more network nodes (such as network node 110a). The servers 135a, 135b, and 135c can also communicate with each other. The servers 135a, 135b, and 135c can communicate using various wireless or wired technologies, such as Ethernet, Wi - Fi, or cellular technologies. For example, the servers 135a and 135b can each, such as by using one or more machine - learning (ML) algorithms, host and train an encoder for use by one or more UEs when encoding information (such as the sensed channel - state feedback from reference signals transmitted by one or more network nodes), as described in more detail elsewhere in this document. The server 135c can, such as by using one or more ML algorithms, host and train a decoder for use by one or more network nodes when decoding information, as described in more detail elsewhere in this document. In some examples, the UE server and the network server can work together to train an encoder for use by the UE when encoding information to be transmitted to a network node. For example, the server 135a can provide the server 135c with input information (such as sensed channel - state feedback) received by the server 135a from one or more UEs. The server 135c can use the received input information to train the decoder and the encoder, and can provide the server 135a with training information for the server 135a to use when training the encoder to be provided to one or more UEs. The server 135a can use the training information to train the encoder, and can transmit encoder parameters for the trained encoder to one or more UEs (such as UE 120a) for use when encoding information to be transmitted to a network node.

[0053] The devices of the wireless network 100 can communicate using the electromagnetic spectrum, which can be subdivided into various categories, bands, channels, etc. according to frequency or wavelength. For example, the devices of the wireless network 100 can communicate using one or more operating bands. In 5G NR, two initial operating bands have been identified as frequency range designations FR1 (410 MHz – 7.125 GHz) and FR2 (24.25 GHz – 52.6 GHz). It should be understood that although a part of FR1 is greater than 6 GHz, in various documents and articles, FR1 is generally (interchangeably) referred to as the "sub - 6 GHz" band. A similar naming issue sometimes occurs with FR2. Although different from the extremely high - frequency (EHF) band (30 GHz – 300 GHz) identified by the International Telecommunication Union (ITU) as the "millimeter - wave" band, FR2 is generally (interchangeably) referred to as the "millimeter - wave" band in various documents and articles.

[0054] The frequency between FR1 and FR2 is generally referred to as the mid-band frequency. Recent 5G NR research has identified the operating bands for these mid-band frequencies as frequency range designator FR3 (7.125 GHz – 24.25 GHz). Bands falling within FR3 can inherit FR1 characteristics and / or FR2 characteristics, and thus can effectively extend the features of FR1 and / or FR2 into the mid-band frequencies. Additionally, higher frequency bands are currently being explored to extend 5G NR operation above 52.6 GHz. For example, three higher operating bands have been identified as frequency range designator FR4a or FR4-1 (52.6 GHz – 71 GHz), FR4 (52.6 GHz – 114.25 GHz), and FR5 (114.25 GHz – 300 GHz). Each of these higher frequency bands falls within the EHF band.

[0055] Considering the above examples, unless otherwise specifically stated, it should be understood that if used herein, terms such as "sub-6 GHz" can be broadly interpreted to mean frequencies that can be less than 6 GHz, can be within FR1, or can include mid-band frequencies. Additionally, unless otherwise specifically stated, it should be understood that if used herein, terms such as "millimeter wave" can be broadly interpreted to mean frequencies that can include mid-band frequencies, can be within FR2, FR4, FR4-a or FR4-1, and / or FR5, or can be within the EHF band. It is contemplated that the frequencies included in these operating bands (e.g., FR1, FR2, FR3, FR4, FR4-a, FR4-1, and / or FR5) can be modified, and the techniques described herein are applicable to those modified frequency ranges.

[0056] In some aspects, a server (e.g., servers 135a, 135b, and / or 135c) can include a communication manager 140. The server may also be referred to herein as a "server device". As described in more detail elsewhere herein, the communication manager 140 can receive, from another server, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model; and train a second model based on selecting one or more weights associated with the second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function. Additionally or alternatively, the communication manager 140 can perform one or more other operations described herein.

[0057] In some aspects, a server (e.g., servers 135a, 135b, and / or 135c) can include a communication manager 150. As described in more detail elsewhere herein, the communication manager 150 can train a first model based on one or more inputs to obtain a trained first model, where the trained first model is associated with one or more activations associated with an output of the trained first model; and transmit a function associated with the trained first model to another server, the function being configured to output one or more activations based on ground truth. Additionally or alternatively, the communication manager 150 can perform one or more other operations described herein.

[0058] As indicated above, Figure 1 is provided as an example. Other examples may differ from the example(s) described with respect to Figure 1 the example(s).

[0059] Figure 2 FIG. 200 is a diagram illustrating an example 200 in which a network node 110 and a UE 120 are in communication in a wireless network 100 in accordance with the present disclosure. The network node 110 can be equipped with a set of antennas 234a through 234t, such as T antennas (T≥1). The UE 120 can be equipped with a set of antennas 252a to 252r, such as R antennas (R≥1). The network node 110 of example 200 includes one or more radio frequency components, such as antennas 234 and a modem 232. In some examples, the network node 110 can include an interface, a communication component, or another component that facilitates communication with the UE 120 or another network node. Some network nodes 110 may not include radio frequency components that facilitate direct communication with the UE 120, such as one or more CUs, or one or more DUs.

[0060] At network node 110, transmit processor 220 may receive data destined for UE 120 (or a group of UEs 120) from data source 212. Transmit processor 220 may select one or more modulation and coding schemes (MCSs) for the UE 120 based at least in part on one or more channel quality indicators (CQIs) received from the UE 120. Network node 110 may process (e.g., encode and modulate) data for the UE 120 based at least in part on the selected MCS(s) for the UE 120 and may provide data symbols to the UE 120. Transmit processor 220 may process system information (e.g., for semi-static resource partitioning information (SRPI)) and control information (e.g., CQI requests, grants, and / or higher layer signaling) and provide overhead symbols and control symbols. Transmit processor 220 may generate reference symbols for reference signals (e.g., cell-specific reference signal (CRS) or demodulation reference signal (DMRS)) and synchronization signals (e.g., primary synchronization signal (PSS) or secondary synchronization signal (SSS)). Transmit (TX) multiple-input multiple-output (MIMO) processor 230 may perform spatial processing (e.g., precoding) on the data symbols, control symbols, overhead symbols, and / or reference symbols, if applicable, and may provide a set of output symbol streams (e.g., T output symbol streams) to a corresponding set of modems 232 (e.g., T modems) (shown as modems 232a through 232t). For example, each output symbol stream may be provided to a modulator component (shown as MOD) of a modem 232. Each modem 232 may process the corresponding output symbol stream (e.g., for OFDM) using the corresponding modulator component to obtain an output sample stream. Each modem 232 may further process (e.g., convert to analog, amplify, filter, and / or up-convert) the output sample stream using the corresponding modulator component to obtain a downlink signal. Modems 232a through 232t may transmit a set of downlink signals (e.g., T downlink signals) via a corresponding set of antennas 234 (e.g., T antennas) (shown as antennas 234a through 234t).

[0061] At the UE 120, an antenna array 252 (shown as antennas 252a through 252r) may receive downlink signals from network node 110 and / or other network nodes 110 and may provide a set of received signals (e.g., R received signals) to a set of modems 254 (e.g., R modems) (shown as modems 254a through 254r). For example, each received signal may be provided to a demodulator component (shown as DEMOD) of a modem 254. Each modem 254 may condition (e.g., filter, amplify, down-convert, and / or digitize) the received signal using the corresponding demodulator component to obtain input samples. Each modem 254 may further process the input samples (e.g., for OFDM) using the demodulator component to obtain received symbols. A MIMO detector 256 may obtain the received symbols from the modems 254, may perform MIMO detection on the received symbols when applicable, and may provide detected symbols. A receive processor 258 may process (e.g., demodulate and decode) the detected symbols, may provide decoded data for the UE 120 to a data sink 260, and may provide decoded control information and system information to a controller / processor 280. The term "controller / processor" may refer to one or more controllers, one or more processors, or a combination thereof. A channel processor may determine a reference signal received power (RSRP) parameter, a received signal strength indicator (RSSI) parameter, a reference signal received quality (RSRQ) parameter, and / or a CQI parameter, etc. In some examples, one or more components of the UE 120 may be included in a housing 284.

[0062] The network controller 130 may include a communication unit 294, a controller / processor 290, and a memory 292. The network controller 130 may include, for example, one or more devices in a core network. The network controller 130 may communicate with the network node 110 via the communication unit 294.

[0063] One or more antennas (e.g., antennas 234a through 234t and / or antennas 252a through 252r) may include, or may be included in, one or more antenna panels, one or more antenna groups, one or more antenna element arrays, and / or one or more antenna arrays, etc. An antenna panel, an antenna group, an antenna element array, and / or an antenna array may include one or more antenna elements (in a single housing or multiple housings), a coplanar antenna element array, a non-coplanar antenna element array, and / or one or more antenna elements coupled to one or more transmit and / or receive components (such as Figure 2 one or more components) of.

[0064] On the uplink, at the UE 120, the transmit processor 264 may receive and process data from the data source 262 and control information from the controller / processor 280 (e.g., for reports including RSRP, RSSI, RSRQ, and / or CQI). The transmit processor 264 may generate reference symbols for one or more reference signals. The symbols from the transmit processor 264 may be precoded by the TX MIMO processor 266 if applicable, further processed by the modem 254 (e.g., for DFT-s-OFDM or CP-OFDM), and transmitted to the network node 110. In some examples, the modem 254 of the UE 120 may include a modulator and a demodulator. In some examples, the UE 120 includes a transceiver. The transceiver may include any combination of antennas 252, modems 254, MIMO detectors 256, receive processors 258, transmit processors 264, and / or TX MIMO processors 266. The transceiver may be used by a processor (e.g., the controller / processor 280) and the memory 282 to perform aspects of any of the methods described herein (e.g., refer to Figure 10 -15).

[0065] At the network node 110, the uplink signals from the UE 120 and / or other UEs may be received by the antenna 234, processed by the modem 232 (e.g., the demodulator component of the modem 232, shown as DEMOD), detected by the MIMO detector 236 if applicable, and further processed by the receive processor 238 to obtain the decoded data and control information transmitted by the UE 120. The receive processor 238 may provide the decoded data to the data sink 239 and the decoded control information to the controller / processor 240. The network node 110 may include a communication unit 244 and may communicate with the network controller 130 via the communication unit 244. The network node 110 may include a scheduler 246 to schedule one or more UEs 120 for downlink communication and / or uplink communication. In some examples, the modem 232 of the network node 110 may include a modulator and a demodulator. In some examples, the network node 110 includes a transceiver. The transceiver may include any combination of antennas 234, modems 232, MIMO detectors 236, receive processors 238, transmit processors 220, and / or TX MIMO processors 230. The transceiver may be used by a processor (e.g., the controller / processor 240) and the memory 242 to perform aspects of any of the methods described herein (e.g., refer to Figure 10 -15).

[0066] References to singular elements are not intended to mean one and only one (unless specifically stated otherwise), but rather "one or more." For example, unless specifically stated otherwise, a reference to an element (e.g., "processor," "controller," "memory," etc.) should be understood to refer to one or more elements (e.g., "one or more processors," "one or more controllers," and / or "one or more memories," etc.). When referring to one or more elements that perform a function (e.g., the steps of a method), one element may perform all of the functions, or more than one element may perform these functions in common. When more than one element performs these functions in common, each function need not be performed by each of these elements (e.g., different functions may be performed by different elements) and / or each function need not be performed only by one element in its entirety (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements that are configured such that another element (e.g., a device) performs a function, one element may be configured such that another element performs all of the functions, or more than one element may be configured in common such that other elements perform these functions.

[0067] In some examples, the servers described herein (e.g., a network server, a UE server, servers 135a, 135b, and / or 135c) may include a bus, a processor, a memory, an input component, an output component, and / or a communication component. The bus may include one or more components that enable wired and / or wireless communication between the various components of the server. For example, the bus may include electrical connections (e.g., wires, traces, and / or leads) and / or a wireless bus. The processor may include a central processing unit, a graphics processing unit, a microprocessor, a controller, a microcontroller, a digital signal processor, a field programmable gate array, an application specific integrated circuit, and / or another type of processing component. The processor may be implemented in hardware, firmware, or a combination of hardware and software. In some examples, the processor may include one or more processors that can be programmed to perform one or more operations or processes described elsewhere herein.

[0068] The memory may include volatile and / or non-volatile memory. For example, the memory may include random access memory (RAM), read-only memory (ROM), hard disk drives, and / or other types of memory (e.g., flash memory, magnetic memory, and / or optical memory). The memory may include internal memory (e.g., RAM, ROM, or hard disk drive) and / or removable memory (e.g., removable via a universal serial bus connection). The memory may be a non-transitory computer-readable medium. The memory may store information related to the operation of the server, one or more instructions, and / or software (e.g., one or more software applications). The input component may enable the server to receive inputs, such as user inputs and / or sensed inputs. For example, the input component may include, by way of example, a touch screen, a keyboard, a keypad, a mouse, buttons, a microphone, switches, sensors, a global positioning system sensor, an accelerometer, a gyroscope, and / or an actuator, etc. The output component may enable the server to provide outputs, such as via a display, a speaker, and / or a light-emitting diode. The communication component may enable the server to communicate with other devices via a wired connection and / or a wireless connection. For example, the communication component may include a receiver, a transmitter, a transceiver, a modem, a network interface card, and / or an antenna.

[0069] The server may perform one or more operations or processes described herein. For example, Figure 12 process 1200 of Figure 13 process 1300 of Figure 12 and / or other processes described herein. For example, a non-transitory computer-readable medium (e.g., the memory) may store an instruction set (e.g., one or more instructions or code) for execution by a processor. The processor may execute the instruction set to perform one or more operations or processes described herein. In some implementations, execution of the instruction set by one or more processors causes the one or more processors and / or the server to perform one or more operations or processes described herein. In some implementations, hardwired circuitry may be used instead of or in combination with these instructions to perform one or more operations or processes described herein. Additionally or alternatively, the processor of the server may be configured to perform one or more operations or processes described herein, such as Figure 13 process 1200 of

[0070] process 1300 of Figure 2Any other component(s) may perform one or more techniques associated with hybrid sequential training for encoder and decoder models, as described in more detail elsewhere herein. In some aspects, the server described herein is the network node 110, is included in the network node 110, or includes Figure 2 one or more components of the network node 110 as shown. In some other aspects, the server described herein is the UE 120, is included in the UE 120, or includes Figure 2 one or more components of the UE 120 as shown. In other aspects, the server described herein may be a device separate from the network node 110 and / or the UE 120 and may be configured to communicate with the network node 110 and / or the UE 120.

[0071] For example, the controller / processor 240 of the network node 110, the controller / processor 280 of the UE 120, and / or Figure 2 any other component(s) may execute or direct, for example, Figure 12 process 1200, Figure 13 process 1300, and / or the operation of other processes as described herein. The memories 242 and 282 may store data and program code for the network node 110 and the UE 120, respectively. In some examples, the memories 242 and / or 282 may include non-transitory computer-readable media storing one or more instructions for wireless communication (e.g., code and / or program code). For example, when the one or more instructions are executed by one or more processors of the server, the network node 110, and / or the UE 120 (e.g., directly executed, or after compilation, transformation, and / or interpretation), the one or more processors, the server, the UE 120, and / or the network node 110 may be caused to execute or direct, for example, Figure 12 process 1200, Figure 13 process 1300, and / or the operation of other processes as described herein. In some examples, executing the instructions may include running the instructions, transforming the instructions, compiling the instructions, and / or interpreting the instructions, etc.

[0072] In some aspects, a server (e.g., servers 135a, 135b, and / or 135c) includes: means for receiving, from another device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model; and / or means for training a second model based on selecting one or more weights associated with the second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function. In some aspects, means for the server to perform the operations described herein may include, for example, one or more of the following: communication manager 140, antenna, modem, MIMO detector, receive processor, transmit processor, TX MIMO processor, controller / processor, input component, output component, communication component, and / or memory, etc.

[0073] In some aspects, a server (e.g., a network server, a UE server, servers 135a, 135b, and / or 135c) includes: means for training a first model based on one or more inputs to obtain a trained first model, the trained first model being associated with one or more activations associated with an output of the trained first model; and / or means for transmitting, to a second device, a function associated with the trained first model, the function being configured to output one or more activations based on a ground truth input. In some aspects, means for the server to perform the operations described herein may include, for example, one or more of the following: communication manager 150, antenna, modem, MIMO detector, receive processor, transmit processor, TX MIMO processor, controller / processor, input component, output component, communication component, and / or memory, etc.

[0074] Although Figure 2 the blocks in are illustrated as different components, the functions described above with respect to these blocks may be implemented using a single hardware, software, or combined component or a combination of various components. For example, the functions described with respect to transmit processor 264, receive processor 258, and / or TX MIMO processor 266 may be performed by or under the control of controller / processor 280.

[0075] As indicated above, Figure 2 is provided as an example. Other examples may be different from the examples described with respect to Figure 2 .

[0076] The deployment of a communication system (such as a 5G NR system) can be arranged with various components or constituent parts in various ways. In a 5G NR system or network, network nodes, network entities, mobility elements of the network, RAN nodes, core network nodes, network elements, base stations, or network equipment can be implemented in a centralized or decomposed architecture. For example, a base station (such as a Node B (NB), evolved NB (eNB), NR base station, 5G NB, access point (AP), TRP, or cell, etc.), or one or more units (or one or more components) performing base station functionality can be implemented as a centralized base station (also referred to as a stand-alone base station or monolithic base station) or a decomposed base station. A "network entity" or "network node" can refer to a decomposed base station or one or more units of a decomposed base station (such as one or more CUs, one or more DUs, one or more RUs, or a combination thereof).

[0077] A centralized base station (e.g., a centralized network node) can be configured to utilize a radio protocol stack physically or logically integrated within a single RAN node (e.g., within a single device or unit). A decomposed base station (e.g., a decomposed network node) can be configured to utilize a protocol stack physically or logically distributed among two or more units (such as one or more CUs, one or more DUs, or one or more RUs). In some examples, a CU can be implemented within a network node, and one or more DUs can be co-located with the CU, or alternatively, can be geographically or virtually distributed across one or more other network nodes. A DU can be implemented to communicate with one or more RUs. Each of the CU, DU, and RU can also be implemented as a virtual unit, such as a virtual central unit (VCU), virtual distributed unit (VDU), or virtual radio unit (VRU), etc.

[0078] Base station type operations or network designs can consider the aggregation characteristics of base station functionality. For example, a decomposed base station can be used in an IAB network, an open radio access network (O-RAN (such as a network configuration initiated by the O-RAN Alliance)), or a virtualized radio access network (vRAN, also referred to as a cloud radio access network (C-RAN)) to facilitate the scaling of a communication system by splitting base station functionality into one or more units that can be individually deployed. A decomposed base station can include functionality implemented across two or more units at various physical locations and functionality implemented for at least one unit that is virtually distributed, which can achieve flexibility in network design. The individual units of a decomposed base station can be configured for wired or wireless communication with at least one other unit of the decomposed base station.

[0079] Figure 3FIG. is an illustration depicting an example disaggregated base station architecture 300 in accordance with the present disclosure. The disaggregated base station architecture 300 may include a CU 310, which may communicate directly with a core network 320 via a backhaul link, or indirectly with the core network 320 through one or more disaggregated control units (such as a near RT RIC 325 via an E2 link, or a non-RT RIC 315 associated with a service management and orchestration (SMO) framework 305, or both). The CU 310 may communicate with one or more DUs 330 via a respective midhaul link (such as via an F1 interface). Each DU 330 may communicate with one or more RUs 340 via a respective fronthaul link. Each RU 340 may communicate with one or more UEs 120 via a respective radio frequency (RF) access link. In some implementations, a UE 120 may be served simultaneously by multiple RUs 340.

[0080] Each of these units (including the CU 310, DU 330, RU 340, and the near RT RIC 325, non-RT RIC 315, and SMO framework 305) may include one or more interfaces, or be coupled to one or more interfaces configured to receive or transmit signals, data, or information (collectively referred to as signals) via a wired or wireless transmission medium. Each of these units, or an associated processor or controller that provides instructions to one or more communication interfaces of a respective unit, may be configured to communicate with one or more of the other units via the transmission medium. In some examples, each of these units may include: a wired interface configured to receive or transmit signals to one or more of the other units over a wired transmission medium; and a wireless interface that may include a receiver, transmitter, or transceiver (such as an RF transceiver) configured to receive or transmit signals, or both, to one or more of the other units over a wireless transmission medium.

[0081] In some aspects, the CU 310 may host one or more higher layer control functions. Such control functions may include Radio Resource Control (RRC) functions, Packet Data Convergence Protocol (PDCP) functions, or Service Data Adaptation Protocol (SDAP) functions, etc. Each control function may be implemented to have an interface configured to communicate signals with other control functions hosted by the CU 310. The CU 310 may be configured to handle user plane functionality (e.g., Central Unit - User Plane (CU - UP) functionality), control plane functionality (e.g., Central Unit - Control Plane (CU - CP) functionality), or a combination thereof. In some implementations, the CU 310 may be logically split into one or more CU - UP units and one or more CU - CP units. When implemented in an O - RAN configuration, the CU - UP units may communicate bidirectionally with the CU - CP units via an interface such as the E1 interface. The CU 310 may be implemented to communicate with the DU 330 as needed for network control and signaling.

[0082] Each DU 330 may correspond to a logical unit that includes one or more base station functions to control the operation of one or more RUs 340. In some aspects, the DU 330 may host one or more of the following, at least in part depending on a functional split (such as the functional split defined by 3GPP): the Radio Link Control (RLC) layer, the Media Access Control (MAC) layer, and one or more high Physical (PHY) layers. In some aspects, one or more high PHY layers may be implemented by one or more modules for forward error correction (FEC) encoding and decoding, scrambling, and modulation and demodulation, etc. In some aspects, the DU 330 may further host one or more low PHY layers (such as implemented by one or more modules for fast Fourier transform (FFT), inverse FFT (iFFT), digital beamforming, or physical random access channel (PRACH) extraction and filtering, etc.). Each layer (which may also be referred to as a module) may be implemented to have an interface configured to communicate signals with other layers (and modules) hosted by the DU 330 or with the control functions hosted by the CU 310.

[0083] Each RU 340 can implement lower layer functionality. In some deployments, the RU 340 controlled by the DU 330 can correspond to a logical node that hosts (such as performs FFT, performs iFFT, digital beamforming, or PRSCH extraction and filtering, etc.) RF processing functions or low PHY layer functions based on a functional split (such as the functional split defined by 3GPP), such as a lower layer functional split. In such an architecture, each RU 340 can be operated to handle over-the-air (OTA) communication with one or more UEs 120. In some implementations, the real-time and non-real-time aspects of the control and user plane communication with the RU(s) 340 can be controlled by the corresponding DU 330. In some scenarios, this configuration can enable each DU 330 and CU 310 to be implemented in a cloud-based RAN architecture, such as a vRAN architecture.

[0084] The SMO framework 305 can be configured to support the deployment and provisioning of RANs with non-virtualized and virtualized network elements. For non-virtualized network elements, the SMO framework 305 can be configured to support the deployment of dedicated physical resources for RAN coverage requirements, which can be managed via an operation and maintenance interface (such as the O1 interface). For virtualized network elements, the SMO framework 305 can be configured to interact with a cloud computing platform (such as the Open Cloud (O-Cloud) platform 390) to perform network element lifecycle management (such as instantiating virtualized network elements) via a cloud computing platform interface (such as the O2 interface). Such virtualized network elements can include, but are not limited to, CU 310, DU 330, RU 340, non-RT RIC 315, and near-RT RIC 315. In some implementations, the SMO framework 305 can communicate with the hardware aspects of a 4G RAN (such as an Open eNB (O-eNB) 311) via the O1 interface. Additionally, in some implementations, the SMO framework 305 can communicate directly with one or more RUs 340 via the corresponding O1 interface. The SMO framework 305 can also include a non-RT RIC 315 configured to support the functionality of the SMO framework 305.

[0085] The non-RT RIC 315 can be configured to include logic functions that implement non-real-time control and optimization of RAN elements and resources, AI / Machine Learning (AI / ML) workflows including model training and updates, or policy-based guidance of applications / features in the near-RT RIC 325. The non-RT RIC 315 can be coupled to or communicate with the near-RT RIC 325 (such as via the A1 interface). The near-RT RIC 325 can be configured to include logic functions that implement near-real-time control and optimization of RAN elements and resources via data collection and actions through an interface (such as via the E2 interface) that connects one or more CUs 310, one or more DUs 330, or both, and the O-eNB to the near-RT RIC 325.

[0086] In some implementations, to generate an AI / ML model to be deployed in the near-RT RIC 325, the non-RT RIC 315 can receive parameters or external enrichment information from an external server. Such information can be utilized by the near-RT RIC 325 and can be received from non-network data sources or from network functions at the SMO framework 305 or the non-RT RIC 315. In some examples, the non-RT RIC 315 or the near-RT RIC 325 can be configured to tune RAN behavior or performance. For example, the non-RT RIC 315 can monitor long-term trends and patterns of performance and employ an AI / ML model to perform corrective actions via the SMO framework 305 (such as reconfiguration via the O1 interface) or via the creation of RAN management policies (such as A1 interface policies).

[0087] As indicated above, Figure 3 is provided as an example. Other examples may be different from the example regarding Figure 3 described.

[0088] Figure 4FIG. 0 is a diagram illustrating an example architecture 400 of a functional framework for radio access network (RAN) intelligence enabled by data collection according to the present disclosure. In some scenarios, the functional framework for RAN intelligence can be realized by further enhancing data collection via use cases and / or examples. For example, the principles or algorithms for RAN intelligence enabled by AI / ML and associated functional frameworks (e.g., AI functionality, and / or input / output of components optimized for enabling AI) have been utilized or learned to identify the benefits of AI-enabled RAN via possible use cases (e.g., compression, beam management, energy saving, load balancing, mobility management, and / or coverage optimization, etc.). In one example, as shown by architecture 400, the functional framework for RAN intelligence can include multiple logical entities, such as a model training host 402, a model inference host 404, a data source 406, and an actuator 408.

[0089] The model inference host 404 can be configured to run an AI / ML model based on inference data provided by the data source 406. The model inference host 404 can utilize the inference data input to generate an output (e.g., a prediction) to the actuator 408. The actuator 408 can be an element or entity of the core network or the RAN. For example, the actuator 408 can be a UE, a network node, a base station (e.g., a gNB), a CU, a DU, and / or an RU, etc. Additionally, the actuator 408 can also depend on the type of task performed by the model inference host 404, the type of inference data provided to the model inference host 404, and / or the type of output generated by the model inference host 404, etc. For example, if the output from the model inference host 404 is associated with beam management, the actuator 408 can be a UE, a DU, or an RU. In other examples, if the output from the model inference host 404 is associated with Tx / Rx scheduling, the actuator 408 can be a CU or a DU.

[0090] After the actuator 408 receives the output from the model inference host 404, the actuator 408 may determine whether to take an action based on the output. For example, if the actuator 408 is a DU or RU, and the output from the model inference host 404 is associated with beam management, the actuator 408 may determine whether to change and / or modify the Tx / Rx beam based on the output. If the actuator 408 determines to take an action based on the output, the actuator 408 may indicate the action to at least one action entity 410. For example, if the actuator 408 determines to change / modify the Tx / Rx beam used for communication between the actuator 408 and the action entity 410 (e.g., UE 120), the actuator 408 may transmit a beam (re)configuration or beam switching indication to the action entity 410. The actuator 408 may modify its Tx / Rx beam based on the beam (re)configuration, such as switching to a new Tx / Rx beam or applying different parameters to the Tx / Rx beam, etc. As another example, the actuator 408 may be a UE, and the output from the model inference host 404 may be associated with beam management. For example, the output may be one or more predicted measurements of one or more beams. The actuator 408 (e.g., UE) may determine to transmit a measurement report (e.g., a layer 1 (L1) RSRP report) to the network node 110.

[0091] The data source 406 may also be configured to collect data that is used as training data for training an ML model or inference data for feeding the ML model inference operation. For example, the data source 406 may collect data from one or more core network and / or RAN entities (which may include the action entity 410) and provide the collected data to the model training host 402 for ML model training. For example, after the action entity 410 (e.g., UE 120) receives the beam configuration from the actuator 408, the action entity 410 may provide performance feedback associated with the beam configuration to the data source 406, where the performance feedback may be used by the model training host 402 to monitor or evaluate the ML model performance (such as whether the output (e.g., prediction) provided to the actuator 408 is accurate). In some examples, if the output provided by the actuator 408 is inaccurate (or the accuracy is lower than an accuracy threshold), the model training host 402 may determine to modify or retrain the ML model used by the model inference host (such as via ML model deployment / update).

[0092] In cross-node machine learning, a neural network can be split into two parts, where the first part includes an encoder of a UE and the second part includes a decoder of a network node. The output of the UE's encoder can be transmitted to the network node as the input of the decoder. For example, the input of the encoder can be channel state information (CSI), such as one or more channel estimates, one or more precoders (e.g., one or more precoding vectors), and / or one or more measurements, etc. The encoder can use a trained AI / ML model to compress the CSI. The output of the encoder model (e.g., the trained AI / ML model) can be transmitted to the network node. The network node can input the received information into the decoder of the network node. The decoder can use a trained AI / ML model to attempt to reconstruct the CSI (e.g., the CSI that is input into the encoder at the UE). To evaluate the CSI compression use case based on machine learning, one or more different types of quantization or dequantization methods can be used, such as vector quantization and / or scalar quantization, etc. In the CSI compression using the bilateral model use case, multiple machine learning models can be trained.

[0093] UEs and network nodes designed, sold, and maintained by different vendors can implement different encoders and decoders to encode and decode information, such as channel state feedback information. A UE server (e.g., server 135a or server 135b) can train the encoder for one or more UEs to implement offline, such as by applying one or more ML algorithms to train the encoder. The UE server can be operated and maintained, for example, by a specific UE vendor, and can determine the encoder parameters for transmission to one or more UEs associated with that specific UE vendor. In some cases, one or more UEs for which the encoder is being trained can transmit input information, such as channel state feedback information, to the UE server.

[0094] A UE server may transmit such input information to a network server, such as server 135c. The network server may train a decoder for offline implementation by one or more network nodes, such as by applying one or more ML algorithms to train the decoder. The network server may be operated and maintained, for example, by a specific network node vendor and may determine decoder parameters to be provided to one or more network nodes associated with the specific network node vendor. In some cases, one or more network nodes for which the decoder is being trained may transmit input information, such as channel state feedback information, to the network server. The network server may further supervise the training of the encoder by one or more UE servers. For example, the network server may receive input information from one or more UE servers or from another source and may use the input information to train both the encoder and the decoder. The network server may then encode the input information using the trained encoder to generate training information. The training information may include both the input information and the output of the encoder, such as the encoded input information. The network server may transmit the training information to one or more UE servers. One or more UE servers may use the training information to perform offline training of the encoder for each respective UE server. Such training may produce one or more encoder parameters for use by one or more UEs when encoding information, and the encoder parameters may be transmitted by one or more UE-side servers to one or more UEs.

[0095] As indicated above, Figure 4 is provided as an example. Other examples may be different from the example described with respect to Figure 4 the example described above.

[0096] Figure 5 FIG. 500 is a diagram illustrating an example architecture associated with AI / ML-based channel state feedback compression in accordance with the present disclosure. As described elsewhere herein, in cross-node machine learning, a neural network may be split into two parts, where the first part includes an encoder 502 of a UE and the second part includes a decoder 504 of a network node. The encoder may include an encoder model, which is an AI / ML model trained to compress CSI. The encoder output at the UE is transmitted to the network node to be provided as an input to the decoder. The decoder may include a decoder model, which is an AI / ML model trained to reconstruct or decompress CSI.

[0097] As Figure 5As shown, the encoder 502 can output a compressed channel state feedback (CSF) or another data signal, which is received as an input at the decoder 504. The decoder 504 can output a reconstructed CSF (e.g., decompressed CSF) or another data signal, such as a precoding vector, etc. In multi-vendor training, each vendor (e.g., a UE vendor or a network node vendor) can be associated with a corresponding server participating in offline training. The UE servers (e.g., server 135a and / or server 135b) can communicate with the network server (e.g., server 135c) during training using server-to-server connections.

[0098] In CSI compression using a two-sided model use case, multiple machine learning models can be trained. In some examples, joint training of the two-sided model at a single-sided / single entity (e.g., UE side or network side) can be utilized. In some examples, joint training of the two-sided model on the network side and the UE side respectively can be utilized. In still some other examples, separate training on the network side and the UE side can be utilized, where the UE-side CSI generation part and the network-side CSI reconstruction part are trained by the UE side and the network side respectively (e.g., separate training can also be referred to as sequential training). "Joint training" can mean that the generation model and the reconstruction model are trained in the same loop for forward propagation and backward propagation. Joint training can be performed at a single node or across multiple nodes both (e.g., through gradient exchange between nodes or servers). Separate training can include sequential training starting from UE-side training, or sequential training starting from network-side training, or parallel training performed by the UE server and the network server.

[0099] As indicated above, Figure 5 is provided as an example. Other examples may be different from the examples regarding Figure 5 described.

[0100] Figure 6 is a diagram illustrating example 600 associated with multi-vendor AI / ML training according to the present disclosure.

[0101] For example, as Figure 6As shown, a first network node (NN1) (e.g., network node 110) may be associated with a first cell and a second network node (NN2) (e.g., network node 110) may be associated with a second cell. A plurality of UEs 120 (e.g., UE 1, UE 2, UE 3, UE 4) may be located within the coverage area of NN1 and / or NN2. In an instance without multi-vendor training, each UE-network node pair may need to utilize different encoder-decoder pairs. Multi-vendor training eliminates the need to utilize different encoder-decoder pairs for each UE-network node pairing. For example, in an instance of multiple UE vendors and one network node vendor, a common network node decoder may be trained to work with multiple UE encoders. Thus, a network node (e.g., NN1) may not need to maintain separate decoder models for each UE located within the coverage area of the cell of that network node. In an example of a single UE vendor and multiple network node vendors, a common UE encoder may be trained to work with multiple network node decoders. In such examples, a UE may not need to maintain separate encoder models for each network node (e.g., such as when a UE moves to a new cell). In an example of multiple UE vendors and multiple network node vendors, UE encoders may be trained to work with multiple network node decoders, and network node decoders may be trained to work with multiple UE encoders. For example, as Figure 6 shown, the respective encoders of UE 1 and UE 2 may be trained to work with the decoder of NN1, while the encoder of UE 4 may be trained to work with the decoder of NN2. However, UE 3 may be located at the cell edge and between NN1 and NN2 such that the encoder of UE 3 may be trained to work with the decoders of both NN1 and NN2. In other words, when UE 3 moves from the coverage area of NN1 to the coverage area of NN2, UE 3 may deploy the same encoder model to communicate with NN1 and NN2 (e.g., where NN1 and NN2 may be associated with different vendors and / or different decoder models). This can reduce the training overhead and / or complexity associated with AI / ML-based CSI compression described herein, as a UE may not need to maintain multiple encoder models for different network node vendors and / or for different network node decoder models. Additionally or alternatively, a network node may not need to maintain multiple decoder models for different UE vendors and / or for different UE encoder models.

[0102] As indicated above, Figure 6 is provided as an example. Other examples may be different from the examples described with respect to Figure 6 above.

[0103] Figure 7A and Figure 7B is a diagram illustrating Examples 700 and 710 associated with concurrent training of an encoder and a decoder model according to the present disclosure. As used herein, joint training or concurrent training occurring at a single device may be referred to as Type 1 training. For example, Type 1 training may be associated with joint training of a bilateral model (e.g., an encoder model and a decoder model) at a single / one-sided entity.

[0104] As Figure 7A shown, an input or ground truth may be provided to an encoder model at a UE (e.g., shown as V Figure 7A in 输入 ). For example, the input may include CSI as described in more detail elsewhere herein. V 输入 may be compressed by the encoder model. The encoder model may output an activation or activation function (e.g., shown as Z in Figure 7A ). An “activation function” or “activation” may refer to the output of a neural network (e.g., the encoder model). For example, the activation function of a node of a neural network defines the output of the node given an input or set of inputs. The UE may transmit and a network node may receive the activation function Z. The network node may provide the activation function Z as an input to the decoder model. The decoder model may provide an output (e.g., shown as V Figure 7A in 输出 ). The output may be a reconstruction of V 输入 and / or a decompression of the activation function Z.

[0105] As Figure 7B shown, Example 710 depicts Type 1 training and model transfer. For example, a device (e.g., a UE server or a network server) may train an encoder model and a decoder model. The device may provide V 输入 and V 输出 to a loss function, which determines the difference between the original input V 输入 of the encoder and the reconstructed version V 输出 of the original input of the decoder. A gradient may be calculated based on the loss function, and the weights of the encoder or decoder may be updated to train the encoder or decoder. As Figure 7B shown, if the joint training occurs at a UE server, the UE server may transmit and the network server may receive an indication of the trained decoder model (e.g., to be provided by the network server to one or more network nodes). As another example, if the joint training occurs at a network server, the network server may transmit and the UE server may receive an indication of the trained encoder model (e.g., to be provided by the UE server to one or more UEs).

[0106] In concurrent training (e.g., type 1 training), both the encoder and the decoder can be jointly trained such that the model weights of both the encoder and the decoder can be jointly optimized. In offline concurrent training, the model can be trained offline and provided to a network node or a UE. However, unilateral concurrent training can allow the trained model to be exposed to a network node or a UE. The joint training can occur at a UE server or a network server. For example, a UE vendor can use its own dataset to train both the encoder and decoder models and can share the trained decoder model with a network server (e.g., associated with a vendor different from the UE vendor). The decoder model shared with this other vendor can reveal or provide relevant information related to the implementation details of components of the UE (e.g., such as a modem of the UE). Similarly, in an example where a network server trains both the encoder and decoder models, the shared encoder model can reveal or provide relevant information related to the implementation details of components of the network node. This information may be revealed in part due to the symmetry that typically exists between the encoder and the decoder. Thus, the trained encoder and decoder may be trade secrets or include proprietary information that a vendor may not want to reveal to another vendor.

[0107] In some other examples, the encoder model and the decoder model can be concurrently trained at different devices (e.g., where the encoder model and the decoder model are trained in the same loop for forward propagation and backpropagation). For example, a UE server can train the encoder model and a network server can train the decoder model. Concurrent training at different devices can be referred to as type 2 training. For example, type 2 training can include the joint training of a bilateral model (e.g., decoder and encoder) on the network side and the UE side respectively. For example, for each forward propagation loop and / or each backpropagation loop, the UE server can generate a forward propagation result (e.g., can generate Z based on providing V to the encoder model). For example, one or more UEs can provide data (e.g., CSI) to the UE server for training the encoder and / or the decoder. The UE server can transmit the forward propagation result (e.g., Z and V) to the network server. The network server can obtain V based on providing Z to the decoder model. The network server can compare V and V. 输入 to generate Z). For example, one or more UEs can provide data (e.g., CSI) to the UE server for training the encoder and / or the decoder. The UE server can transmit the forward propagation result (e.g., Z and V) to the network server. 输入 ) to the network server. The network server can obtain V based on providing Z to the decoder model. 输出 . The network server can compare V 输出 and V 输入The loss function is used to generate the backpropagation results (e.g., gradients). The network server can transmit and the UE server can receive the backpropagation results (e.g., gradients). The UE server can train the encoder model based on the backpropagation results (e.g., gradients). For example, the UE server can update one or more weights of the neural network of the encoder model based on the backpropagation results (e.g., gradients). After training the model, the UE server can transmit the trained encoder model to one or more UEs. Similarly, the network server can transmit the trained decoder model to one or more network nodes. The UEs and network nodes can use the trained models to perform inference, as described in more detail elsewhere in this document.

[0108] Type 2 training ensures that confidential and / or proprietary information is not shared between the UE server and the network server during training (e.g., using distributed training at different devices instead of at a single device as in type 1 training). Additionally, type 2 training can be associated with improved training of the models because these models are trained concurrently in the same loop of forward and backpropagation. However, type 2 training is performed concurrently at different devices. For example, a training session can be established between the UE server and the network server to perform type 2 training. Thus, type 2 training can be associated with restrictions on the timing of this training (e.g., because a training session between the UE server and the network server is required to perform type 2 training).

[0109] As indicated above, Figure 7A and 7B are provided as examples. Other examples can be different from the examples described with respect to Figure 7A and 7B described.

[0110] Figure 8 is a diagram illustrating example 800 associated with sequential training for encoder and decoder models according to the present disclosure. As used herein, sequential training or separate training may be referred to as type 3 training. For example, type 3 training can be associated with separate training of bilateral models (e.g., encoder model and decoder model) at different entities. For example, Figure 8 depicts network-driven sequential training. However, type 3 training can include UE-driven (e.g., UE server-driven) sequential training in a similar manner as described herein.

[0111] As Figure 8 shown, multiple UE encoders can be trained based on the trained network node decoder. For example, the network server can be in accordance with the combination of Figure 7A and 7Bbe trained in a similar manner as described (e.g., using an encoder model at a network server). The network server can transmit and the UE server can receive a dataset. The dataset can include one or more inputs (e.g., one or more Vs) used to train a decoder model. 输入 and / or one or more outputs of the encoder (e.g., one or more Z functions). This can enable different UE servers to use the dataset to train the encoder model. For example, as Figure 8 shown, the UE server can provide V from the dataset 输入 as an input to the encoder model. The UE server can provide the output obtained from the encoder model (e.g., Z UE ) to a loss function, along with the output (e.g., Z) corresponding to the input (e.g., V 输入 ) from the dataset. The loss function can output gradients used by the UE server to update one or more weights of the encoder model, as described in more detail elsewhere herein. For example, training the UE encoder can be achieved by minimizing the loss between Z (e.g., the output of the network node encoder) and the output Z UE of the UE encoder. Thus, type 3 training enables offline individual training at different devices. Additionally, type 3 training can occur at different times at different devices, providing additional flexibility for the training of the encoder and decoder models (e.g., compared to type 2 training described elsewhere herein).

[0112] As indicated above, Figure 8 is provided as an example. Other examples may differ from the example described with respect to Figure 8 .

[0113] Figure 9A and Figure 9B are diagrams illustrating examples 900 and 910 associated with vector quantization in accordance with the present disclosure.

[0114] In vector quantization, an input vector can be quantized and mapped to one or more vectors in a quantization codebook. In some examples, the quantization codebook can include vectors of size 2 or 4, where each entry can be represented by 2 bits or another number of bits. However, in other examples, the quantization codebook can include vectors of different sizes.

[0115] As Figure 9A shown, the input V 输入 can be input into an encoder model, which produces an encoder output Z E . The output Z E can be quantized to produce a quantized output Z q . The quantized output Z q can be processed by a decoder model to attempt to reconstruct V输入 , where the decoder output is V 输出 . As Figure 9B shown, to perform quantization, the quantizer may receive the encoder output Z E and partition Z E into sub-vectors of size d subsets (e.g., 2 or 4). Based on the quantization codebook, the sub-vectors (e.g., Z E,0 , Z E,1 ) are quantized to produce quantized sub-vectors (e.g., Z q,0 , Z q,1 ), where the quantized sub-vectors are mapped to one of the vectors in the codebook. To perform the codebook-based mapping, the quantizer maps the value of the quantized sub-vector to two values in the codebook (e.g., one of the K values of the codebook). For example, the quantizer may map the input to the closest quantized value in the codebook. Then, the quantized sub-vectors are combined to form the quantized output Z q .

[0116] As indicated above, Figure 9A and 9B are provided as examples. Other examples may be different from those described with respect to Figure 9A and 9B .

[0117] As described above, different training techniques can be used to train the encoder model and the decoder model for CSI compression. For example, type 2 training can be used to ensure that confidential and / or proprietary information is not shared between the UE server and the network server during training (e.g., using distributed training at different devices instead of at a single device as in type 1 training). However, type 2 training is performed concurrently at different devices. For example, a training session can be established between the UE server and the network server to perform type 2 training. Thus, type 2 training can be associated with restrictions on the timing of the training (e.g., because a training session between the UE server and the network server is required to perform type 2 training). Type 3 training can be used to provide additional flexibility in the timing of performing the training (e.g., by performing separate training at different devices). However, in some cases, type 2 training may be associated with improved results and / or accuracy of the trained model compared to type 3 training (e.g., because the models in type 2 training are trained concurrently and in the same loop of forward propagation and backward propagation, rather than using the above dataset). Thus, the device performing the training can choose between improved training results and / or accuracy (e.g., by performing type 2 training) or increased flexibility in the timing at which the training occurs (e.g., by performing type 3 training).

[0118] Some of the techniques and apparatuses described herein implement hybrid sequential training of encoder and decoder models. For example, a first device may transmit and a second device may receive an indication of a function associated with a trained model (e.g., a trained encoder model or a trained decoder model) associated with the first device. For example, the first device may offline train the first model in a manner similar to type 3 training. The first device may transmit to the second device a function that emulates a forward propagation path and a backward propagation path to facilitate concurrent training of the second model at the second device. For example, the function may be an application programming interface (API), a software program, an instruction set, code, and / or another function.

[0119] For example, the first device may be a network server and the first model may be a decoder model. The second device may be a UE server and the second model may be an encoder model. The network server may transmit to the UE server a function that accepts an activation function (e.g., Z) and a ground truth (e.g., V 输入 ) as inputs and outputs one or more gradients (e.g., to emulate the backward propagation path of the trained decoder model). The UE server may use the one or more gradients to train the encoder model (e.g., to update one or more weights of the encoder model at least in part based on the one or more gradients). As another example, the first device may be a UE server and the first model may be an encoder model. The second device may be a network server and the second model may be a decoder model. The UE server may transmit and the network server may receive a function that receives a ground truth (e.g., V 输入 ) as an input and outputs an activation function Z (e.g., to emulate the forward propagation path of the trained encoder model). The network server may use the activation function and the ground truth to train the decoder model (e.g., by providing the activation function and the ground truth to a loss function and using the gradient of the loss function to update the weights of the decoder model).

[0120] As a result, the encoder model and the decoder model can be trained using the forward propagation path and the backward propagation path in the same training loop (e.g., in a manner similar to type 2 training), and also be trained sequentially and / or separately. For example, the function provided by the first device to the second device can enable the forward propagation path and the backward propagation path (e.g., which are fixed at the first device) to be simulated at the second device for emulating joint or concurrent training. This can improve the accuracy of the training of the encoder and / or decoder model (e.g., by training concurrently and in the same loop for forward propagation and backward propagation). Additionally, this can increase the flexibility regarding the timing at which the training occurs (e.g., because the encoder model and the decoder model can be trained separately and / or at different times). For example, a training session may not be established between the UE server and the network server to jointly train the encoder model and the decoder model.

[0121] Figure 10 is an illustration depicting Example 1000 associated with hybrid sequential training for an encoder and a decoder model in accordance with the present disclosure. As Figure 10 shown, network node 110 (e.g., a base station, CU, DU, and / or RU) can communicate with UE 120. In some aspects, network node 110 and UE 120 can be part of a wireless network (e.g., wireless network 100). UE 120 and network node 110 may have established a wireless connection prior to the operations Figure 10 shown. As Figure 10 shown, UE 120 can communicate with UE server 1005 (e.g., server 135a or server 135b). UE server 1005 can be associated with the vendor of UE 120. Similarly, network node 110 can communicate with network server 1010 (e.g., server 135c). This network server 1010 can be associated with the vendor of network node 110.

[0122] As described herein, the operations performed by UE 120 and / or UE server 1005 can be referred to as "UE - side" operations. Similarly, the operations performed by network node 110 and / or network server 1010 can be referred to as "network - side" operations. In some aspects, one or more (or all) of the operations described herein as being performed by UE server 1005 can be performed by UE 120. Similarly, one or more (or all) of the operations described herein as being performed by network server 1010 can be performed by network node 110 (or another network node).

[0123] In some aspects, actions described herein as being performed by network node 110 may be performed by multiple different network nodes. For example, a configuration action may be performed by a first network node (e.g., a CU or a DU), while a radio communication action may be performed by a second network node (e.g., a DU or an RU). As used herein, a network node 110 "transmitting" communication to a UE 120 may refer to a direct transmission (e.g., from network node 110 to UE 120) or an indirect transmission via one or more other network nodes or devices. For example, if network node 110 is a DU, the indirect transmission to UE 120 may include the DU transmitting the communication to the RU and the RU transmitting the communication to UE 120. Similarly, a UE 120 "transmitting" communication to a network node 110 may refer to a direct transmission (e.g., from UE 120 to network node 110) or an indirect transmission via one or more other network nodes or devices. For example, if network node 110 is a DU, the indirect transmission to network node 110 may include the UE 120 transmitting the communication to the RU and the RU transmitting the communication to the DU.

[0124] As indicated by reference numeral 1015, the network server 1010 may train a decoder model associated with the network node 110. For example, the network server 1010 may train the decoder model in a similar manner as described elsewhere herein (such as in connection with type 1 and / or type 3 training). For example, the network server 1010 may receive from the network node 110, the UE server 1005, and / or one or more UEs 120 one or more data sets to be used as inputs to train the decoder model. For example, the one or more data sets may include CSI. The network server 1010 may deploy an encoder model at the network server 1010. The encoder model may be configured to output an activation function (e.g., Z) based on an input or ground truth (e.g., V 输入 ) provided to the encoder model. The network server 1010 may provide the activation function as an input to the decoder model. The decoder model may output V 输出 , which is a reconstruction of the input or ground truth (e.g., V 输入 ) provided to the encoder model. The network server 1010 may provide the input or ground truth (e.g., V 输入 ) to be provided to the encoder model and the output V 输出 to be provided to the loss function. The loss function may compare V 输入 with V 输出Compare. The network server 1010 can obtain gradients based on the output of the loss function. The network server 1010 can train the decoder model based on the gradients. For example, the network server 1010 can use the gradients to update one or more weights of the neural network of the decoder model (e.g., attempting to minimize the loss function). The network server 1010 can perform one or more training loops in a similar manner to update the weights of the decoder model until the output of the loss function meets the training threshold. For example, the network server 1010 can perform one or more training loops until the difference between the output V 输出 between sufficiently reconstructs the input or ground truth V provided to the encoder model 输入 .

[0125] As shown by reference numeral 1020, the network server 1010 can generate a function based on the trained decoder model. The function can be an API, instruction set, code, software program, and / or another function. The function can be configured to output one or more gradients based on the activated input and input. For example, based on the training of the decoder model, the network server 1010 can configure the function to use the information obtained via the training loop and / or based on the loss function to simulate the forward and backward propagation paths of the decoder model. For example, the function can be configured to mimic or simulate the forward and backward propagation paths of the trained decoder model when executed by a device (such as the UE server 1005). For example, the function can be configured to accept an activation (e.g., Z) and a ground truth (e.g., V 输入 ) as inputs and return gradients as outputs (e.g., which can be used to update the weights of the encoder model, as described in more detail elsewhere in this document). In other words, the function can be configured to provide the results of the backward propagation path associated with the trained decoder model (e.g., for the training loop).

[0126] In some aspects, the network server 1010 can generate the function based on the trained decoder model. For example, the network server 1010 can determine, during the training process of the decoder model, from various activations (e.g., Z) and ground truths (e.g., V 输入)The obtained gradient. The network server 1010 may configure the function to provide a given gradient based on a given activation and / or ground truth input to the function (e.g., using information obtained via a training loop and / or based on a loss function such that the encoder model can use the forward propagation path and the backpropagation path in the same training loop as the decoder model (e.g., in a manner similar to type 2 training), while also enabling the decoder model and the encoder model to be trained sequentially and / or separately). Additionally or alternatively, the function may be pre-configured (e.g., by a vendor associated with the network server 1010). In such examples, the network server 1010 may obtain the function from the memory of the network server 1010.

[0127] In some aspects, the network server 1010 (and / or network node 110) may determine a codebook for vector quantization associated with the compressed CSI, as described in more detail elsewhere herein (such as in conjunction with Figure 9A and Figure 9B ). For example, the network server 1010 (and / or network node 110) may train a vector quantization model as part of training the decoder model. For example, the network server 1010 (and / or network node 110) may train a quantizer associated with vector quantization. As part of the training of the decoder model, a quantization codebook (e.g., a vector codebook or a scalar codebook) may be determined at the network server 1010 (and / or network node 110). In some aspects, a function (e.g., an API or other function) generated by the network server 1010 may include a vector quantization component. For example, the function may be configured to simulate the effect of quantization on activations or other information output by or input to the trained decoder model.

[0128] In some aspects, the network server 1010 may generate a function associated with multiple decoder models. For example, the network server 1010 may train multiple decoder models (e.g., in a manner similar to that described in more detail elsewhere herein). In some aspects, the multiple decoder models may be associated with respective UE vendors. As another example, the multiple decoder models may be associated with respective types of CSI (e.g., a first decoder model may be associated with a precoding vector, a second decoder model may be associated with a channel estimate, etc.). As another example, the multiple decoder models may be associated with respective channel conditions. As another example, the multiple decoder models may be associated with respective CSI sizes (e.g., the size of the CSI to be communicated between the network node 110 and the UE 120). The network server 1010 may generate a function configured to simulate the forward propagation path and the backpropagation path of the multiple trained decoder models.

[0129] As shown by reference numeral 1025, the network server 1010 may transmit and the UE server 1005 may receive the function (e.g., which is associated with a trained decoder model). For example, the network server 1010 and the UE server 1005 may establish a connection (e.g., a wireless connection or a wired connection). The function may be transmitted from the network server 1010 to the UE server 1005 via the connection.

[0130] As shown by reference numeral 1030, the UE server 1005 may use the function to train an encoder model. For example, the UE server 1005 may train the encoder model based on selecting or updating one or more weights associated with the encoder model using the one or more gradients. In some aspects, the one or more gradients may be obtained via an output from the encoder (e.g., Z) and an input to the encoder (e.g., V 输入 ). For example, the one or more gradients may be obtained based on inputting one or more activation functions and one or more input functions (e.g., ground truth) into the function. For example, the UE server 1005 may train the encoder model in a manner similar to type 2 training, as described in more detail above. However, the UE server 1005 may input the one or more activations and one or more input functions (e.g., ground truth) into the function received from the network server 1010, rather than providing the one or more activation functions and one or more input functions (e.g., ground truth) to the network server 1010 (e.g., as in the case of type 2 training). The function may simulate the forward propagation path of the trained decoder model (e.g., providing activation functions into the decoder model and obtaining V 输出 ) and the backpropagation path (e.g., providing gradients based on the output of the loss function). Thus, the encoder model may be trained using the forward propagation path and the backpropagation path in the same training loop (e.g., in a manner similar to type 2 training), while also being trained sequentially and / or separately, which is different from type 2 training.

[0131] In some aspects, as described above, the function may include a vector quantization component. For example, the function may be configured to simulate the effect of quantization on activations or other information output by or input to the trained decoder model. In such examples, the UE server 1005 may use the function to train a quantizer and / or a vector quantization model. In other examples, the UE server 1005 (or UE 120) may determine a codebook for vector quantization associated with compressed CSI, as described in more detail elsewhere herein (such as in connection with Figure 9A and Figure 9B)。For example, the UE server 1005 (and / or the UE 120) can train a vector quantization model as part of training an encoder model. For example, the UE server 1005 (and / or the UE 120) can train a quantizer associated with vector quantization. As part of the training of the encoder model, a quantization codebook (e.g., a vector codebook or a scalar codebook) can be determined at the UE server 1005 (and / or the UE 120). In such examples, the input provided to a function (e.g., an API) can include quantized activations output by the encoder model (e.g., which are quantized using vector quantization and / or the quantization codebook determined by the UE server 1005 and / or the UE 120). In other words, the quantizer can be trained with the encoder model (e.g., and the function may not simulate the effect of such quantization).

[0132] In some aspects, as described above, the function can be associated with multiple trained decoder models. In such examples, training the encoder model can include providing an indication of an identifier associated with the decoder model to the function. For example, the input to the function (e.g., an API) can include a model identifier (e.g., which is associated with a given encoder model and / or decoder model). The function can be configured to provide information based on the model identifier provided to the function. In some examples, the UE server 1005 can train a single encoder model to be operable with each of the multiple trained decoder models. In other examples, the UE server 1005 can train multiple encoder models to be operable with corresponding decoder models from multiple trained decoder models (e.g., if the function is associated with N trained decoder models, the UE server 1005 can train N encoder models).

[0133] In some aspects, the UE server 1005 can receive another function (e.g., a second function) from another network server (e.g., another network server 1010). For example, the other network server can be associated with a network node vendor different from the vendor associated with the network server 1010. The UE server 1005 can use the first function (e.g., the first function received from the network server 1010) and use the second function (e.g., the second function received from another network server 1010) to train an encoder. In other words, the UE server 1005 can use multiple functions provided by network servers associated with different vendors to train an encoder model. In this way, the trained encoder model can be configured to operate with trained decoders associated with the multiple functions (e.g., in a similar manner as described in conjunction with Figure 6 ).

[0134] As indicated by reference numeral 1035, the UE server 1005 may transmit and the UE 120 may receive an indication of the trained encoder model. For example, the UE 120 may download the trained encoder model from the UE server 1005 (e.g., the trained encoder model is trained using a function associated with the decoder model of the network node 110). Similarly, as indicated by reference numeral 1040, the network server 1010 may transmit and the network node 110 may receive an indication of the trained decoder model. For example, the network node 110 may download the trained decoder model from the network server 1010.

[0135] As indicated by reference numeral 1045, the UE 120 and the network node 110 may communicate using the trained encoder model and the trained decoder model, respectively. For example, the UE 120 may obtain CSI to be transmitted to the network node 110. The UE 120 may input the CSI into the trained encoder model. The trained encoder model may output an activation function (e.g., the compressed CSI). In some aspects, the UE 120 may quantize the activation function output by the trained encoder model (e.g., using a quantization codebook and / or vector quantization). The UE 120 may transmit and the network node 110 may receive the activation function output by the trained encoder model (e.g., the compressed CSI). In some aspects, the UE 120 may transmit and the network node 110 may receive a quantized representation of the activation function output by the trained encoder model (e.g., the compressed CSI). The network node 110 may input the activation function into the trained decoder model. The trained decoder model may output the decompressed CSI, which is a reconstruction of the CSI input to the encoder model (e.g., at the UE 120).

[0136] As a result, the encoder model and the decoder model may be trained using the forward propagation path and the backward propagation path in the same training loop (e.g., in a manner similar to type 2 training), while also being trained sequentially and / or separately. For example, the function provided by the first device to the second device may enable the forward propagation path and the backward propagation path (e.g., which are fixed at the first device) to be simulated at the second device for simulated joint or concurrent training. This may improve the accuracy of the training of the encoder and / or decoder model (e.g., by training concurrently and in the same loop for forward propagation and backward propagation). Additionally, this may increase the flexibility regarding the timing at which the training occurs (e.g., because the encoder model and the decoder model may be trained separately and / or at different times). For example, a training session may not be established between the UE server and the network server to jointly train the encoder model and the decoder model.

[0137] As indicated above,Figure 10 are provided as examples. Other examples may be different from those Figure 10 described.

[0138] Figure 11 is a diagram illustrating Example 1100 associated with hybrid order training for encoder and decoder models according to the present disclosure. As Figure 10 shown, network node 110 (e.g., a base station, CU, DU, and / or RU) may communicate with UE 120. In some aspects, network node 110 and UE 120 may be part of a wireless network (e.g., wireless network 100). UE 120 and network node 110 may have established a wireless connection Figure 11 before the operations shown. As Figure 11 shown, UE 120 may communicate with UE server 1005 in a similar manner as described above in connection with Figure 10 . Similarly, network node 110 may communicate with network server 1010 in a similar manner as described above in connection with Figure 10 .

[0139] As indicated by reference numeral 1105, UE server 1005 may train an encoder model associated with UE 120. For example, UE server 1005 may train the encoder model in a similar manner as described elsewhere herein in connection with Type 1 training and / or Type 3 training. For example, UE server 1005 may receive from UE 120 one or more data sets to be used as inputs to train the encoder model. For example, the one or more data sets may include CSI. UE server 1005 may deploy a decoder model at UE server 1005. The decoder model may be configured to output a reconstructed CSI (e.g., V 输出 ) based on an input to an activation function (e.g., Z). UE server 1005 may provide a ground truth (e.g., V 输入 ) as an input to the encoder model. The encoder model may output an activation function (e.g., Z). UE server 1005 may input the activation function into the decoder model. The decoder model may output a reconstruction of the ground truth (e.g., V 输出 ). UE server 1005 may use a loss function to compare V 输出 with V 输入Compare and determine the gradient. The UE server 1005 can use this gradient to update one or more weights of the encoder model (e.g., to minimize a loss function). For example, the UE server 1005 can use this gradient to update one or more weights of the neural network of the encoder model (e.g., attempting to minimize a loss function). The UE server 1005 can perform one or more training loops in a similar manner to update the weights of the encoder model until the output of the loss function meets a training threshold. For example, the UE server 1005 can perform one or more training loops until the difference between the output V 输出 between fully reconstructs the input or ground truth V provided to the encoder model 输入 .

[0140] As shown by reference numeral 1110, the UE server 1005 can generate a function based on the trained encoder model. The function can be an API, an instruction set, code, a software program, and / or another function. The function can be configured to output an activation function (e.g., Z) based on the input of an input function (e.g., ground truth, V 输入 ). For example, when executed by a device (such as the network server 1010), the function can be configured to mimic or simulate the forward and backward propagation paths of the trained encoder model. For example, based on the training of the decoder model, the network server 1010 can configure the function to simulate the forward and backward propagation paths of the decoder model using the information obtained via the training loop and / or based on the loss function. For example, when executed by a device, the function can be configured to accept the ground truth (e.g., V 输入 ) as input and return the activation function (e.g., Z) as output (e.g., which can be used as input for training the decoder model, as described in more detail elsewhere herein). In other words, the function can be configured to provide the forward propagation path results associated with the trained encoder model (e.g., for the training loop). Thus, the decoder model can be trained using the forward and backward propagation paths in the same training loop (e.g., in a manner similar to type 2 training), while also being trained sequentially and / or separately, which is different from type 2 training.

[0141] In some aspects, the UE server 1005 can generate the function based on the trained encoder model. For example, the UE server 1005 can determine, during the training process of the encoder model, from the ground truth (e.g., V 输入)Obtained activation function. The UE server 1005 can configure this function to provide a given activation function based on a given true value input to the function (e.g., using information obtained via a training loop and / or based on a loss function, such that the decoder model can use the forward propagation path and the backward propagation path in the same training loop as the encoder model (e.g., in a manner similar to type 2 training), while also enabling the decoder model and the encoder model to be trained sequentially and / or individually). Additionally or alternatively, this function can be pre-configured (e.g., by a vendor associated with the UE server 1005). In such examples, the UE server 1005 can obtain this function from the memory of the UE server 1005.

[0142] In some aspects, the UE server 1005 (and / or the UE 120) can determine a codebook for vector quantization associated with the compressed CSI, as described in more detail elsewhere in this document (such as in conjunction with Figure 9A and Figure 9B ). For example, the UE server 1005 (and / or the UE 120) can train a vector quantization model as part of training the encoder model. For example, the UE server 1005 (and / or the UE 120) can train a quantizer associated with vector quantization. As part of the training of the encoder model, a quantization codebook (e.g., a vector codebook or a scalar codebook) can be determined at the UE server 1005 (and / or the UE 120). In some aspects, a function (e.g., an API or other function) generated by the UE server 1005 can include a vector quantization component. For example, this function can be configured to simulate the effect of quantization on the activation or other information output by the trained encoder model or input to the trained decoder model.

[0143] In some aspects, the UE server 1005 can generate a function associated with multiple encoder models. For example, the UE server 1005 can train multiple encoder models (e.g., in a manner similar to that described in more detail elsewhere in this document). In some aspects, these multiple encoder models can be associated with corresponding network node vendors. As another example, these multiple encoder models can be associated with corresponding types of CSI (e.g., the first encoder model can be associated with a precoding vector, the second encoder model can be associated with a channel estimate, etc.). As another example, these multiple encoder models can be associated with corresponding channel conditions. As another example, these multiple encoder models can be associated with corresponding CSI sizes (e.g., the size of the CSI to be communicated between the network node 110 and the UE 120). The UE server 1005 can generate a function configured to simulate the forward propagation path and the backward propagation path of these multiple trained encoder models.

[0144] As indicated by reference numeral 1115, the UE server 1005 may transmit and the network server 1010 may receive the function (e.g., which is associated with a trained encoder model). For example, the network server 1010 and the UE server 1005 may establish a connection (e.g., a wireless connection or a wired connection). The function may be transmitted from the UE server 1005 to the network server 1010 via the connection.

[0145] As indicated by reference numeral 1120, the network server 1010 may use the function to train a decoder model. For example, the network server 1010 may train the decoder model by selecting or updating one or more weights associated with the decoder model based on one or more gradients obtained from a loss function, as described in more detail elsewhere herein. The one or more gradients may be obtained based on inputting one or more input functions (e.g., ground truth) into the function. For example, the network server 1010 may train the decoder model in a manner similar to type 2 training. However, instead of receiving the one or more activation functions and one or more input functions (e.g., ground truth) from the UE server 1005 (e.g., as in the case of type 2 training), the network server 1010 may obtain one or more activation functions and / or one or more input functions (e.g., ground truth) based on the function received from the UE server 1005. The function may simulate the forward propagation path and the backward propagation path of the trained encoder model. Thus, the decoder model may be trained using the forward propagation path and the backward propagation path in the same training loop as the encoder model (e.g., in a manner similar to type 2 training), while also being trained sequentially and / or separately, which is different from type 2 training.

[0146] In some aspects, as described above, the function may include a vector quantization component. For example, the function may be configured to simulate the effect of quantization on the activations or other information output by or input to the trained decoder model. In such examples, the network server 1010 may use the function to train a quantizer and / or a vector quantization model. In other examples, the network server 1010 (or network node 110) may determine a codebook for vector quantization associated with the compressed CSI, as described in more detail elsewhere herein (such as in connection with Figure 9A and Figure 9B)。For example, the network server 1010 (and / or the network node 110) may train a vector quantization model as part of training an encoder model. For example, the network server 1010 (and / or the network node 110) may train a quantizer associated with vector quantization. As part of the training of the decoder model, a quantization codebook (e.g., a vector codebook or a scalar codebook) may be determined at the network server 1010 (and / or the network node 110). In such examples, the input provided to a function (e.g., an API) may include quantized activations output by the function (e.g., which are quantized using vector quantization and / or the quantization codebook determined by the network server 1010 and / or the network node 110). In other words, the quantizer may be trained with the decoder model (e.g., and the function may not simulate the effect of such quantization).

[0147] In some aspects, as described above, the function may be associated with multiple trained encoder models. In such examples, training the encoder model may include providing an indication of an identifier associated with the decoder model and / or the encoder model (from multiple trained encoder models) to the function. For example, the input to the function (e.g., an API) may include a model identifier (e.g., which is associated with a given encoder model and / or decoder model). The function may be configured to provide information based on the model identifier provided to the function. In some examples, the network server 1010 may train a single decoder model to be operable with each of the multiple trained encoder models. In other examples, the network server 1010 may train multiple decoder models to be operable with corresponding encoder models from the multiple trained encoder models (e.g., if the function is associated with N trained encoder models, the network server 1010 may train N decoder models).

[0148] In some aspects, the network server 1010 may receive another function (e.g., a second function) from another UE server (e.g., another UE server 1005). For example, the other UE server may be associated with a UE vendor different from the vendor associated with the UE server 1005. The network server 1010 may use the first function (e.g., the first function received from the UE server 1005) and the second function (e.g., the second function received from another UE server) to train the decoder model. In other words, the network server 1010 may use multiple functions provided by UE servers associated with different vendors to train the decoder model. In this way, the trained decoder model may be configured to operate with the trained encoders associated with the multiple functions (e.g., in a manner similar to that described in conjunction with Figure 6 the description).

[0149] As indicated by reference numeral 1125, the UE server 1005 may transmit and the UE 120 may receive an indication of a trained encoder model. For example, the UE 120 may download the trained encoder model from the UE server 1005. Similarly, as indicated by reference numeral 1130, the network server 1010 may transmit and the network node 110 may receive an indication of a trained decoder model (e.g., which is trained using a function associated with the encoder model of the UE 120). For example, the network node 110 may download the trained decoder model from the network server 1010.

[0150] As indicated by reference numeral 1135, the UE 120 and the network node 110 may communicate using the trained encoder model and the trained decoder model, respectively. For example, the UE 120 may obtain CSI to be transmitted to the network node 110. The UE 120 may input the CSI into the trained encoder model. The trained encoder model may output an activation function (e.g., the compressed CSI). In some aspects, the UE 120 may quantize the activation function output by the trained encoder model (e.g., using a quantization codebook and / or vector quantization). The UE 120 may transmit and the network node 110 may receive the activation function output by the trained encoder model (e.g., the compressed CSI). In some aspects, the UE 120 may transmit and the network node 110 may receive a quantized representation of the activation function output by the trained encoder model (e.g., the compressed CSI). The network node 110 may input the activation function into the trained decoder model. The trained decoder model may output the decompressed CSI, which is a reconstruction of the CSI input to the encoder model (e.g., at the UE 120).

[0151] As a result, the encoder model and the decoder model may be trained using forward and backward propagation paths in the same training loop (e.g., in a manner similar to type 2 training), while also being trained sequentially and / or separately. For example, a function provided by a first device may enable the forward and backward propagation paths (e.g., which are fixed at the first device) to be simulated at a second device for simulated joint or concurrent training. This may improve the accuracy of training of the encoder and / or decoder model (e.g., by training concurrently and in the same loop for forward and backward propagation). Additionally, this may increase the flexibility regarding the timing at which training occurs (e.g., because the encoder model and the decoder model may be trained separately and / or at different times). For example, a training session may not be established between the UE server and the network server to jointly train the encoder model and the decoder model.

[0152] As indicated above, Figure 11is provided as an example. Other examples may be different from the example Figure 11 described.

[0153] Figure 12 is a diagram illustrating an example process 1200 performed, for example, by a first device in accordance with the present disclosure. The example process 1200 is an example in which a first device (e.g., a server, UE server 1005, network server 1010, UE 120, and / or network node 110) performs operations associated with hybrid-order training for encoder and decoder models.

[0154] As Figure 12 shown, in some aspects, process 1200 may include: receiving, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model (block 1210). For example, the first device (e.g., using the Figure 14 communication manager 140 and / or receiving component 1402 depicted in) may receive, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model, as described above.

[0155] As Figure 12 further shown, in some aspects, process 1200 may include: training a second model based on selecting one or more weights associated with the second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function (block 1220). For example, the first device (e.g., using the Figure 14 communication manager 140 and / or model training component 1408 depicted in) may train a second model based on selecting one or more weights associated with the second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function, as described above.

[0156] Process 1200 may include additional aspects, such as any individual aspect or any combination of aspects described below and / or in connection with one or more other processes described elsewhere herein.

[0157] In a first aspect, process 1200 includes transmitting the second model to a UE or network node after training the second model.

[0158] In a second aspect, alone or in combination with the first aspect, the second model is configured to output compressed CSI, the one or more activations include compressed CSI, and the trained first model is configured to output CSI based on an input of the compressed CSI, the one or more inputs include CSI.

[0159] In a third aspect, either alone or in combination with one or more of the first and second aspects, process 1200 includes training a vector quantization model using the one or more gradients.

[0160] In a fourth aspect, either alone or in combination with one or more of the first to third aspects, the function is configured to perform vector quantization associated with the output of the function.

[0161] In a fifth aspect, either alone or in combination with one or more of the first to fourth aspects, the function is associated with a plurality of trained first models, and training the second model includes providing an identifier associated with the trained first model as an input to the function.

[0162] In a sixth aspect, either alone or in combination with one or more of the first to fifth aspects, training the second model further includes training the second model to be configured to operate with each of the plurality of trained first models.

[0163] In a seventh aspect, either alone or in combination with one or more of the first to sixth aspects, training the second model further includes training a plurality of second models (including the second model) to be configured to operate with corresponding trained first models from the plurality of trained first models.

[0164] In an eighth aspect, either alone or in combination with one or more of the first to seventh aspects, the function is a first function, and process 1200 includes: receiving an indication of a second function associated with another trained first model from a third device, and training the second model includes using the first function and the second function to train the second model.

[0165] In a ninth aspect, either alone or in combination with one or more of the first to eighth aspects, the function is an API.

[0166] In a tenth aspect, either alone or in combination with one or more of the first to ninth aspects, the first device is a server associated with a UE, the trained first model is a decoder model, and the second model is an encoder model (e.g., in a similar manner as depicted and described in connection with Figure 10 )). In some aspects, the function may be configured to use an activation function (e.g., Z) and a ground truth (e.g., V 输入 ) as inputs, and the function may output one or more gradients (e.g., to simulate the forward and backward propagation paths of the decoder model). The one or more gradients may be used to update one or more weights of the encoder model.

[0167] In an eleventh aspect, either alone or in combination with one or more of the first to tenth aspects, the first device is a server associated with a network node, the trained first model is an encoder model, and the second model is a decoder model (e.g., in a similar manner as depicted and described in conjunction with Figure 11 ). In some aspects, the function may use a ground truth (e.g., V 输入 ) as an input, and the function may output an activation function (e.g., Z). The output of the function (e.g., activation function Z) may be used as an input to the decoder model to train the decoder model.

[0168] In a twelfth aspect, either alone or in combination with one or more of the first to eleventh aspects, the first device is a UE or a network node.

[0169] In a thirteenth aspect, either alone or in combination with one or more of the first to twelfth aspects, the function is configured to simulate a forward propagation path and a backward propagation path of the trained first model based on the one or more gradients.

[0170] Although Figure 12 illustrates example blocks of process 1200, in some aspects, process 1200 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to the blocks depicted in Figure 12 . Additionally or alternatively, two or more blocks of process 1200 may be executed in parallel.

[0171] Figure 13 is a diagram illustrating an example process 1300 performed, for example, by a first device in accordance with the present disclosure. Example process 1300 is an example in which a first device (e.g., a server, UE server 1005, network server 1010, UE 120, and / or network node 110) performs operations associated with hybrid order training for an encoder and a decoder model.

[0172] As Figure 13 shown, in some aspects, process 1300 may include: training a first model based on one or more inputs to obtain a trained first model, the trained first model being associated with one or more activations associated with an output of the trained first model (block 1310). For example, the first device (e.g., using the communication manager 150 and / or the model training component 1508 depicted in Figure 15 ) may train a first model based on one or more inputs to obtain a trained first model, the trained first model being associated with one or more activations associated with an output of the trained first model, as described above.

[0173] As Figure 13As further shown, in some aspects, process 1300 may include: transmitting to a second device a function associated with a trained first model, the function being configured to output one or more activations based on ground truth inputs (block 1320). For example, a first device (e.g., using Figure 15 the communication manager 150 and / or the transmission component 1504 depicted in) may transmit to a second device a function associated with a trained first model, the function being configured to output one or more activations based on ground truth inputs, as described above.

[0174] Process 1300 may include additional aspects, such as any single aspect or any combination of aspects described below and / or in connection with one or more other processes described elsewhere herein.

[0175] In a first aspect, process 1300 includes transmitting the trained first model to a UE or a network node after training the first model.

[0176] In a second aspect, either alone or in combination with the first aspect, the trained first model is configured to output compressed CSI or to output CSI based on an input of compressed CSI.

[0177] In a third aspect, either alone or in combination with one or more of the first and second aspects, process 1300 includes using the trained first model to train a vector quantization model.

[0178] In a fourth aspect, either alone or in combination with one or more of the first to third aspects, the function is configured to perform vector quantization associated with the output of the function.

[0179] In a fifth aspect, either alone or in combination with one or more of the first to fourth aspects, the function is an API.

[0180] In a sixth aspect, either alone or in combination with one or more of the first to fifth aspects, the first device is a server associated with a network node, the first model is a decoder model, and the second device is associated with a UE and an encoder model (e.g., in a similar manner as depicted and described in connection with Figure 10 ). In some aspects, the function may be configured to use an activation function (e.g., Z) and a ground truth (e.g., V 输入 ) as inputs, and the function may output one or more gradients (e.g., to simulate the forward and backward propagation paths of the decoder model). The one or more gradients may be used to update one or more weights of the encoder model.

[0181] In a seventh aspect, either alone or in combination with one or more of the first to sixth aspects, the first device is a server associated with a UE, the first model is an encoder model, and the second device is associated with a network node and a decoder model (e.g., in a similar manner as depicted and described in connection with Figure 11 ). In some aspects, the function may use a ground truth (e.g., V 输入 ) as an input, and the function may output an activation function (e.g., Z). The output of the function (e.g., activation function Z) may be used as an input to the decoder model to train the decoder model.

[0182] In an eighth aspect, either alone or in combination with one or more of the first to seventh aspects, the first device is a network node or a UE.

[0183] In a ninth aspect, either alone or in combination with one or more of the first to eighth aspects,

[0184] the function is configured to simulate the forward propagation path and the backward propagation path of the first model.

[0185] Although Figure 13 illustrates example blocks of process 1300, in some aspects, process 1300 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to the blocks depicted in Figure 13 . Additionally or alternatively, two or more blocks of process 1300 may be executed in parallel.

[0186] Figure 14 is a diagram of an example apparatus 1400 for wireless communication in accordance with the present disclosure. Apparatus 1400 may be the first device, or the first device may include apparatus 1400. In some aspects, the first device may be a server, UE server 1005, network server 1010, UE 120, and / or network node 110. In some aspects, apparatus 1400 includes a receiving component 1402 and a transmitting component 1404, which may be in communication with each other (e.g., via one or more buses and / or one or more other components). As shown, apparatus 1400 may use the receiving component 1402 and the transmitting component 1404 to communicate with another device 1406 (such as a UE, a base station, or another wireless communication device). As further shown, apparatus 1400 may include a communication manager 140. The communication manager 140 may include a model training component 1408, among others.

[0187] In some aspects, apparatus 1400 may be configured to perform one or more operations described herein in connection with Figure 10 and 11 . Additionally or alternatively, apparatus 1400 may be configured to perform one or more processes described herein (such as Figure 12In some aspects, Figure 14 The apparatus 1400 and / or one or more components shown in FIG. 1 may include a combination of Figure 2 Additionally or alternatively, Figure 14 One or more of the components shown may be combined in Figure 2 Additionally or alternatively, one or more components in the component set may be implemented at least in part as software stored in a memory. For example, a component (or a portion of a component) may be implemented as instructions or code stored in a non-transitory computer-readable medium and may be executed by a controller or processor to perform the function or operation of the component.

[0188] The receiving component 1402 may receive communications (such as reference signals, control information, data communications, or combinations thereof) from the device 1406. The receiving component 1402 may provide the received communications to one or more other components of the device 1400. In some aspects, the receiving component 1402 may perform signal processing (such as filtering, amplification, demodulation, analog-to-digital conversion, demultiplexing, deinterleaving, demapping, equalization, interference cancellation, or decoding, etc.) on the received communications and may provide the processed signals to one or more other components of the device 1400. In some aspects, the receiving component 1402 may include combining Figure 2 One or more antennas, modems, demodulators, MIMO detectors, receive processors, controllers / processors, memories, or combinations thereof of the described first device.

[0189] Transmission component 1404 may transmit communications (such as reference signals, control information, data communications, or a combination thereof) to device 1406. In some aspects, one or more other components of device 1400 may generate communications and may provide the generated communications to transmission component 1404 for transmission to device 1406. In some aspects, transmission component 1404 may perform signal processing (such as filtering, amplification, modulation, digital-to-analog conversion, multiplexing, interleaving, mapping, encoding, etc.) on the generated communications and may transmit the processed signals to device 1406. In some aspects, transmission component 1404 may include combining Figure 2 One or more antennas, modems, modulators, transmit MIMO processors, transmit processors, controllers / processors, memories, or combinations thereof of the described first device. In some aspects, the transmitting component 1404 can be co-located with the receiving component 1402 in a transceiver.

[0190] The receiving component 1402 may receive, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients based on an activation input and an input. The model training component 1408 may train a second model based on selecting one or more weights associated with the second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function.

[0191] The transmitting component 1404 may transmit the second model to a UE or a network node after training the second model.

[0192] The model training component 1408 may use the one or more gradients to train a vector quantization model.

[0193] Figure 14 The number and arrangement of components shown are provided as an example. In practice, there may be additional components, fewer components, different components, or differently arranged components compared to those shown in Figure 14 In addition, Figure 14 two or more components shown in Figure 14 may be implemented within a single component, or Figure 14 a single component shown in Figure 14 may be implemented as multiple distributed components. Additionally or alternatively,

[0194] Figure 15 is a diagram of an example apparatus 1500 for wireless communication in accordance with the present disclosure. The apparatus 1500 may be a first device, or the first device may include the apparatus 1500. In some aspects, the first device may be a server, a UE server 1005, a network server 1010, a UE 120, and / or a network node 110. In some aspects, the apparatus 1500 includes a receiving component 1502 and a transmitting component 1504, which may be in communication with each other (e.g., via one or more buses and / or one or more other components). As shown, the apparatus 1500 may use the receiving component 1502 and the transmitting component 1504 to communicate with another device 1506 (such as a UE, a base station, or another wireless communication device). As further shown, the apparatus 1500 may include a communication manager 150. The communication manager 150 may include one or more of a model training component 1508 and / or a function generation component 1510, etc.

[0195] In some aspects, the apparatus 1500 may be configured to perform herein in connection with Figure 10 and 11One or more operations described. Additionally or alternatively, apparatus 1500 may be configured to perform one or more processes described herein (such as Figure 13 process 1300) or combinations thereof. In some aspects, Figure 15 apparatus 1500 and / or one or more components shown in Figure 2 may include one or more components of a first device described in connection with Figure 15 One or more components shown in Figure 2 may be implemented within one or more components described in connection with

[0196] Receiving component 1502 may receive a communication (such as a reference signal, control information, data communication, or a combination thereof) from apparatus 1506. Receiving component 1502 may provide the received communication to one or more other components of apparatus 1500. In some aspects, receiving component 1502 may perform signal processing on the received communication (such as filtering, amplification, demodulation, analog-to-digital conversion, demultiplexing, deinterleaving, demapping, equalization, interference cancellation, or decoding, etc.), and may provide the processed signal to one or more other components of apparatus 1500. In some aspects, receiving component 1502 may include one or more antennas, modems, demodulators, MIMO detectors, receiving processors, controller / processors, memories, or combinations thereof of a first device described in connection with Figure 2

[0197] Transmitting component 1504 may transmit a communication (such as a reference signal, control information, data communication, or a combination thereof) to apparatus 1506. In some aspects, one or more other components of apparatus 1500 may generate a communication and may provide the generated communication to transmitting component 1504 for transmission to apparatus 1506. In some aspects, transmitting component 1504 may perform signal processing on the generated communication (such as filtering, amplification, modulation, digital-to-analog conversion, multiplexing, interleaving, mapping, coding, etc.), and may transmit the processed signal to apparatus 1506. In some aspects, transmitting component 1504 may include one or more antennas, modems, modulators, transmit MIMO processors, transmit processors, controller / processors, memories, or combinations thereof of a first device described in connection with Figure 2 In some aspects, transmitting component 1504 may be co-located with receiving component 1502 in a transceiver.

[0198] ​The model training component 1508 can train a first model based on one or more inputs to obtain a trained first model, and the trained first model is associated with one or more activations associated with the output of the trained first model. The transmission component 1504 can transmit a function associated with the trained first model to a second device, and the function is configured to output one or more activations based on a ground truth input.

[0199] The transmission component 1504 can transmit the trained first model to a UE or a network node after training the first model.

[0200] The model training component 1508 can use the trained first model to train a vector quantization model.

[0201] The function generation component 1510 can generate the function at least in part based on training the first model.

[0202] Figure 15 The number and arrangement of the components shown are provided as an example. In practice, there may be additional components, fewer components, different components, or differently arranged components compared to those shown. Additionally, Figure 15 there may be additional components, fewer components, different components, or differently arranged components compared to those shown. Additionally, Figure 15 two or more of the components shown may be implemented in a single component, or Figure 15 a single component shown may be implemented as multiple distributed components. Additionally or alternatively, Figure 15 a set of components (e.g., one or more components) shown may perform one or more functions described as being performed by Figure 15 another set of components shown.

[0203] Some aspects of the present disclosure are provided below:

[0204] Aspect 1: A wireless communication method performed by a first device, including: receiving, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model; and training a second model based on selecting one or more weights associated with the second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function. This enables the trained first model and the second model to be trained using a forward propagation path and a backward propagation path in the same training loop (e.g., in a manner similar to type 2 training), while also being trained sequentially and / or separately.

[0205] Aspect 2: The method of aspect 1, further including: transmitting the second model to a user equipment (UE) or a network node after training the second model. This increases the flexibility regarding the timing at which training occurs.

[0206] Aspect 3: The method as in any one of Aspects 1-2, wherein the second model is configured to output compressed channel state information (CSI), the one or more activations include the compressed CSI, and wherein the trained first model is configured to output CSI based on an input of the compressed CSI, the one or more inputs include CSI. This improves the accuracy of the CSI compression model (e.g., by training the models concurrently and in the same loop for forward and backward propagation).

[0207] Aspect 4: The method as in any one of Aspects 1-3, further comprising: training a vector quantization model using the one or more gradients.

[0208] Aspect 5: The method as in any one of Aspects 1-3, wherein the function is configured to perform vector quantization associated with an output of the function.

[0209] Aspect 6: The method as in any one of Aspects 1-5, wherein the function is associated with a plurality of trained first models, and wherein training the second model comprises: providing an identifier associated with the trained first model as an input to the function. This enables a single function to simulate forward and backward propagation paths for multiple trained models, thereby saving resources that would otherwise be used to configure, transmit, and / or use multiple functions for the multiple trained models.

[0210] Aspect 7: The method as in Aspect 6, wherein training the second model further comprises: training the second model to be configured to operate with each of the plurality of trained first models. This enables the second model to be trained to operate with the plurality of trained models, thereby saving resources that would otherwise be used to configure, transmit, and / or use multiple models for the multiple trained models.

[0211] Aspect 8: The method as in Aspect 6, wherein training the second model further comprises: training a plurality of second models including the second model to be configured to operate with corresponding trained first models from the plurality of trained first models.

[0212] Aspect 9: The method as in any one of Aspects 1-8, wherein the function is a first function, the method further comprising: receiving an indication of a second function associated with another trained first model from a third device, and wherein training the second model comprises: using the first function and the second function to train the second model.

[0213] Aspect 10: The method as in any one of Aspects 1-9, wherein the function is an application programming interface (API).

[0214] Aspect 11: A method as in any one of aspects 1 - 10, wherein the first device is a server associated with a user equipment (UE), wherein the trained first model is a decoder model, and wherein the second model is an encoder model.

[0215] Aspect 12: A method as in any one of aspects 1 - 10, wherein the first device is a server associated with a network node, wherein the trained first model is an encoder model, and wherein the second model is a decoder model.

[0216] Aspect 13: A method as in any one of aspects 1 - 10, wherein the first device is a user equipment (UE) or a network node.

[0217] Aspect 14: A wireless communication method performed by a first device, comprising: training a first model based on one or more inputs to obtain a trained first model, the trained first model being associated with one or more activations associated with an output of the trained first model; and transmitting to a second device a function associated with the trained first model, the function being configured to output one or more activations based on a ground truth input.

[0218] Aspect 15: The method of aspect 14, further comprising: transmitting the trained first model to a user equipment (UE) or a network node after training the first model.

[0219] Aspect 16: The method of any one of aspects 14 - 15, wherein the trained first model is configured to output compressed channel state information (CSI) or output CSI based on an input of the compressed CSI.

[0220] Aspect 17: The method of any one of aspects 14 - 16, further comprising: using the trained first model to train a vector quantization model.

[0221] Aspect 18: The method of any one of aspects 14 - 16, wherein the function is configured to perform vector quantization associated with an output of the function.

[0222] Aspect 19: The method of any one of aspects 14 - 18, wherein the function is an application programming interface (API).

[0223] Aspect 20: The method of any one of aspects 14 - 19, wherein the first device is a server associated with a network node, wherein the first model is a decoder model, and wherein the second device is associated with a user equipment (UE) and an encoder model.

[0224] Aspect 21: A method as in any one of aspects 14 - 19, wherein the first device is a server associated with a user equipment (UE), wherein the first model is an encoder model, and wherein the second device is associated with a network node and a decoder model.

[0225] Aspect 22: A method as in any one of aspects 14 - 19, wherein the first device is a network node or a user equipment (UE).

[0226] Aspect 23: An apparatus for wireless communication at a device, comprising: one or more processors; one or more memories coupled to the one or more processors; and instructions stored in the one or more memories and executable by the one or more processors to cause the device to perform a method as in one or more of aspects 1 to 13.

[0227] Aspect 24: A device for wireless communication, comprising one or more memories and one or more processors coupled to the one or more memories, the one or more processors configured to perform a method as in one or more of aspects 1 to 13.

[0228] Aspect 25: A device for wireless communication, comprising at least one means for performing a method as in one or more of aspects 1 - 13.

[0229] Aspect 26: A non - transient computer - readable medium storing code for wireless communication, the code comprising instructions executable by one or more processors to perform a method as in one or more of aspects 1 to 13.

[0230] Aspect 27: A non - transient computer - readable medium storing an instruction set for wireless communication, the instruction set comprising one or more instructions that, when executed by one or more processors of a device, cause the device to perform a method as in one or more of aspects 1 - 13.

[0231] Aspect 28: An apparatus for wireless communication at a device, comprising: one or more processors; one or more memories coupled to the one or more processors; and instructions stored in the one or more memories and executable by the one or more processors to cause the device to perform a method as in one or more of aspects 14 to 22.

[0232] Aspect 29: A device for wireless communication, comprising one or more memories and one or more processors coupled to the one or more memories, the one or more processors configured to perform a method as in one or more of aspects 14 to 22.

[0233] Aspect 30: A device for wireless communication, comprising at least one means for performing the method of one or more of Aspects 14 - 22.

[0234] Aspect 31: A non-transitory computer-readable medium storing code for wireless communication, the code comprising instructions executable by one or more processors to perform the method of one or more of Aspects 14 to 22.

[0235] Aspect 32: A non-transitory computer-readable medium storing a set of instructions for wireless communication, the set of instructions comprising one or more instructions that, when executed by one or more processors of a device, cause the device to perform the method of one or more of Aspects 14 - 22.

[0236] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the aspects to the precise forms disclosed. Modifications and variations may be made in light of the above disclosure or may be obtained by practicing the aspects.

[0237] As used herein, the term "component" is intended to be broadly construed as a combination of hardware and / or hardware and software. "Software" shall be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, procedures, and / or functions, etc., whether referred to in software, firmware, middleware, microcode, hardware description language, or other terms. As used herein, a "processor" is implemented with hardware and / or a combination of hardware and software. It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware and / or a combination of hardware and software. The actual specific control hardware or software code used to implement these systems and / or methods does not limit the aspects. Thus, the operation and behavior of these systems and / or methods are described herein without reference to specific software code, as those skilled in the art will understand that software and hardware can be designed to implement these systems and / or methods at least in part based on the description herein.

[0238] As used herein, depending on the context, "meeting a threshold" may mean a value greater than a threshold, greater than or equal to a threshold, less than a threshold, less than or equal to a threshold, equal to a threshold, not equal to a threshold, etc.

[0239] Although specific feature combinations are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of the various aspects. Many of these features may be combined in ways not specifically recited in the claims and / or not disclosed in the specification. The disclosure of the various aspects includes each dependent claim in combination with every other claim in the claim set. As used herein, a phrase that recites "at least one" of a list of items means any combination of these items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, a + b, a + c, b + c, and a + b + c, as well as any combination having multiple identical elements (e.g., a + a, a + a + a, a + a + b, a + a + c, a + b + b, a + c + c, b + b, b + b + b, b + b + c, c + c, and c + c + c, or any other ordering of a, b, and c).

[0240] Elements, acts, or instructions used herein should not be construed as critical or essential unless expressly described as such. Also, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Further, as used herein, the article "the" is intended to include one or more items referred to in conjunction with the article "the" and may be used interchangeably with "one or more." Additionally, as used herein, the terms "set" and "group" are intended to include one or more items and may be used interchangeably with "one or more." In instances where only one item is intended, the phrase "only one" or similar language is used. Also, as used herein, the terms "having," "containing," "including," etc. are intended to be open-ended terms that do not limit the elements they modify (e.g., an element "having" A may also have B). Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise expressly stated. Also, as used herein, the term "or" when used in a series is intended to be inclusive and may be used interchangeably with "and / or" unless otherwise expressly stated (e.g., when used in conjunction with "any of" or "only one of").

Claims

1. A first device for wireless communication, comprising: one or more memories; and one or more processors coupled to the one or more memories, the one or more processors being configured to: receive, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model; and train the second model based on selecting one or more weights associated with a second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function.

2. The first device according to claim 1, wherein the one or more processors are further configured to: transmit the second model to a user equipment (UE) or a network node after training the second model.

3. The first device according to claim 1, wherein the second model is configured to output compressed channel state information (CSI), the one or more activations include the compressed CSI, and wherein the trained first model is configured to output CSI based on an input of the compressed CSI, the one or more inputs include the CSI.

4. The first device according to claim 1, wherein the one or more processors are further configured to: train a vector quantization model using the one or more gradients.

5. The first device according to claim 1, wherein the function is configured to perform vector quantization associated with an output of the function.

6. The first device according to claim 1, wherein the function is associated with a plurality of trained first models, and wherein, for training the second model, the one or more processors are configured to: provide an identifier associated with the trained first model as an input to the function.

7. The first device according to claim 6, wherein, for training the second model, the one or more processors are configured to: train the second model to be configured to operate with each of the plurality of trained first models.

8. The first device according to claim 6, wherein, for training the second model, the one or more processors are configured to: train a plurality of second models including the second model to be configured to operate with a respective trained first model from the plurality of trained first models.

9. The first device according to claim 1, wherein the function is a first function, wherein the one or more processors are further configured to: receive an indication of a second function associated with another trained first model from a third device, and wherein, for training the second model, the one or more processors are configured to: train the second model using the first function and the second function.

10. The first device according to claim 1, wherein the function is an application programming interface (API).

11. The first device according to claim 1, wherein the first device is a server associated with a user equipment (UE), wherein the trained first model is a decoder model, and wherein the second model is an encoder model.

12. The first device according to claim 1, wherein the first device is a user equipment (UE) or a network node.

13. The first device according to claim 1, wherein the function is configured to simulate a forward propagation path and a backward propagation path of the trained first model based on the one or more gradients.

14. A first device for wireless communication, comprising: one or more memories; and one or more processors coupled to the one or more memories, the one or more processors being configured to: train a first model based on one or more inputs to obtain a trained first model, an output of the trained first model being associated with one or more activations; and transmit a function associated with the trained first model to a second device, the function being configured to output one or more activations based on a ground truth input.

15. The first device according to claim 14, wherein the one or more processors are further configured to: transmit the trained first model to a user equipment (UE) or a network node after training the first model.

16. The first device according to claim 14, wherein the trained first model is configured to output compressed channel state information (CSI) or output CSI based on an input according to the compressed CSI.

17. The first device according to claim 14, wherein the one or more processors are further configured to: train a vector quantization model using the trained first model.

18. The first device according to claim 14, wherein the function is configured to perform vector quantization associated with an output of the function.

19. The first device according to claim 14, wherein the first device is a first server associated with a network node, wherein the first model is a decoder model, and wherein the second device is a second server associated with a user equipment (UE).

20. A wireless communication method performed by a first device, comprising: receiving, from a second device, a function associated with a trained first model, the function being configured to output one or more gradients associated with the trained first model; and training a second model based on selecting one or more weights associated with the second model using the one or more gradients, the one or more gradients being obtained based on inputting one or more activations and one or more inputs into the function.

21. The method according to claim 20, further comprising: transmitting the second model to a user equipment (UE) or a network node after training the second model.

22. The method according to claim 20, wherein the second model is configured to output compressed channel state information (CSI), the one or more activations include the compressed CSI, and Wherein the trained first model is configured to output CSI based on an input of the compressed CSI, and the one or more inputs include the CSI.

23. The method of claim 20, further comprising: training a vector quantization model using the one or more gradients.

24. The method of claim 20, wherein the function is configured to perform vector quantization associated with an output of the function.

25. The method of claim 20, wherein the function is associated with a plurality of trained first models, and wherein training the second model comprises: providing an identifier associated with the trained first model as an input to the function.

26. The method of claim 25, wherein training the second model further comprises: training the second model to be configured to operate with each of the plurality of trained first models.

27. The method of claim 25, wherein training the second model further comprises: training a plurality of second models including the second model to be configured to operate with a corresponding trained first model from the plurality of trained first models.

28. A wireless communication method performed by a first device, comprising: training a first model based on one or more inputs to obtain a trained first model, the trained first model being associated with one or more activations associated with an output of the trained first model; and transmitting, to a second device, a function associated with the trained first model, the function being configured to output one or more activations based on a ground truth input.

29. The method of claim 28, further comprising: transmitting the trained first model to a user equipment (UE) or a network node after training the first model.

30. The method of claim 28, wherein the trained first model is configured to output compressed channel state information (CSI) or output CSI based on an input of the compressed CSI.