Obtaining parameters for a learning model
By employing learning models with a feature and classifier component structure, the inefficiencies in wireless communication systems are addressed, reducing resource usage and optimizing parameter transmission for different domains.
Patent Information
- Application Number
- PCT/IB2025/052958
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-04
- Filing Date
- 2025-03-20
- Publication Date
- 2025-08-21
AI Technical Summary
Existing wireless communication systems face inefficiencies in signaling overhead and memory storage due to the transmission and storage of multiple learning models trained for different domains, which have a large numerical quantity of parameters, leading to increased resource usage.
Implement learning models with a feature component and a classifier component, where the feature component transforms input data into features and the classifier component identifies patterns, allowing for the sharing of parameters across different domains, reducing the need to transmit full sets of parameters for each model.
This approach reduces signaling overhead and processing resources by minimizing the transmission of parameters for the classifier component, thereby optimizing time-frequency resource usage and memory storage.
Smart Images

Figure IB2025052958_21082025_PF_FP_ABST
Abstract
Description
Lenovo Ref. No. SMM920240006-WO-PCT 1 OBTAINING PARAMETERS FOR A LEARNING MODEL RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Application Serial No. 63 / 574,830 filed April 4, 2024, entitled “OBTAINING PARAMETERS FOR A LEARNING MODEL,” the disclosure of which is incorporated by reference herein in its entirety. TECHNICAL FIELD
[0002] The present disclosure relates to wireless communications, and more specifically to learning model techniques for classification. BACKGROUND
[0003] A wireless communications system may include one or multiple network communication devices, which may be otherwise known as network equipment (NE), supporting wireless communications for one or multiple user communication devices, which may be otherwise known as user equipment (UE), or other suitable terminology. The wireless communications system may support wireless communications with one or multiple user communication devices by utilizing resources of the wireless communication system (e.g., time resources (e.g., symbols, slots, subframes, frames, or the like)) or frequency resources (e.g., subcarriers, carriers, or the like). Additionally, the wireless communications system may support wireless communications across various radio access technologies including third generation (3G) radio access technology, fourth generation (4G) radio access technology, fifth generation (5G) radio access technology, among other suitable radio access technologies beyond 5G (e.g., sixth generation (6G)). SUMMARY
[0004] An article “a” before an element is unrestricted and understood to refer to “at least one” of those elements or “one or more” of those elements. The terms “a,” “at least one,” “one or more,” and “at least one of one or more” may be interchangeable. As used herein, including in the claims, “or” as used in a list of items (e.g., a list of items prefaced by a phrase such as “at least one of” or “one or more of” or “one or both of”) indicates an inclusive list such that, for example, a list of at least one of A, B, or C means A or B or C or AB or AC or BC or ABC (i.e., A and B and C). Also, Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 2 as used herein, the phrase “based on” shall not construed as a reference to a closed set of conditions. For example, an example step that is described as “based on condition A” may be based on both a condition A and a condition B without departing from the scope of the present disclosure. In other words, as used herein, the phrase “based on” shall be construed in the same manner as the phrase “based at least in part on.” Further, as used herein, including in the claims, a “set” may include one or more elements.
[0005] A first device for wireless communication is described. The first device may be configured to, capable of, or operable to perform one or more operations as described herein. For example, the first device may be configured to, capable of, or operable to obtain a set of learning models, where each learning model of the set of learning models is associated with a first set of one or more parameters, each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and select, for executing at the first device, at least one learning model of the set of learning models or transmit, to a second device, a message including the first set of one or more parameters and at least one second set of one or more parameters of the set of second sets of one or more parameters.
[0006] A processor (e.g., a standalone processor chipset, or a component of a first device) for wireless communication is described. The processor may be configured to, capable of, or operable to perform one or more operations as described herein. For example, the processor may be configured to, capable of, or operable to obtain a set of learning models, where each learning model of the set of learning models is associated with a first set of one or more parameters, each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and select, for executing at the first device, at least one learning model of the set of learning models or transmit, to a second device, a message including the first set of one or more parameters and at least one second set of one or more parameters of the set of second sets of one or more parameters.
[0007] A method performed or performable by a first device for wireless communication is described. The method may include obtaining a set of learning models, where each learning model of the set of learning models is associated with a first set of one or more parameters, each learning Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 3 model of the set of learning models includes a set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and selecting, for executing at the first device, at least one learning model of the set of learning models or transmitting, to a second device, a message including the first set of one or more parameters and at least one second set of one or more parameters of the set of second sets of one or more parameters.
[0008] In some implementations of the first device, the processor, and the method described herein, the subset of layers includes at least one layer associated with classification, an additional subset of layers of the respective set of layers includes a set of layers associated with feature extraction, and the set of layers associated with the feature extraction correspond to the first set of one or more parameters. In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to select the at least one learning model of the set of learning models, and obtain, as output from the set of layers associated with the feature extraction, one or more attributes of a data sample based on providing the data sample as input to the at least one learning model of the set of learning models. In some implementations of the first device, the processor, and the method described herein, to obtain the set of learning models, the first device, the processor, and the method may further be configured to, capable of, or operable to train, by minimizing a loss function using a subset of data samples of a set of data samples, a learning model of the set of learning models to determine the first set of one or more parameters and an initial second set of one or more parameters of the set of second sets of one or more parameters, where the trained learning model is associated with the first set of one or more parameters and the initial second set of one or more parameters, and train, by minimizing the loss function using respective remaining subsets of data samples of the set of data samples, one or more remaining learning models of the set of learning models to determine respective second sets of one or more parameters associated with the one or more remaining learning models, where the one or more remaining learning models are associated with the first set of one or more parameters and the respective second sets of one or more parameters.
[0009] In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 4 operable to determine the set of data samples of the second device. In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to receive an additional message that indicates the set of data samples. In some implementations of the first device, the processor, and the method described herein, the set of data samples includes a set of subsets of training data samples corresponding to each learning model of the set of learning models, and respective subsets of training data samples of the set of subsets of training data samples are associated with one or more different characteristics. In some implementations of the first device, the processor, and the method described herein, the one or more different characteristics include at least one of a mean of the respective subsets of training data samples, a standard deviation of the respective subsets of training data samples, a variance of the respective subsets of training data samples, or a probability distribution of the respective subsets of training data samples. In some implementations of the first device, the processor, and the method described herein, the set of subsets of training data samples are associated with at least one of a respective network condition between the first device and the second device, respective parameters associated with the first device, respective parameters associated with the second device, or respective channel characteristics associated with a channel between the first device and the second device. In some implementations of the first device, the processor, and the method described herein, the set of data samples includes at least one of a labeled data sample including a data sample for input to respective learning models of the set of learning models and an indication of an expected output from the respective learning models, or an unlabeled data sample including a data sample for input to the respective learning models of the set of learning models.
[0010] In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to select the at least one learning model of the set of learning models, where the at least one learning model is associated with the first set of one or more parameters and the at least one second set of one or more parameters, and generates, based on providing data as input to the at least one learning model of the set of learning models, output from the at least one learning model. In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 5 update the at least one second set of one or parameters associated with the at least one learning model to an additional second set of one or more parameters of the set of second sets of one or more parameters, where the additional second set of one or more parameters is associated with an additional learning model of the set of learning models. In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to transmit, to the second device, the message including the first set of one or more parameters and the at least one second set of one or more parameters of the set of second sets of one or more parameters, and transmit, to the second device and after the message is transmitted, an additional message including an additional second set of one or more parameters of the set of second sets of one or more parameters, where the first set of one or more parameters and the additional second set of one or more parameters are associated with an additional learning model of the set of learning models.
[0011] In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to receive, from the second device, an additional message that indicates for the first device to transmit the message, and transmit, responsive to the additional message and to the second device, the message including the first set of one or more parameters and the at least one second set of one or more parameters of the set of second sets of one or more parameters. In some implementations of the first device, the processor, and the method described herein, the additional message requests at least one of the first set of one or more parameters or the at least one second set of one or more parameters of the plurality of second sets of one or more parameters. In some implementations of the first device, the processor, and the method described herein, to transmit the message, the first device, the processor, and the method may further be configured to, capable of, or operable to determine one or more conditions are satisfied associated with the message, where the one or more conditions are associated with at least one of a network condition between the first device and the second device, one or more parameters associated with the first device, one or more parameters associated with the second device, or one or more channel characteristics associated with a channel between the first device and the second device.
[0012] A first device for wireless communication is described. The first device may be configured to, capable of, or operable to perform one or more operations as described herein. For Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 6 example, the first device may be configured to, of, or operable to transmit, to a second device, a first message that requests at least one of a first set of one or more parameters associated with respective learning models of a set of learning models or at least one second set of one or more parameters associated with at least one learning model of the set of learning models, where each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and receive, responsive to the first message and from the second device, a second message including at least one of the first set of one or more parameters or the at least one second set of one or more parameters.
[0013] A processor (e.g., a standalone processor chipset, or a component of a first device) for wireless communication is described. The processor may be configured to, capable of, or operable to perform one or more operations as described herein. For example, the processor may be configured to, capable of, or operable to transmit, to a second device, a first message that requests at least one of a first set of one or more parameters associated with respective learning models of a set of learning models or at least one second set of one or more parameters associated with at least one learning model of the set of learning models, where each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and receive, responsive to the first message and from the second device, a second message including at least one of the first set of one or more parameters or the at least one second set of one or more parameters.
[0014] A method performed or performable by a first device for wireless communication is described. The method may include transmitting, to a second device, a first message that requests at least one of a first set of one or more parameters associated with respective learning models of a set of learning models or at least one second set of one or more parameters associated with at least one learning model of the set of learning models, where each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and receiving, responsive to the first message and from the second device, a second Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 7 message including at least one of the first set of or more parameters or the at least one second set of one or more parameters.
[0015] In some implementations of the first device, the processor, and the method described herein, the subset of layers includes at least one layer associated with classification, an additional subset of layers of the respective set of layers includes a set of layers associated with feature extraction, and the set of layers associated with the feature extraction correspond to the first set of one or more parameters. In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to obtain, as output from the set of layers associated with the feature extraction, one or more attributes of a data sample based on providing the data sample as input to the at least one learning model associated with the first set of one or more parameters and the at least one second set of one or more parameters. In some implementations of the first device, the processor, and the method described herein, the at least one second set of one or more parameters is determined based on the at least one learning model being trained using a subset of data samples of a set of data samples. In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to transmit a third message that indicates the set of data samples.
[0016] In some implementations of the first device, the processor, and the method described herein, the set of data samples includes a set of subsets of training data samples corresponding to each learning model of the set of learning models, and respective subsets of training data samples of the set of subsets of training data samples are associated with one or more different characteristics. In some implementations of the first device, the processor, and the method described herein, the one or more different characteristics include at least one of a mean of the respective subsets of training data samples, a standard deviation of the respective subsets of training data samples, a variance of the respective subsets of training data samples, or a probability distribution of the respective subsets of training data samples. In some implementations of the first device, the processor, and the method described herein, the set of subsets of training data samples are associated with at least one of a respective network condition between the first device and the second device, respective parameters associated with the first device, respective parameters associated with the second device, or respective channel characteristics associated with a channel between the first device and the second Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 8 device. In some implementations of the first the processor, and the method described herein, the set of data samples includes at least one of a labeled data sample including a data sample for input to a learning model of the set of learning models and an indication of an expected output from the learning model, or an unlabeled data sample including a data sample for input to the learning model of the set of learning models.
[0017] In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to generate, based on providing data as input to the at least one learning model, output from the at least one learning model based on the at least one learning model including the first set of one or more parameters and the at least one second set of one or more parameters. In some implementations of the first device, the processor, and the method described herein, the first device, the processor, and the method may further be configured to, capable of, or operable to transmit, to the second device, a third message that requests an additional second set of one or more parameters associated with an additional learning model of the set of learning models, and receive, responsive to the third message, a fourth message including the additional second set of one or more parameters, where the additional learning model includes the first set of one or more parameters and the additional second set of one or more parameters. In some implementations of the first device, the processor, and the method described herein, to transmit the first message, the first device, the processor, and the method may further be configured to, capable of, or operable to determine one or more conditions are satisfied associated with the first message, where the one or more conditions are associated with at least one of a network condition between the first device and the second device, one or more parameters associated with the first device, one or more parameters associated with the second device, or one or more channel characteristics associated with a channel between the first device and the second device. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 illustrates an example of a wireless communications system in accordance with aspects of the present disclosure.
[0019] Figure 2 illustrates an example of a learning model diagram, in accordance with aspects of the present disclosure. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 9
[0020] Figure 3 illustrates an example of a communications system, in accordance with aspects of the present disclosure.
[0021] Figure 4 illustrates an example of a signaling diagram, in accordance with aspects of the present disclosure.
[0022] Figure 5 illustrates an example of a UE in accordance with aspects of the present disclosure.
[0023] Figure 6 illustrates an example of a processor in accordance with aspects of the present disclosure.
[0024] Figure 7 illustrates an example of an NE in accordance with aspects of the present disclosure.
[0025] Figure 8 illustrates a flowchart of a method performed by a UE in accordance with aspects of the present disclosure.
[0026] Figure 9 illustrates a flowchart of a method performed by an NE in accordance with aspects of the present disclosure. DETAILED DESCRIPTION
[0027] A wireless communications system may include one or more devices, such as UEs and NEs, among other devices, that transmit and receive signaling. The devices can implement one or more learning models, which may also be referred to as machine learning (ML) models and / or artificial intelligence (AI) models, to perform tasks related to the signaling. In some examples, the devices can implement the learning models to select a transmit beam and a receive beam, referred to as a beam pair, that result in relatively high signal strength and signal quality (e.g., greater than a threshold value) at a receiving device (e.g., a base station and / or UE, or other NE). In some examples, an NE or other device trains the learning models by determining one or more parameters of the learning models, which can include weights, biases, and / or other parameters that define respective layers of the learning models. During the training of the learning model, parameters are selected, such that the learning model maps (e.g., associates) a set of beam measurements (e.g., a received signal strength and / or quality from a set of beams) to one or more beam indices of beams that result in the relatively high signal strength and signal quality. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 10
[0028] The learning models are executed implemented) by one or more devices using the determined or selected parameters. However, the devices can input data to the learning models that has different statistical characteristics than the dataset over which the learning models are trained. In some cases, the NE can train multiple different learning models for respective domains over which the learning models make inferences or predictions, where a domain represents data for a network scenario, a network configuration, one or more parameters of a network, a physical propagation medium, and / or a device behavior. For example, a domain can include a dataset for a network configuration (e.g., a defined beam codebook, a defined set of multiple input multiple output (MIMO) configuration parameters, a type of scheduling device, different types of link adaptation and power control), for a device behavior (e.g., a device mobility, an orientation of the device, a quality of service (QoS) to support communications at the device), for different traffic patterns, for physical conditions of the propagation medium (e.g., channels with rich multipath vs. sparse channels, indoor vs. outdoor), among other examples, at the time of data collection or at the time of a simulation to generate the data in the dataset. The NE can store the different learning models for transmission to one or more other devices (e.g., UEs with reduced memory storage capabilities when compared with the NE). Additionally, or alternatively, the one or more other devices can store the different learning models. The other devices can use respective learning models for the different domains. However, transmitting and / or storing learning models for different domains results in increased signaling overhead and memory storage usage, as well as inefficient use of time- frequency resources due to the learning models having a relatively large (e.g., greater than a threshold value) numerical quantity of parameters.
[0029] As described herein, to reduce signaling overhead and processing related to transmission and storage of parameters of multiple learning models for different domains, a device can implement learning models that include a feature component and a classifier component. The feature component transforms input data into respective features or feature vectors, which are defined by a measurable or observable characteristic of a data sample. For example, if the input data includes beam measurement data samples, then the feature component can extract (e.g., obtain, retrieve, acquire) different types of beam measurements (e.g., received signal strength indicator (RSSI), signal-to-noise ratio (SNR), or signal-to-interference-plus-noise ratio (SINR)) from the data samples. The classifier component identifies patterns in input data that are indicative of different Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 11 classes or labels. For example, if the input data the different types of beam measurements extracted by the feature component, then the classifier component can output beam indices, as an example, for a beam pair that has a greatest signal strength and / or quality.
[0030] In some examples, a learning model can include multiple layers, where the feature component includes a subset of layers, and the classifier component includes another subset of layers. A layer of the learning model is defined by one or more parameters, which can include weights, biases, and / or other parameters. The subset of layers of a classifier component can include a relatively small numerical quantity of parameters (e.g., less than a threshold value) when compared with the subset of layers of a feature component. Multiple learning models (e.g., for different domains) can have a same set of parameters for layers of the feature component, while the learning models have unique sets of parameters for the layers of the classifier component. In some examples, the device can transmit an indication of an initial learning model that includes the parameters for the layers of the feature component and initial parameters for layers of the classifier component in a message to another device that is implementing the learning model. In some other examples, the device can select a learning model that includes the parameters for the layers of the feature component and initial parameters for layers of the classifier component for implementation.
[0031] In some examples, such as for subsequent updates to the learning model, the device can transmit different sets of parameters for the layers of a classifier component (e.g., without transmitting the parameters for the layers of the feature component again), as the parameters for the layers of the feature component are common to multiple learning models (e.g., the same for the learning models). The device and / or another device that is implementing (e.g., processing, executing) the learning model can switch (e.g., update, modify, change) out the parameters for the layers of the classifier component to use different learning models for respective domains. The device transmitting the parameters for the layers of the classifier component, where multiple learning models for different domains share parameters for the layers of the feature component, reduces signaling overhead, as well as processing and storage resources when compared with transmitting a full set of parameters for the multiple learning models. For example, the numerical quantity of parameters for the layers of the classifier component is less than a numerical quantity of parameters in the full set of parameters for a learning model, which takes up fewer time-frequency resources, leads to reduced processing, and leads to reduced memory storage. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 12
[0032] Although the learning models are as being trained and implemented for wireless communications, the learning models can additionally, or alternatively, be trained and implemented for any task that uses classification. Example tasks that use classification include, but are not limited to, wireless communications, image recognition, natural language processing, and bio-medical imaging, among other tasks. The learning models can be switched or communicated from one device to another device for implementation to perform the task.
[0033] Reference is made herein to receiving, transmitting, or communicating data or information, such as signaling communication resources and / or communications that are transmitted or received between devices. It is to be appreciated that other terms may be used interchangeably with communicating, such as signaling, transmitting, receiving, outputting, forwarding, retrieving, obtaining, and so forth. Similarly, other terms may be used interchangeably with transmitting (e.g., communicating, signaling, outputting, forwarding, and so forth), and other terms may be used interchangeably with receiving (e.g., communicating, retrieving, obtaining, and so forth).
[0034] Aspects of the present disclosure are described in the context of a wireless communications system.
[0035] Figure 1 illustrates an example of a wireless communications system 100 in accordance with aspects of the present disclosure. The wireless communications system 100 may include one or more NE 102, one or more UE 104, and a core network (CN) 106. The wireless communications system 100 may support various radio access technologies. In some implementations, the wireless communications system 100 may be a 4G network, such as an LTE network or an LTE-Advanced (LTE-A) network. In some other implementations, the wireless communications system 100 may be a NR network, such as a 5G network, a 5G-Advanced (5G-A) network, or a 5G ultrawideband (5G-UWB) network. In other implementations, the wireless communications system 100 may be a combination of a 4G network and a 5G network, or other suitable radio access technology including Institute of Electrical and Electronics Engineers (IEEE) 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), IEEE 802.20. The wireless communications system 100 may support radio access technologies beyond 5G, for example, 6G. Additionally, the wireless communications system 100 may support technologies, such as time division multiple access (TDMA), frequency division multiple access (FDMA), or code division multiple access (CDMA), etc. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 13
[0036] The one or more NE 102 may be throughout a geographic region to form the wireless communications system 100. One or more of the NE 102 described herein may be or include or may be referred to as a network node, a base station, an access point (AP), a network element, a network function, a network entity, network infrastructure (or infrastructure), a radio access network (RAN), a NodeB, an eNodeB (eNB), a next-generation NodeB (gNB), or other suitable terminology. An NE 102 and a UE 104 may communicate via a communication link, which may be a wireless or wired connection. For example, an NE 102 and a UE 104 may perform wireless communication (e.g., receive signaling, transmit signaling) over a Uu interface.
[0037] An NE 102 may provide a geographic coverage area for which the NE 102 may support services for one or more UEs 104 within the geographic coverage area. For example, an NE 102 and a UE 104 may support wireless communication of signals related to services (voice, video, packet data, messaging, broadcast, etc.) according to one or multiple radio access technologies. In some implementations, an NE 102 may be moveable, for example, a satellite associated with a non-terrestrial network (NTN). In some implementations, different geographic coverage areas associated with the same or different radio access technologies may overlap, but the different geographic coverage areas may be associated with different NE 102.
[0038] The one or more UEs 104 may be dispersed throughout a geographic region of the wireless communications system 100. A UE 104 may include or may be referred to as a remote unit, a mobile device, a wireless device, a remote device, a subscriber device, a transmitter device, a receiver device, or some other suitable terminology. In some implementations, the UE 104 may be referred to as a unit, a station, a terminal, or a client, among other examples. Additionally, or alternatively, the UE 104 may be referred to as an Internet-of-Things (IoT) device, an Internet-of- Everything (IoE) device, or machine-type communication (MTC) device, among other examples.
[0039] A UE 104 may be able to support wireless communication directly with other UEs 104 over a communication link. For example, a UE 104 may support wireless communication directly with another UE 104 over a device-to-device (D2D) communication link. In some implementations, such as vehicle-to-vehicle (V2V) deployments, vehicle-to-everything (V2X) deployments, or cellular-V2X deployments, the communication link may be referred to as a sidelink. For example, a UE 104 may support wireless communication directly with another UE 104 over a PC5 interface. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 14
[0040] An NE 102 may support with the CN 106, or with another NE 102, or both. For example, an NE 102 may interface with other NE 102 or the CN 106 through one or more backhaul links (e.g., S1, N2, N6, or other network interface). In some implementations, the NE 102 may communicate with each other directly. In some other implementations, the NE 102 may communicate with each other indirectly (e.g., via the CN 106). In some implementations, one or more NE 102 may include subcomponents, such as an access network entity, which may be an example of an access node controller (ANC). An ANC may communicate with the one or more UEs 104 through one or more other access network transmission entities, which may be referred to as a radio heads, smart radio heads, or transmission-reception points (TRPs).
[0041] The CN 106 may support user authentication, access authorization, tracking, connectivity, and other access, routing, or mobility functions. The CN 106 may be an evolved packet core (EPC), or a 5G core (5GC), which may include a control plane entity that manages access and mobility (e.g., a mobility management entity (MME), an access and mobility management functions (AMF)) and a user plane entity that routes packets or interconnects to external networks (e.g., a serving gateway (S-GW), a packet data network (PDN) gateway (P-GW), or a user plane function (UPF)). In some implementations, the control plane entity may manage non-access stratum (NAS) functions, such as mobility, authentication, and bearer management (data bearers, signal bearers, etc.) for the one or more UEs 104 served by the one or more NE 102 associated with the CN 106.
[0042] The CN 106 may communicate with a packet data network over one or more backhaul links (e.g., via an S1, N2, N6, or other network interface). The packet data network may include an application server. In some implementations, one or more UEs 104 may communicate with the application server. A UE 104 may establish a session (e.g., a protocol data unit (PDU) session, or the like) with the CN 106 via an NE 102. The CN 106 may route traffic (e.g., control information, data, and the like) between the UE 104 and the application server using the established session (e.g., the established PDU session). The PDU session may be an example of a logical connection between the UE 104 and the CN 106 (e.g., one or more network functions of the CN 106).
[0043] In the wireless communications system 100, the NEs 102 and the UEs 104 may use resources of the wireless communications system 100 (e.g., time resources (e.g., symbols, slots, subframes, frames, or the like) or frequency resources (e.g., subcarriers, carriers)) to perform Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 15 various operations (e.g., wireless . In some implementations, the NEs 102 and the UEs 104 may support different resource structures. For example, the NEs 102 and the UEs 104 may support different frame structures. In some implementations, such as in 4G, the NEs 102 and the UEs 104 may support a single frame structure. In some other implementations, such as in 5G and among other suitable radio access technologies, the NEs 102 and the UEs 104 may support various frame structures (i.e., multiple frame structures). The NEs 102 and the UEs 104 may support various frame structures based on one or more numerologies.
[0044] One or more numerologies may be supported in the wireless communications system 100, and a numerology may include a subcarrier spacing and a cyclic prefix. A first numerology (e.g., ^=0) may be associated with a first subcarrier spacing (e.g., 15 kHz) and a normal cyclic prefix. In some implementations, the first numerology (e.g., ^=0) associated with the first subcarrier spacing (e.g., 15 kHz) may utilize one slot per subframe. A second numerology (e.g., ^=1) may be associated with a second subcarrier spacing (e.g., 30 kHz) and a normal cyclic prefix. A third numerology (e.g., ^=2) may be associated with a third subcarrier spacing (e.g., 60 kHz) and a normal cyclic prefix or an extended cyclic prefix. A fourth numerology (e.g., ^=3) may be associated with a fourth subcarrier spacing (e.g., 120 kHz) and a normal cyclic prefix. A fifth numerology (e.g., ^=4) may be associated with a fifth subcarrier spacing (e.g., 240 kHz) and a normal cyclic prefix.
[0045] A time interval of a resource (e.g., a communication resource) may be organized according to frames (also referred to as radio frames). Each frame may have a duration, for example, a 10 millisecond (ms) duration. In some implementations, each frame may include multiple subframes. For example, each frame may include 10 subframes, and each subframe may have a duration, for example, a 1 ms duration. In some implementations, each frame may have the same duration. In some implementations, each subframe of a frame may have the same duration.
[0046] Additionally, or alternatively, a time interval of a resource (e.g., a communication resource) may be organized according to slots. For example, a subframe may include a number (e.g., quantity) of slots. The number of slots in each subframe may also depend on the one or more numerologies supported in the wireless communications system 100. For instance, the first, second, third, fourth, and fifth numerologies (i.e., ^=0, ^=1, ^=2, ^=3, ^=4) associated with respective subcarrier spacings of 15 kHz, 30 kHz, 60 kHz, 120 kHz, and 240 kHz may utilize a single slot per Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 16 subframe, two slots per subframe, four slots eight slots per subframe, and 16 slots per subframe, respectively. Each slot may include a number (e.g., quantity) of symbols (e.g., OFDM symbols). In some implementations, the number (e.g., quantity) of slots for a subframe may depend on a numerology. For a normal cyclic prefix, a slot may include 14 symbols. For an extended cyclic prefix (e.g., applicable for 60 kHz subcarrier spacing), a slot may include 12 symbols. The relationship between the number of symbols per slot, the number of slots per subframe, and the number of slots per frame for a normal cyclic prefix and an extended cyclic prefix may depend on a numerology. It should be understood that reference to a first numerology (e.g., ^=0) associated with a first subcarrier spacing (e.g., 15 kHz) may be used interchangeably between subframes and slots.
[0047] In the wireless communications system 100, an electromagnetic (EM) spectrum may be split, based on frequency or wavelength, into various classes, frequency bands, frequency channels, etc. By way of example, the wireless communications system 100 may support one or multiple operating frequency bands, such as frequency range designations FR1 (410 MHz – 7.125 GHz), FR2 (24.25 GHz – 52.6 GHz), FR3 (7.125 GHz – 24.25 GHz), FR4 (52.6 GHz – 114.25 GHz), FR4a or FR4-1 (52.6 GHz – 71 GHz), and FR5 (114.25 GHz – 300 GHz). In some implementations, the NEs 102 and the UEs 104 may perform wireless communications over one or more of the operating frequency bands. In some implementations, FR1 may be used by the NEs 102 and the UEs 104, among other equipment or devices for cellular communications traffic (e.g., control information, data). In some implementations, FR2 may be used by the NEs 102 and the UEs 104, among other equipment or devices for short-range, high data rate capabilities.
[0048] FR1 may be associated with one or multiple numerologies (e.g., at least three numerologies). For example, FR1 may be associated with a first numerology (e.g., ^=0), which includes 15 kHz subcarrier spacing; a second numerology (e.g., ^=1), which includes 30 kHz subcarrier spacing; and a third numerology (e.g., ^=2), which includes 60 kHz subcarrier spacing. FR2 may be associated with one or multiple numerologies (e.g., at least 2 numerologies). For example, FR2 may be associated with a third numerology (e.g., ^=2), which includes 60 kHz subcarrier spacing; and a fourth numerology (e.g., ^=3), which includes 120 kHz subcarrier spacing.
[0049] In some cases, the NE 102 (e.g., a base station) is equipped with multiple antennas, such as ^ antennas, in the form of one or more antenna array panels. One or more UEs 104, such as ^ Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 17 UEs 105, can have one or more antenna array with multiple antenna elements in each panel. In some variations, the devices in the wireless communications system 100 can use millimeter wave (mmWave) frequencies, which can be referred to as a frequency range 2 (FR2). In the mmWave frequency range, the devices can communicate using one or more narrow beams formed using multiple antenna elements available at the NE 102 and the UEs 104, such as to reduce propagation losses at the mmWave frequencies and to maintain a threshold signal strength. The total number beams formed by the NE 102 (e.g., across all of the antenna panels of the NE 102) is ^ and a UE 104 forms a number of beams, ^. Such beams are directional in nature (e.g., providing directional gain) and can have a narrow beam width (e.g., a beam width that is less than a threshold value). For achieving a reliable communication link between the NE 102 and the UE 104, the NE 102 and / or the UE 104 can select a beam for the NE 102 to use and a corresponding beam for the UE 104 to use. That is, the NE 102 and / or the UE 104 can select a beam pair that includes a receive beam for a receiving device to use for receiving signaling and a transmit beam for a transmitting device to use for transmitting signaling, such that the beam pair results in a signal strength at the receiving device that satisfies a threshold value.
[0050] In some examples, the process of selecting a beam pair can be referred to as a beam search procedure or a beam selection procedure. For signaling from a UE 104 to an NE 102 (e.g., downlink signaling), the beam selection includes selecting a transmit beam at the NE 102 and a corresponding receive beam at the UE 104. For signaling from the NE 102 to the UE 104 (e.g., uplink signaling), the beam selection includes selecting a transmit beam at the UE 104 and a corresponding receive beam at the NE 102. A beam corresponds to a non-zero power (NZP) channel state information-reference signal (CSI-RS) resource of at least one NZP CSI-RS resource set, where the NZP CSI-RS resource set is configured with a value of a higher-layer parameter set to repetition.
[0051] Conventionally, for beam selection, a device (e.g., the UE 104 and / or the NE 102) can search over a set of possible beams and select a beam with a maximum signal strength, which is referred to as an exhaustive search. In some examples, such as for downlink signaling, an NE 102 can send one or more reference signals using a set of possible transmit beams. A UE 104 can perform one or more measurements to obtain a received signal strength (e.g., a reference signal receive power (RSRP) and / or a SINR for layer 1 (L1)) for the beams in the set of possible transmit Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 18 beams. In some examples, L1, which can also to as the physical layer, supports transmission of data bits over a physical communication medium. The UE 104 can select a transmit beam that provides a highest RSRP and / or SINR. The UE 104 can transmit a message to the NE 102 that includes an indication of a beam index for the selected transmit beam. Additionally, or alternatively, the UE 104 can select a receive beam by fixing the transmit beam at the NE 102 and measuring the received signal strength for a set of possible receive beams (e.g., sweeping the receive beams at the UE 104). Thus, an exhaustive search includes selecting a beam (e.g., a best beam, a beam with a highest signal strength), but increases latency and signaling overhead related to transmitting and / or receiving signaling on the set of possible beams for performing measurements.
[0052] In some examples, a device (e.g., the UE 104 and / or the NE 102) can perform beam-pair prediction, in which the device predicts, infers, or otherwise determines a transmit and receive beam pair. For downlink signaling, the NE 102 uses the transmit beam to transmit signaling and the UE 104 uses the receive beam to receive the signaling. For uplink signaling, the UE 104 uses the transmit beam to transmit signaling and the NE 102 uses the receive beam to receive the signaling. In some variations, the device can implement one or more learning models to perform the beam selection and / or beam prediction. The learning models can include, but are not limited to, one or more ML models and / or one or more AI models. A device can train the learning models using a training data set obtained either through simulations or through field trials (e.g., considering either one or, at most, a finite set of physical cell-sites, network configurations, and / or wireless channel characteristics).
[0053] When a dataset includes both an input data sample and a corresponding output data sample for data samples in the dataset, the dataset is referred to as a labeled dataset. When the dataset includes the input data sample without the corresponding output data sample, the dataset is referred to as an unlabeled dataset. In some examples, a device can implement supervised learning techniques to train the learning models. For example, the NE 102 can train the learning models to map input data to output labels based on example input-output pairs provided during training. In supervised learning, the learning model computes a mapping function from input features to output labels by observing a dataset that includes labeled examples, such that the learning models can generalize the mapping to make accurate predictions on new, unseen data. Although supervised learning techniques are described, the device can additionally, or alternatively, implement any other Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 19 type of training techniques to train the including, but not limited to, unsupervised learning techniques and semi-supervised learning techniques, among others. In unsupervised learning, the learning models detects patterns, structures, or relationships within training data. Unlike supervised learning, there are no explicit output labels provided during training. Semi- supervised learning leverages both labeled and unlabeled data during a training procedure. The learning models use the labeled examples, while also using the structure of the unlabeled data to improve a performance of the learning models.
[0054] In some examples, data samples ^^^ = ^^x^, y^^^^^^^ denote the set of labeled trainingdata samples, where x^ ∈ ^ and y^ ∈ ^ denote an input sample and the corresponding label ordesired output data sample from a learning model for input x^. In some variations, that y^is referred to as a prediction and / or inference for input x^. In unsupervised learning, the training data setincludes an unlabeled data set ^^^ = ^x^^^^^^ .
[0055]
[0056] A learning model defines a mapping or a function, ^^, where ^^: ^ → ^. In somevariations, ^ = ^^^, ^^, … , ^^^^ denotes a set of modelare learned during theprocess oflearning model parameters include, but are not limited to, weights, biases, and activation function parameters, among other parameters. Weight parameters represent a strength of connections between neurons in different layers of the learning model. Biases are additional parameters added to neurons in the learning model that provide for the learning model to capture offset or bias in input data. Activation function parameters can include slope parameters or parameters defining a shape of an activation function in a parametric activation function. Determining values of model parameters that provide for a greatest accuracy, lowest latency, etc. (e.g., highest performance) of the learning model (e.g., determining ^ using the data set ^^^) is referred to as training the model or learning the model. When y^is discrete valued and assumes finitely many values (e.g., when |^|, the cardinality of the set ^ is finite), then the learningmodel is referred to as a classifier model. When y^ assumes continuous values (e.g., when ^ ⊆ ℝand |^| = ∞), then the learning model is referred to as a regression model. Supervised learningand / or training of a learning model includes minimizing a loss function ℒ. For example, the set of parameters (^) is determined by solving Equation 1: Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 20 (1)
[0057] In some examples, the deviceprediction and / or temporal beam prediction. Spatial beam prediction can include predicting a direction or orientation of a beam (e.g., a beam with a greatest signal strength) for beamforming or beam steering. Beamforming is a technique used to focus radio frequency energy in a defined direction, providing for increased signal strength, improved communication quality, and reduced interference. Spatial beam prediction can include analyzing channel conditions, device mobility, environmental obstacles, and interference sources, among other factors, to select a beam direction for transmission or reception of signaling. Temporal beam prediction can include predicting a direction or orientation of a beam (e.g., a beam with a greatest signal strength) for beamforming or beam steering over time. While spatial beam prediction focuses on selecting the beam direction based on current channel conditions and environmental factors, temporal beam prediction accounts for how the channel conditions may change over time.
[0058] In some cases, such as for spatial beam prediction, a device can select a beam (e.g., a best beam, a beam with a highest signal strength) out of ^ beams based on a set of current beam measurements rather than historical beam measurements. A value of ^, can vary according to one or more defined values (e.g., in a range from 8 to 256). The beams can span over an angular space, with each beam directed towards a different azimuth and / or elevation angle. Exhaustive beam selection includes performing ^ beam measurements (e.g., signal strength measurements, L1-RSRP measurements, and L1-SINR measurements, among other measurements). A device can use a supervised learning method to train a learning model, such as a deep neural network (DNN), using a labeled training data set. For example, the device can train the learning model to determine a beam with a highest performance (e.g., a highest signal strength or other metric, referred to as a best beam or optimal beam) when the device provides ^+beam measurements as input to the learning model,where ^+ < ^.
[0059] In some examples, indices for ^+number of beams over which measurements are performed are selected and fixed (e.g., prior to performing the spatial beam prediction). The ^+number of beams includes a set of possible beams or beams that are available for a device to use.The training data includes ^ ≫ 1 labeled samples, w ^ / ^^ here an % sample of the training data set canFirm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 21be expressed as 0)1^ , 1^ , … , 1^ *, 78, a ^ +2 ^ 456^ ^ nd 19 , : = 0, 1, … , ^ − 1, denotes one or moremeasurements signal strength, L1-RSRP, and / or L1-SINR) over the :^ / beam and 7^is thebeam index for with a highest performance (e.g., greatest signal strength, othermetrics) for the values of beam measurements )1^ ^ ^ ^ / 2 , 1^ , … , 1456^ *. In some examples, the %training labeled data sample, 0)1^ ^ ^ ^ ^ ^2 , 1^ , … , 1 sample is )12 , 1^ , … , 1456^ *and the corresponding label training, a spatial beam learningmodel finds a mapping ^12, 1^, … , 1456^^ and the best beam index7. When deployed, the learning model is supplied with ^+beam measurements and the learning model determines a beam index for a beam with a highest performance. Thus, the learning model developed for spatial beam prediction can be defined as a classification model.
[0060] In some examples, for spatial beam prediction a device can select a single beam (e.g., with a highest performance) out of ^ available spatial beams based on a current set of beam measurements. For temporal beam prediction, the input data can include one or more historical (e.g., past) selected beam indices and beam measurements. A learning model (e.g., a beam-pair DNN) can predict (e.g., infer) a beam with a highest performance, or ^ beams with a highestperformance for each of the τ ≥ 1 time slots in the future. Thus, the learning model developed fortemporal beam prediction can be defined as a classification model.
[0061] A device can implement a learning model for beam prediction (e.g., spatial and / or temporal beam prediction) for a defined set of one or more cell-sites, network configurations, channel types, and / or antenna parameters for which the learning model is trained. The learning model can perform the prediction in a relatively short duration (e.g., less than a threshold duration) when provided with ^+beam measurements. However, the learning model may not be generalized or adaptable to data distributions different from the data distributions of the training data. Forexample, x^ ∈ ^ and y^ ∈ ^ denote an input sample and the corresponding label or desired output,prediction, or inference from a learning model. In some cases, x^is a scalar or a one or multi- dimensional vector and y^is a scalar or a one or multi-dimensional vector. For beam prediction, x can be an ^+-dimensional vector including ^+beam measurement values and @ is a scalar value that indicates an index of a beam with a highest performance for a set of beam measurement values. In some cases, a device can provide additional, or alternative, inputs to a learning model used for Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 22 beam prediction, such as assistance location of the UE 104) and the output, @, caninclude indices of ^ ≥ 1 beams with a highest performance instead of a single beam index. ^ and^ denote the input sample space and the output sample space (e.g., label space), respectively.
[0062] A learning model defines a mapping or a function, ^^, where ^^: ^ → ^. In somevariations, ^ denotes a set of model parameters that are process of training themodel. When y^is discrete valued and assumes a finitely (e.g., when |^|, the cardinality of the set ^ is finite), then the learning model can be referred to as a classifier model. A device can develop a learning model by minimizing a loss function based on a training data set that includes labeled samples and / or unlabeled samples, resulting in supervised learning (e.g., training), unsupervised learning, and / or semi-supervised learning.
[0063] In some examples, a device can implement or deploy a trained learning model, such as to make predictions or inferences. The learning model can be expected to perform with a same level of accuracy and / or precision by providing desired inferences or predictions as seen during the training and testing phase of the learning models before deployment. However, in a real-world wireless network, a learning model can make predictions or inferences using input data that has different statistical characteristics than the dataset over which the learning model is trained. A learning model that uses input data with different statistical characteristics than the dataset over which the learning model is trained can output erroneous (e.g., incorrect) inferences or predictions. To improve accuracy of the inferences and / or the predictions, a device can train a learning model with a generalized ability to make inferences and / or predictions over many different data distributions (e.g., data distributions that are different from training data distributions and encountered at a time of inference). However, developing a learning model that generalizes the possible domains and outputs a desired performance across the domains may be difficult for the device, especially, in the context of wireless networks with varying statistics of data distributions.
[0064] In some examples, a source domain refers to a set of data samples over which a learning model 202 is trained. For example, when the learning model 202 is trained on data samples ^A=^^xA, yB^^^ ∼ DA , ^A is referred to as the sour A B^ ^ ^^^ EF ce domain. With x ∈ ^ and y ∈ ^ , the jointprobability density(PDF) DAEF : ^ × ^ → ℝH denotes the source distribution. In somevariations, a learning model 202 is trained on more than one source domains or over a dataset that Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 23 includes data samples coming from more than source domains, which leads to multiple sourcedomains, ^AI = ^)xAI , yBJ ^KL AI^ ^ *^^^^ ∼ DEF , % = 1, … , 1. A target domain refers to a set of data samples overwhich a inferences and / or predictions. Thus, when a learning model 202makes^ ^ ^ ^ ^^ ^ ^ ^^ a corresponding label y^, where ^ = ^ x^, y^ ^^^^ ∼ DEF , ^ iscalled the target domain and the pdf D^ HEF : ^ × ^ → ℝ is called target distribution.
[0065] In some examples, the learning model 202 is trained on a source domain having thedistribution DA HEF : ^ × ^ → ℝ . Once the learning model is deployed at a device, the learning modelcan make inferences and / or predictions in the target domain, where a joint PDF is given byD^EF : ^ × ^ → ℝH and D^EF ≠ DAEF . As DA^x, y^ = DA^x^DA^y|x^ and D^^x, y^ = D^^x^D^^y|x^,^ A A^ ^ ^^ ^ A^ ^ ^^ ^ A^ ^ ^^ ^D y|x^ ≠ D ^y|x^ or D ^x^ ≠ D ^x^ and D ^y|x^ ≠ D ^y|x^, such that DEF ≠ DEF . Thus, thelearning model can make predictions and / or inferences using new data distributions when the input distribution D^x^ changes, the label distribution D^y^ changes, or the conditional distribution D^@|x^changes. The learning model, ^^, learned and / or trained using the source domain data samples having the joint PDF DEAF may not perform well when making inferences or predictions using input data samples from a target domain having a different distribution DE^F . In some examples, a device can use a learning model to make inferences and / or predictions in multiple (e.g.,P) target data domains, ^9 = Q)x9^ , y9^ *R^S^^^ ∼ D9 9 9TEF , : = 1, … , P, where DEF ≠ DEF , for : ≠ :T, 1 ≤:, :T ≤ P.
[0066] For a learning model developed for beam selection, the data input to the learning model includes beam measurements and the output of the learning model is a beam with a highestperformance. Thus, D^x^ is a joint marginal density of the beam measurements ^12, 1^, … , 14+6^^supplied as input to the learning model used for beam selection. The D^x^can change from one geographic location to another or from one cell-site to another due to a change in the physical channel characteristics between the two geographic locations and / or cell-sites. Additionally, or alternatively, the values of the beam measurements can change due to measurement noise, which can be caused by multiple factors (thermal noise, hardware imperfections of related modules in the transceiver chain, differences in the hardware and the corresponding measurement sensitivity acrossdifferent user devices, etc.). Thus, the distribution of beam measurements ^12, 1^, … , 14+6^^ (e.g.,Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 24 the distribution D^x^of the input data samples learning model), is vulnerable to considerable variations. The relation and / or mapping between the beam measurements and the corresponding beam with a highest performance can change depending on the physical characteristics of the propagation medium. For example, the mapping between beam measurements and the beam with a highest performance changes from outdoor scenarios to indoor scenarios. Thus, the distribution of output conditioned on input, D^y|x^, can change for beam selection.
[0067] In some examples, a device can train a separate learning model for respective domains of the P domains. The device can identify different possible domains for a defined task (e.g., beam prediction, CSI prediction, AI based position, ML based position, AI based receiver, ML based receiver, etc.). For spatial beam prediction, there can be multiple data domains corresponding to indoor users, outdoor users, different codebooks (e.g., that have different beam widths or radiation patterns), rural environments, urban environments, and dense urban environments which influence the physical propagation environment, different orientations of the UE 104, etc. In some examples, ata set ^ = Q)x^ , y^ *R^Sa d 9 9 9^^^corresponds to a data domain, (e.g., a particular scenario or configuration of the wireless channel along with a particular set of network and / or UE parameters or a configuration), with DE9F as the joint distribution of data samples x^9and y^9.
[0068] For beam prediction, the device identifies different possible data domains and prepares atraining data set ^9 = Q)x ^S9 9 9^ , y^ *R^^^ ∼ DEF , where : = 1, … , P (e.g., the index : runs through thepossible targetdomains, scenarios, and / or configurations). The device trains a learning model(e.g., a DNN), ^^S , using the data set ^9 = Q)x9^ , y9^ *R^S^^^ ∼ D9EF , : = 1, … , P, where P is the numberof data domains identified for the task under consideration (e.g., beam. The device stores the learning models at a node or device that implements (e.g., uses) the learning model to make inferences or predictions. The node and / or device can either be a network node and / or device (e.g., a base station or network node responsible for determining the position based on the measurements reported by an edge node) or an edge node and / or device, such as a UE. For example, if the learning model 202 is a UE side model for assisting the UE to predict a beam to use through one or more UE measurements (e.g., reference signal receive power (RSRP) values), then the UE can store the P models. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 25
[0069] In some cases, training and testing a learning model can use a relatively large amount (e.g., greater than a threshold value) of computational resources, energy, and power supply. Additionally, or alternatively, training and testing a learning model can use a relatively large amount (e.g., greater than a threshold value) of memory, which can exceed an amount of memory, computational resources, energy, and power available to an edge device. Thus, a network node (e.g., an NE 102 or a centralized facility having no resource constraints) can train a learning model. The network node can transmit the trained learning models to the other devices where they are to be deployed or implemented. Transmitting multiple learning models to the node and / or device at which the models are going to be deployed leads to a significant cost of network resources (e.g., bandwidth, energy, time etc.) due to the learning model including a relatively large numerical quantity (e.g., hundreds or thousands) of parameters that are sent over a noisy wireless channel.
[0070] Additionally, or alternatively, storing multiple learning models is difficult at edge devices, which can have reduced memory or storage capabilities when compared with an NE 102 (e.g., due to having small form factors). Thus, the learning model can be stored at a central node and transferred to the edge device. The edge device storing a relatively small numerical quantity of learning models (e.g., less than a threshold numerical quantity of learning models) per task and receiving other learning models from an NE 102 or other device via signaling can lead to increased latency related to transfer of the learning models and implementation of the learning models. For example, a learning model can include a relatively large numerical quantity of parameters (e.g., greater than a threshold numerical quantity of parameters) and signaling including the parameters can consume a relatively significant amount of time (e.g., greater than a threshold duration) to provide the model to the edge device from a central node. Learning models with a relatively small numerical quantity of parameters (e.g., less than a threshold value) do not provide greater than a threshold prediction and / or inference performance. Further, compressing and / or reducing the numerical quantity of parameters of a learning model (e.g., a size of the learning model) through techniques such as network pruning, knowledge distillation, model compression (e.g., using Kronecker factorization, quantization of model parameters, etc.) do not effectively reduce the size of a DNN.
[0071] Conventionally, to train a learning model to make inferences and / or predictions across different data distributions that result from wireless channels and changes in network parameters, a device can use a mixed dataset. For example, a device (e.g., an NE 102) can train the model with a Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 26 training dataset that includes data samples from possible data distributions that the learning model can expect to encounter after deployment. However, a learning model trained with multiple data distributions has a sub-optimal performance across the different distributions, while a learning model trained for a specific data distribution would provide a better inference performance. Other conventional techniques include model adaptation, in which a device trains a learning model for one or two target domains and then adapts and / or adjusts the parameters of the learning model when the learning model is deployed for another target domain. However, model adaptation includes retraining the learning model with training samples from a new target domain, which increases the use of computational, energy, power, and memory resources at the device retraining the learning model. After deploying the learning model at a device and / or node, the device and / or node may be unable to adapt the learning model. If the learning model is deployed at an edge device with a relatively small form factor (e.g., a size of the edge device is less than a threshold size), then the edge device may be unable to adapt the model by retraining due to reduced capabilities related to computational, energy, power, and memory resources at the edge device. In some other conventional techniques to train a learning model to make inferences and / or predictions across different data distributions, a device can implement model fine-tuning or refining, which includes retraining the learning model.
[0072] Additionally, or alternatively, a device can implement multiple learning models with model transfer and switching. However, transmitting parameters for different learning model to provide for an edge device to switch models leads to increased signaling overhead every time a learning model is transferred from a node to another node (e.g., from an NE 102 to a UE 104). A learning model with a relatively large numerical quantity of parameters (e.g., greater than a threshold) can have improved performance relative to a learning model with less than the threshold numerical quantity of parameters but also results in a relatively high signaling overhead due to transmission of the parameters to device that implement or deploy the learning model. Techniques to reduce the size of the learning model, such as model compression, could reduce signaling overhead and latency. However, model compression may not be effective in reducing the signaling overhead and latency to value that satisfies one or more thresholds. In some variations, a device can develop multiple learning models that provide a relatively high level of inference and / or prediction performance over a specific data distribution or target domain (e.g., a specific scenario and / or configuration of a network and / or wireless channel), while ensuring that less than a threshold numerical quantity of parameters of the Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 27 learning models differ, thereby reducing the of signaling overhead when transferring the different learning models and minimizing the memory for storing the learning models.
[0073] In some examples, the device can develop a learning model for different target data domain (e.g., target data domains corresponds to a specific scenario, configuration, or parameter of the network, physical propagation medium, and / or device behavior). For example, the device can develop a learning model for different network configurations (e.g., different beam codebooks, different MIMO configuration parameters, different types of scheduling devices, different types of link adaptation and power control), for different device behaviors (e.g., different device mobility capabilities, different orientations of the device, different QoSs supported communications at the device), for different traffic patterns, for different physical conditions of the propagation medium (e.g., channels with rich multipath vs. sparse channels, indoor vs. outdoor), among other examples, at the time of data collection or at the time of a simulation to generate the data in the dataset. For beam prediction (e.g., spatial beam prediction and / or temporal beam prediction), the device can reduce signaling overhead related to transfer of multiple learning models, reduce a memory footprint of the multiple learning models, and reduce the latency related to transmitting the learning models, without reducing the size of the learning models to below a threshold value (e.g., which can degrade inference and / or prediction performance).
[0074] According to implementations, one or more of the NEs 102 and the UEs 104 are operable to implement various aspects of the techniques described with reference to the present disclosure. For example, an NE 102 (e.g., a base station) communicates a message that includes a first set of parameters that are common to multiple learning models and a second set of parameters that are unique to a learning model. The first set of parameters can indicate one or more layers of a feature component of a learning model, while the second set of parameters can indicate one or more layers of a classifier component of a learning model. In some variations, the second set of parameters includes a lower numerical quantity (e.g., less than a threshold numerical quantity) of parameters than the first set of parameters. The NE 102 can transmit the message to the UE 104. The UE 104 can implement a learning model that includes layers indicated by the first set of parameters and the second set of parameters. Additionally, or alternatively, the NE 102 can implement the learning model that includes layers indicated by the first set of parameters and the Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 28 second set of parameters. For example, the UE and / or the NE 102 can use the learning model for beam management.
[0075] Figure 2 illustrates an example of a learning model diagram 200 in accordance with aspects of the present disclosure. In some examples, the learning model diagram 200 implements aspects of the wireless communications system 100. For example, the learning model diagram 200 can implement aspects of, or can be implemented by, a UE and / or an NE, which may be examples of a UE 104 and an NE 102 as described with reference to Figure 1. The learning model diagram 200 illustrates an example of one or more learning models 202. In some cases, a learning model 202 can include a single learning model 202. In some other cases, a learning model 202 can include multiple learning models 202.
[0076] In some examples, the learning models 202 can be examples of ML models and / or AI models. For example, the one or more learning models 202 can be an example of a neural network with multiple layers. A neural network is a computational model that includes multiple layers of artificial neurons, which may be referred to as nodes or units, organized into an input layer 204, one or more feature extraction layers 206, one or more classification layers 208, and an output layer 210. An input layer 204 passes input data to subsequent layers and may not include learnable parameters. An output layer 210 represents a prediction generated by the neural network and may not include learnable parameters. The feature extraction layers 206 and / or the classification layers 208 can include one or more hidden layers. A node in a hidden layer receives an input signal, performs a mathematical operation on the input data, and produces an output signal, which is then passed on to other nodes in subsequent layers. The connections between nodes are represented by weighted edges, which determine the strength of the connection between nodes. During a training process, the weights are adjusted using input-output pairs from a training dataset, with the goal of minimizing a defined loss or error function. Neural networks are capable of learning complex patterns and representations from data, enabling them to perform a wide range of tasks, including classification, regression, clustering, pattern recognition, and sequence generation. Example neural networks include, but are not limited to, an autoencoder, a convolutional neural network (CNN), a recurrent neural network (RNN), a generative adversarial network (GAN), a long short-term memory (LSTM) network or any other type of neural network. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 29
[0077] In some variations, a device in a communications system (e.g., a UE and / or an NE) can implement the one or more learning models 202 for beam prediction, as described with reference to Figure 1. For example, the device can implement a learning model 202 with a feature component 212 and a classifier component (e.g., one of the classifier component 214-a, the classifier component 214-b, through the classifier component 214-T). In some cases, there are P tinct target data domains and ^^^ = Q)x9^ , y9^ *R^Sdis 9^^^ ∼ D9EF , is the labeled training data set for atarget domain :, with DEFas a joint distribution of input samples Qx^S 9 ^9R^^^and the corresponding S labels or expected predictions, inferences, and / or output Q@^9R^^^^ over the domain :, where : =1, … , P. As the domains have different statistical characteristics, D9 9TEF ≠ DEF , for : ≠ :T, 1 ≤ :, :T ≤P.8] In some cases, a data set ^ = Q)x9^ , y9^ *R^S[007 9^^^corresponds to a data domain (e.g., a specific scenario and / or configuration) of the wireless channel, with DE9F as the joint distribution of data samples x^9and y^9. Example data domains include, but are not limited to, data domains for different types of channels and / or communication scenarios (for a line-of-sight (LoS) channel, a non-line-of-sight (NLoS) channel, an indoor environment, an outdoor environment, an indoor hotspot, 3-dimensional (3D) urban micro (UMi) channels, 3D urban macro (UMa) channels, etc.), data domains for different network and / or UE configurations or parameters (antenna array specifications at an NE and / or a UE, a beam codebook, a UE orientation, UE mobility, etc.), among other examples.
[0079] In some examples, the learning models 202 can include site-specific learning models 202. For example, one or more target domains may correspond to site-specific scenarios, which leads to site-specific learning models 202 and / or device specific learning models 202 for device specific scenarios. For site-specific models, different learning models 202 are developed for different cell-sites or geographic locations, which result in different data distributions due to their distinct physical characteristics. For example, a cell-site in a rural location can have different channel characteristics than a cell-site in a dense urban environment, thereby resulting in two different data distributions DE9F and DE9TF . In some examples, the learning models 202 can include device specific learning models 202. The data distributions can change due to device specific Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 30 parameters and / or scenarios. For example, the channel experienced by a device that is in an outdoor environment can have different statistical characteristics than that of a device in an indoor environment. The mobility of a device and a physical orientation of the device can create different data distributions. Further, different UEs having different capabilities, such as a measurement sensitivity of a UE, can benefit from using multiple learning models 202.
[0080] A device can develop or train P learning models 202, ^^S , : = 1, … , P, where respectivelearning models 202 have a threshold performance (e.g., a threshold accuracy and precision) over at Sleast one target domain using the training data sets ^9^^ = Q)x9^ , y9^ *R^^^^ , : = 1, … , P. The devicecan implement the learning models 202 for atask, such as for spatial beam prediction, temporal beam prediction, a receiver, and / or predicting channel information (e.g., a channel quality indicator (CQI), a rank indicator (RI), and / or a modulation and coding scheme (MCS) index). A classification task can include any task that categorizes or labels input data as a defined class or category based on features or attributes of the input data. A learning model 202 for spatial beam prediction can predict either a single beam with a highest performance or ^ beams with highest performances out of a set of M beams based on a current set of beam measurements. A learning model 202 for temporal beam prediction can predict either a single beam with a highestperformance or ^ beams with highest performances for a future V ≥ 1 time slots based on thecurrent and past (e.g., historical) data. The current and past data can include current and past beam measurements and / or current and past beam indices with the highest performance. A learning model 202 at a receiver infers or predicts transmitted messages, symbols, and / or bits based on the signals at the output of the wireless channel. A learning model 202 can predict channel information (e.g., a CQI, a RI, and / or an MCS index) if the output of the learning model is a discrete value (e.g., where there are a finite number of values).
[0081] A learning model 202 for beam prediction (e.g., a BP-DNN), that predicts or infers a beam with a highest performance based on input data samples (e.g., beam measurements and / or assistance information that includes a location of the UE, among other information), can be a classifier. For example, a number of possible outputs from a learning model 202 for beam prediction are finite in number and are discrete, as predicting a beam with a highest performance includes determining the index of the beam with the highest performance. Thus, a learning model Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 31 202 for beam prediction can be trained to a set of input data samples given at any instance of inference into a beam with a highest performance from a set of available beams. In some cases, the learning model 202 for beam prediction can output ^ beams with highest performances. Forexample, with ^ = 2, the learning model 202 for beam prediction can infer and / or predict twobeams, where one of the beams has a highest performance and the other beam has a second-highest performance. In some examples, the learning model 202 for beam prediction includes a classifierwith a number of classes increased from ^ to to ^ × ^^ − 1^ × … × ^^ − ^ + 1^, where ^ is atotal number of candidate and / or available beams to select from. In temporal beam prediction, the input data includes historical and / or past beam indices for beams with highest performances and beam measurements. The learning model 202 for beam prediction (e.g., a BP-DNN) predicts and / orinfers a beam with a highest performance, or ^ beams with highest performances, for τ ≥ 1 timeslots in the future. Thus, the learning model 202 for beam prediction can be developed as a classifier.
[0082] In some variations, the learning model 202 includes a feature component 212 and a classifier component (e.g., at least one of a classifier component 214-a through a classifier component 214-T). The feature component 212 and the classifier component can represent sub- networks of a neural network and can include one or more layers. For example, the feature component 212 can include an input layer 204 and one or more feature extraction layers 206, which may be examples of hidden layers. In some variations, one or more last layers of a learning model (e.g., the output layer 210 and / or a layer prior to an output layer 210) can include at least one classification layer 208. In some cases, a classifier component 214-a through a classifier component 214-T can include a classification layer 208 and the output layer 210. In some other cases, the classification layer 208 can be part of the output layer 210, such that the classifier component 214-a through a classifier component 214-T can include an output layer 210 that includes the classification layer 208. That is, the classification layer 208 can be a specialized component within the output layer 210 that performs a final classification based on the learned features from the feature component 212.
[0083] Although the learning model 202 is illustrated as including a single classification layer 208, the learning model 202 can include any numerical quantity of classification layers 208. Similarly, although the learning model 202 is illustrated as including a classification layer 208 that Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 32 is separate from the output layer 210, the layer 208 and the output layer 210 can be a single layer. A classification layer 208 maps one or more learned features from one or more preceding layers (e.g., features extracted by the feature extraction layers 206 of the feature component 212) to one or more output classes or labels of the output layer 210. The output layer 210 can include neurons representing different classes or categories, and the activation of these neurons indicates the likelihood or probability of the input data belonging to a class. In some variations, the output from the feature component 212 is used as input data and is provided to a classifier component of a learning model 202 (e.g., at least one of the classifier component 214-a through the classifier component 214-T). The classifier component outputs a classification (e.g., the likelihood or probability of the input data belonging to a class) of input data provided to the input layer 204 of the feature component 212.
[0084] A learning model 202 ^^(e.g., a classifier network) can be divided into two components, or two sub-networks. The twonetworks can include a feature extractor sub-network, YZ, that computes and / or extracts the features from the input data samples. Additionally, or alternatively, the two sub-networks can include one or more last layers (e.g., a classifier sub-network), [\, which performs the classification by computing distance between the extracted features and the set ofclassification vectors \ = ^]^, … , ]4^. In some examples, Z denotes a set of weights and / orparameters of the feature extractor sub-network of the learning model 202 and \ denotes the set ofweights and / or parameters of the classifier sub-network. Thus, ^ = ^Z, \^ denotes the set of allweights and / or parameters of the learning model 202, ^^. The feature component 212 of the learning model 202 is separated from the classifier component (e.g., one of the classifier component 214-a through the classifier component 214-T) of the learning model 202 that includes classification vectors.
[0085] With a given feature extractor sub-network YZand the classifier sub-network [\that includes classification vectors as weights, the learning model 202 classifies each inputinto one of the ^ classes based on Equation 2: cde )f^cZ^g^,hi^*(2)where o is a distance-function. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 33
[0086] A device can develop multiple models 202 for beam prediction (e.g., BP-DNN models) for respective target domains of the P target data domains by changing the parameters (weights, biases, etc.) of the classifier component without changing the parameters of the feature component 212. The device can train and / or learn a learning model 202, ^^S , for one of the targetdata domains. That is, the device can fix : = p such that p ∈ ^1, … , P^ and determine the learningmodel 202, ^ q , by training over the r r ^q^ / r r^ p labeled training data set ^^^ = Q)x^ , y^ *R^^^ ∼ DEF . Thus,the learning model 202 includes trained or learned values of the parameters ^s = ^Zs, \s^ for thep^ / target data domain. In other words, to determine ^^q, the device determines a feature extractor sub-network, YZs, and a classifier sub-network, [\s, for the p^ / target data domain. The device can use supervised learning techniques to train and / or learn ^^qby minimizing a loss function ℒ overthe data set ^r = Q)xr r ^q ss s^^ ^ , y^ *R^^^ . For example, the set of parameters ^ = ^Z , \ ^ is determinedby solving^s = $^'%& ℒ )^', ^r^^ *. (3)
[0087] In some examples, theacross the target data domains (e.g., by keeping YZS = YZs for : = 1, … , P). The device can learnand / or determine a set of classification vectors \9 = ^]9 9^, … , ]4 ^ for the remaining P − 1 targetata domains using the corresponding data sets ^ ^Sd 9 9 9^^ = Q)x^ , y^ *R^^^ , : = 1, … , P, : ≠ p (e.g., in asimilar manner as a prototypical classificationdata).
[0088] For example, the device can determine and / or learn a learning model 202 ^^S for : ≠p, : = 1, … , P by fixing the feature extractor sub-network to the determined and / or learned featureextractor sub-network for the p^ / target data domain. The device can train and / or learn the classifiersub-network (e.g., a last layer of different learning models 202), [\S , for : = 1, … , P, : ≠ p (e.g.,each of the remaining P − 1 target domains), by using theSQ)x9 9^ *R^: = … : ≠ p. The device can determine, learn, and / or train the last layervectors or prototypes) of the learning model 202, \9=^]9, … , ]94 ^ for : = 1, … , P, : ≠ p (e.g., each of the remaining P − 1 target.Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 34example, the device can train and / or learn the sub-network, [\S , for : = 1, … , P, : ≠ p,(e.g., the classification vectors \9 = ^]9, … , 9^ ]4 ^ or prototypes of a BP-DNNnetwork), while fixing the feature extractor sub-network YZS = YZs for : = 1, … , P, such that thelearning model 202, ^^S , is trained for implementation over a target data domain ^^9^ =Q)x^ , y^ *R^S9 9^^^ ∼ D9EF , for : = 1, … , P.99 9 ^Sdomain data set ^^^ = Q)x^ , y^ *R^^^includes labeled datasamples. Based on the labeled data samples )x9^ , y9^ * in :^ / target domain, the parameters for theclassifier sub-network, [\S, or optimal values of the classification vectors and / or prototypes \9=^]9 9 ^^, … , ]4 ^ for the : / target domain (e.g., for : = 1, … , P, : ≠ p) can be learned and / orata samples ^ 9 9 ^Sdetermined with labeled d 9 9^^ = Q)x^ , y^ *R q^^^ ∼ DEF . In some examples, YZS = YZ for: = 1, … , P. The device inputs the labeled data samples from : / target domain Q)x9 9 ^S ^^ , y^ *R^^^, YZS,and \r = ^]r^ , … , ]r4 ^ as training data to the learning model 202. The learning model 202 outputsa set of vectors \9 = ^]9, … , ]9^ 4 ^ (e.g., the classifier sub-network, [\S). The devicetializes tu = hqini iv q , $ = 1, … , ^, where tb is the support set for class $. Then, for % = 1 →weights of the last layer of a learning model 202 or a classifier sub-network, [\S , of the learning model 202 ^^S ) according to Equation 4: ]9b = ^ ∑ { , $ = 1, … , ^. (4)In some examples, YZ= YZq for : = … ,respective learning model 202 is the same, where each learning model 202 is optimized for a datadomain :, : = 1, … , P. The feature extraction sub-network for respective learning models 202 canbe denoted by YZ, where YZ = YZn = YZ| = ⋯ = YZ~ . The P learning models 202 trained and / orlearned for target data domains 1, … , P, differ in the weights and / or parameters of the classifier sub-Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 35 network, which can include a last layer of the models 202. That is, the P learning models 202 have a same feature extractor sub-network and the classification vectors of the classifier sub- network are trained (e.g., optimized) for the target data distributions represented by different targetdomain data sets ^9^^ , : = 1, … , P.
[0091] In some variations, a greater portion of the layers and corresponding parameters of a learning model 202 are part of the feature extractor sub-network than for the classifier sub-network. For example, a number of parameters that determine the classifier sub-network are a fraction of the number of parameters of the feature extraction sub-network of the learning model 202. Thus, a device can store P learning models 202 by storing one feature extractor sub-network YZand Pclassifier sub-networks, [\S , where : = 1, … , P. Storing a single feature extractor sub-network anddifferent classifier sub-networks leads tousage of storage and memory at the device. Similarly, to transfer P different learning models 202 for the P different data domains (e.g., data distributions), the device can transfer one feature extractor sub-network YZand P classifier sub-networks, [\S , where : = 1, … , P. Transmitting or transferring a single feature extractor sub-network and different classifier sub-networks leads to reduced usage of time-frequency resources, as well as decreased signaling overhead to indicate the learning models 202 to other devices.
[0092] Additionally, or alternatively, transmitting or transferring a single feature extractor sub- network and different classifier sub-networks leads to reduced latency. In some cases, once the feature extractor sub-network is available at another device, such as a UE or any node that deploys the learning model 202 for inference, transferring an additional or new learning model 202 includes transferring the classifier sub-network, which includes less than a threshold numerical quantity of parameters. The device training an initial learning model 202 includes computationally expensiveprocessing of computing gradients and back propagation. For the remaining P − 1 learning models202, the device trains the classifier sub-network, which does not include gradient computations and backpropagation through the learning models 202 leading to improved efficiency of training multiple learning models 202 by reducing the training time and computational resources used during training. The device can train and transfer a single feature extractor sub-network and different classifier sub-networks for any learning model 202 implemented for a classification task. Thus, the device can develop and / or train multiple domain specific learning models 202 at a receiver that infer and / or predict transmitted messages, symbols, and / or bits based on the signals at Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 36 the output of the wireless channel. alternatively, the device can develop and / or train multiple domain specific learning models 202 for predicting a CQI, a RI, and / or an MCS index, as long as the output of a learning model 202 (e.g., CQI, RI, and / or MCS) is a discrete value of a finite number of values.
[0093] Figure 3 illustrates an example of a wireless communications system 300 in accordance with aspects of the present disclosure. In some examples, the wireless communications system 300 implements aspects of the wireless communications system 100 and the learning model diagram 200. For example, the wireless communications system 300 includes a UE 104 and an NE 102, which may be examples of a UE 104 and an NE 102 as described with reference to Figure 1. In some examples, an NE 102 may be in wireless communications with one or more other devices in the wireless communications system 300. For example, the UE 104 transmits signaling to the NE 102 via an uplink communication link 302, which may be an example of a communication link as described with reference to Figure 1. In some other examples, the NE 102 transmits signaling to the UE 104 via a downlink communication link 304, which may be an example of a communication link, as described with reference to Figures 1 and 2. The signaling between the UE 104 and the NE 102 may include control signaling and / or data transmissions.
[0094] The devices in the wireless communications system 300 can implement one or more learning models (e.g., machine earning models and / or AI models) for various tasks, where the learning models can be examples of the learning models 202, as described with reference to Figure 2. For example, an NE 102 and / or a UE 104 can implement or deploy one or more learning models for a classification task, such as for beam prediction, predicting a CQI, a RI, and / or an MCS index, and as a wireless receiver, among others example classification tasks. The learning models can include multiple layers, where respective layers include one or more parameters. The layers can include, but are not limited to, an input layer, at least one feature extraction layer (e.g., hidden layers), at least one classification layer, and an output layer. In some variations, the classification layer can be a component of the output layer. In some examples, the parameters for hidden layers and / or the classification layer of a learning model can include weights and / or biases that define the relationship between nodes or neurons in the hidden layer and / or the classification layer.
[0095] In some examples, a device in the wireless communications system 300 can train one or more of the learning models to obtain or generate the parameters of the different layers of the Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 37 learning models. In some variations, the NE train the one or more learning models, while the NE 102 and the UE 104 implement the learning models. For example, the NE 102 can generate the parameters that define the learning models and can transmit the parameters to the UE 104, where the UE 104 implements a learning model that includes the parameters. In some other variations, the NE 102 and / or the UE 104 can both train and implement the learning models.
[0096] Conventionally, an NE 102 can develop (e.g., train) multiple learning models for a defined task, where respective learning models have different parameters for the different layers. The NE 102 can develop a learning model for different target domains for the defined task. The different target domains can represent a specific scenario, configuration, parameter of the network, physical propagation medium, or device behavior. That is, a target domain represents data for a scenario, configuration, parameter of a network, a physical propagation medium, and / or a device behavior. The NE 102 can store parameters of the different learning models for transmission to one or more other devices (e.g., UEs 104 or other devices with reduced memory storage capabilities when compared with the NE 102). Additionally, or alternatively, the one or more other devices can receive and store the parameters of different learning models. The other devices can use respective learning models for the different target domains. For example, a device can provide data samples from the target domain as input to at least one learning model. The learning model can provide an output that the device uses for a task. However, transmitting and / or storing parameters of learning models for different target domains results in increased signaling overhead and memory storage usage, as well as inefficient use of time-frequency resources due to the learning models having a relatively large (e.g., greater than a threshold value) numerical quantity of parameters.
[0097] To reduce signaling overhead, memory storage usage, and time-frequency resource usage, the devices in the wireless communications system 300 can implement learning models with a common feature component (e.g., a same feature component) and different classifier components. For example, in an initial training process (e.g., for a first learning model), an NE 102 can develop or train a full learning model by obtaining parameters for the different layers, including the layers of the feature component and the layers of the classifier component. For subsequent learning models, the NE 102 can develop or train a learning model by fixing the parameters of the feature component to the parameters obtained during the initial training process. The NE 102 can obtain different parameters for the layers of the classifier component for each learning model. At 306, the NE 102 Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 38 generates feature extraction parameters that to multiple learning models and different classification parameters for respective learning models. The feature extraction parameters include a set of one or more parameters for layers of a feature component. Classification parameters for a single learning model include a set of one or more parameters for layers of a feature component of the learning model.
[0098] For example, the multiple learning models share the feature extraction parameters, such that the multiple learning models have a same feature component, and have different classification parameters, such that different learning models have different classifier components. A numerical quantity of classification parameters can be less than a numerical quantity of feature extraction parameters. Additionally, or alternatively, the numerical quantity of classification parameters can be less than a threshold value. Thus, training multiple learning models by obtaining a single set of feature extraction parameters and multiple respective sets of classification parameters provides for reduced processing and memory usage for storing the parameters for the different models at the NE 102.
[0099] In some examples, the UE 104 can transmit one or more data samples 308 to the NE 102. For example, the UE 104 can transmit control signaling and / or a data transmission to the NE 102 via the uplink communication link 302 that includes the data samples 308. The NE 102 can use the data samples 308 to generate the parameters for the different learning models. The content of the data samples 308 can be based on the task for which the learning models are to be used. The data samples 308 can include one or more measurements performed by the UE 104, including RSRP measurements, CSI measurements, and beam measurements, among others. For example, if the task includes beam prediction, then the data samples 308 can include one or more beam measurements, such as beamforming gain measurements, beamwidth measurements, and SINR measurements, among other measurements.
[0100] Additionally, or alternatively, the NE 102 can obtain or otherwise determine the data samples 308, such as by obtaining the data samples 308 from a simulation and / or receiving signaling from another device that indicates the data samples 308. The NE 102 can use the data samples 308 as a training dataset to obtain multiple learning models by generating parameters for respective learning models. In some examples, the data samples 308 can include labeled and / or unlabeled training data samples, such as for supervised learning, unsupervised learning, and / or Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 39 semi-supervised learning. A labeled data include two components, where a first component is a data sample for input to a learning model and a second component is a label of the input sample (e.g., an expected output sample from the learning model). An unlabeled data sample can include a single component, where the single component is an input sample to a learning model.
[0101] In some variations, the NE 102 can divide or group the data samples 308 into subsets of data samples 308. For example, the NE 102 can divide the data samples 308 into subsets of data samples 308 with different statistical characteristics (mean, standard deviation, variance, probability distribution, etc.). That is, the NE 102 divides the data samples 308 such that a subset of data samples 308 is representative of a target domain for a task with the defined statistical characteristics. Additionally, or alternatively, the data samples 308 can be preconfigured or defined, such that the data samples 308 are divided into subsets of data samples 308 prior to being obtained by the NE 102. The NE 102 can train a same numerical quantity of learning models as a numerical quantity of subsets of data samples 308 (e.g., a respective learning model for the target domains for the task).
[0102] In some examples, the NE 102 can transmit an indication of the learning models to the UE 104 via the downlink communication link 304. For example, the NE 102 can transmit a feature extraction parameters and classification parameters indication 310 to the UE 104. The feature extraction parameters and classification parameters indication 310 can include control signaling with one or more fields with values that include the parameters for one or more layers of a feature component of at least one learning model and parameters for one or more layers of a classifier component of the learning model. The UE 104 can use the parameters for layers of the feature component and the parameters for layers of the classifier component to generate a learning model for implementation and / or deployment. For example, the UE 104 can use a learning model with the parameters for layers of the feature component and the parameters for layers of the classifier component to perform a task by providing additional data samples as input to the learning model and receiving output from the learning model.
[0103] In some examples, the UE 104 can transmit a request for the learning model parameters 312. The request can include a message in uplink control signaling that includes one or more fields that indicate to the NE 102 to transmit the feature extraction parameters and classification parameters indication 310. In some variations, the request can indicate one or more learning models Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 40 (e.g., for one or more defined target domains). some other variations, the request can indicate for the NE 102 to transmit the parameters for layers of the feature component, the parameters for layers of the classifier component, or both for a learning model. In some examples, the UE 104 can transmit multiple requests for the learning model parameters 312. For example, the UE 104 can request an initial message from the NE 102 that indicates both the parameters for layers of the feature component and the parameters for layers of the classifier component. Additionally, or alternatively, the UE 104 can request one or more subsequent messages from the NE 102 that indicates parameters for layers of respective classifier components for one or more learning models, where the learning models share a common set of feature extraction parameters and have unique classifier parameters. The NE 102 can transmit an updated classification parameters indication 314 that includes respective classification parameters for one or more learning models, where the NE 102 transmits the feature extraction parameters to the UE 104 prior to the updated classification parameters indication 314.
[0104] Additionally, or alternatively, the NE 102 can train the one or more learning models for executing at the NE 102. The NE 102 can select a trained learning model to use for a task. For example, the NE 102 can select a learning model and can provide data samples as input to the learning model. The NE 102 can use the output from the learning model for a defined task, as described with reference to Figure 1. Although the NE 102 is described as training the learning models, the UE 104 or another device in the wireless communications system 300 can train the one or more learning models.
[0105] Figure 4 illustrates an example of a signaling diagram 400 in accordance with aspects of the present disclosure. In some examples, the signaling diagram 400 may implement aspects of the wireless communications system 100, the learning model diagram 200, and the wireless communications system 300. The signaling diagram 400 may illustrate an example of a device 402-a transmitting hidden layer parameters and normalization layer parameters for at least one learning model to a device 402-b. In some cases, the device 402-a may be an example of an NE 102 and the device 402-b may be an example of a UE 104, as described with reference to Figures 1 through 3. Alternative examples of the following may be implemented, where some processes are performed in a different order than described or are not performed. In some cases, processes may include additional features not mentioned below, or further processes may be added. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 41
[0106] In some examples, at 404, the b transmits a request for one or more learning model parameters (e.g., ML model parameters and / or AI model parameters) to the device 402-a. The request for the learning model parameters can include an indication of a learning model and / or an indication of a target domain for which the device 402-b is to implement the learning model. In some cases, the device 402-b transmits a message that requests at least one of a first set of one or more parameters or at least one second set of one or more parameters of a learning model.
[0107] In some cases, at 406, the device 402-b transmits one or more data samples to the device 402-a. In some variations, the device 402-b can transmit the data samples prior to transmitting the request for the learning model parameters. Additionally, or alternatively, the device 402-b can transmit the data samples after transmitting the request for the learning models. The device 402-b can obtain the data samples by performing one or more measurements. For example, the data samples can include one or more beam measurements, such as beamforming gain measurements, beamwidth measurements, and SINR measurements, among other measurements. In some other cases, the device 402-a determines the data samples independent of the device 402-b (e.g., without receiving signaling or other indications of the data samples from the device 402-b). For example, the device 402-a can perform measurements to determine the data samples, perform a simulation that generates the data samples, and / or can receive the data samples from another device (e.g., a device other than the device 402-b).
[0108] In some variations, the data samples can include subsets of training data samples. For example, the NE 102 can divide a training dataset into subsets of training data samples. In some other examples, the NE 102 obtains a training dataset that is already divided into the subsets of training data samples. The training dataset includes subsets of training data samples for respective target domains for a task, and the NE 102 trains a learning model using the subsets of training data samples. Thus, the NE 102 trains a same numerical quantity of leaning models as a numerical quantity of subsets of training data samples. The NE 102 can divide the subsets of training data based on the characteristics of the subsets of training data. For example, the subsets of training data have different characteristics including, but not limited to, a mean of the respective subsets of training data samples, a standard deviation of the respective subsets of training data samples, a variance of the respective subsets of training data samples, or a probability distribution of the respective subsets of training data samples. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 42
[0109] The subsets of training data at least one of a respective network condition between the device 402-a and the device 402-b, respective parameters associated with the device 402-a, respective parameters associated with the device 402-b, or respective channel characteristics of a channel between the device 402-a and the device 402-b. The training data samples can include measurements or other parameters related to a network used for communications (e.g., transmitting and receiving signaling) between the device 402-a and the device 402-b, such as signal strength of the network, interference in the network, channel congestion in the network, a data rate for the network (e.g., throughput), latency in the network, among other network conditions. Additionally, or alternatively, the training data samples can include one or more parameters of a device (e.g., of the device 402-a and / or the device 402-b), such as a transmit power for signaling sent from the device, an antenna gain for the device, a modulation scheme used by the device, a data rate supported by the device, and a duplexing mode used by the device, among other parameters. Example channel characteristics can include, but are not limited to, attenuation, propagation delay, multipath fading, noise and interference, bandwidth, channel capacity, and Doppler shift, among others characteristics.
[0110] In some examples, the data samples include labeled data samples and / or unlabeled data samples. A labeled data sample includes a data sample for input to a learning model and an indication of an expected output from the learning model (e.g., a label). An unlabeled data sample includes a data sample for input to the learning model (e.g., without a corresponding label).
[0111] At 408, the device 402-a obtains a set of learning models. For example, the device 402-a generates a first set of parameters common to a set of learning models and respective second sets of parameters for different learning models in the set of learning models. That is, the device 402-a can perform an initial training process that includes generating a first set of parameters and a second parameters for a first learning model. The first set of parameters are common to multiple different learning models, such that any subsequent learning models that the device 402-a trains share the first set of parameters. The learning models have different second sets of parameters. For example, a learning model in a set of learning models is determined by the first set of parameters and one of the second sets of parameters.
[0112] In some examples, each learning model includes a respective set of layers. In some examples, the first set of parameters defines a subset layers that are used for feature extraction and Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 43 the second sets of parameters defines a subset that are used for classification. The subset of layers that are used for classification can include at least one layer, while the subset of layers that are used for feature extraction can include multiple layers. The learning models can include a feature component and respective classifier components, as described with reference to Figure 2. Examples parameters in the first set of parameters include, but are not limited to, weights, biases, and activation function parameters for the subset layers that are used for feature extraction. Example parameters in the second set of parameters include, but are not limited to, weights, biases, and activation function parameters for the subset layers that are used for classification.
[0113] In some examples, at 410, the device 402-a trains learning models using data samples to obtain the set of learning models. For example, the device 402-a trains an initial learning model using a subset of data samples from a dataset to determine the first set of parameters and an initial second set of parameters. The device 402-a trains the initial learning model by minimizing a loss function using the subset of data samples. In some cases, one or more layers of the initial learning model are defined by the first set of parameters and the initial second set of parameters. For example, the initial second set of parameters define one or more classification layers of the initial learning model. The first set of parameters are common to multiple learning models, while the initial second set of parameters are unique to the initial learning model.
[0114] In some examples, after training the initial learning model, the device 402-a trains one or more remaining learning models of a set of learning models for a task. The device 402-a can use respective remaining subsets of data samples of the dataset to determine second sets of parameters for the remaining learning models. For example, the device 402-a trains one or more remaining learning models by minimizing the loss function using the respective remaining subsets of data samples. The learning models in the set (e.g., the initial learning model and the remaining learning models) share a common set of first parameters and have unique second sets of parameters. The first set of parameters can include any numerical quantity of parameters. Similarly, the second set of parameters can include any numerical quantity of parameters. A numerical quantity of parameters in the first set of parameters can be greater than a numerical quantity of parameters in the second set of parameters. The numerical quantity of parameters in the second set of parameters can satisfy a threshold value (e.g., be less than a threshold value), such that the signaling overhead for Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 44 transmitting a second set of parameters to device is relatively small (e.g., less than a threshold value).
[0115] In some examples, at 412, the device 402-a can select at least one learning model of the set of learning models for executing at the device 402-a. In some examples, the device 402-b selects the learning model based on one or more conditions being satisfied. For example, the device 402-b can select a learning model based on a network condition between the device 402-a and the device 402-b being satisfied, a threshold value for one or more parameter values for the device 402-a being satisfied, a threshold value for one or more parameter values for the device 402-b being satisfied, or one or more channel characteristics for a channel between the device 402-a and the device 402-b being satisfied.
[0116] Additionally, or alternatively, at 414, the device 402-a can transmit a message to the device 402-b that includes the first set of parameters and at least one second set of parameters. In some variations, the device 402-a transmits the message responsive to receiving the request at 404. In some examples, the device 402-b transmits the message if one or more conditions are satisfied. For example, the device 402-b can determine a network condition between the device 402-a and the device 402-b is satisfied, one or more parameter values for the device 402-a are satisfied, one or more parameter values for the device 402-b are satisfied, or one or more channel characteristics for a channel between the device 402-a and the device 402-b are satisfied.
[0117] In some cases, at 416, the device 402-a can obtain output from a learning model. That is, the device 402-a can use the selected learning model for a task in addition to, or as an alternative to, transmitting the parameters for one or more learning models for executing at the device 402-b. The device 402-a can provide data samples as input to the learning model. The learning model can output information that the device 402-a can use for a task. The output can include a beam pair prediction, or any other output that can be used for a classification task (e.g., a receiver task and / or an information prediction task). In some examples, the device 402-a can obtain one or more attributes of a data sample as output from the layers of the selected learning model that are trained for feature extraction by providing a data sample as input to the learning model. The attributes of the data sample are provided as input to one or more layers that are trained for classification. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 45
[0118] If the device 402-a transmits the for a learning model to the device 402-b, then, at 418, the device 402-b can obtain output from a learning model. For example, the device 402-b can generate a learning model using the first set of parameters and a second set of parameters received at 412. The device 402-b can provide data samples as input to the learning model. The learning model can output information that the device 402-b can use for a task. The output can include a beam pair prediction, or any other output that can be used for a classification task (e.g., a receiver task and / or an information prediction task). In some examples, the device 402-b can obtain one or more attributes of a data sample as output from the layers of the selected learning model that are trained for feature extraction by providing a data sample as input to the learning model. The attributes of the data sample are provided as input to one or more layers that are trained for classification.
[0119] In some examples, at 420, the device 402-b transmits a request for updated learning model parameters. For example, the device 402-b can detect a change in a target domain and can request a new learning model to use for the target domain. The change in the target domain can include, but is not limited to, a change in network conditions between the device 402-a and the device 402-b, a change in parameters associated with the device 402-a, a change in parameters associated with the device 402-b, or a change in channel characteristics of a channel between the device 402-a and the device 402-b. The request for updated learning model parameters can include an indication of a learning model and / or an indication of the detected change.
[0120] In some cases, at 422, the device 402-a transmits an updated second set of parameters indication to the device 402-b. The device 402-b uses a learning model with the first set of parameters and the updated second set of parameters. For example, the device 402-b switches the second sets of parameters from an existing second set of parameters to the new second set of parameters to switch a learning model. That is, the device 402-b switches a classifier component of a learning model by switching from one of the second sets of parameters to another one of the second sets of parameters. The updated second set of parameters indication can include the second set of parameters for a learning model and not the first set of parameters, as the first set of parameters is common to the learning models.
[0121] Additionally, or alternatively, the device 402-a can detect a change in a target domain and can update a second set of parameters for a learning model implemented by the device 402-a. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 46 For example, the device 402-a can switch the set of parameters from an existing set of parameters to a new set of parameters to update a learning model from an existing learning model to a new learning model. That is, device 402-a can develop multiple learning models to use and can switch the learning model by switching one or more last layers (e.g., a classification component) of the learning model.
[0122] Figure 5 illustrates an example of a UE 500 in accordance with aspects of the present disclosure. The UE 500 may include a processor 502, a memory 504, a controller 506, and a transceiver 508. The processor 502, the memory 504, the controller 506, or the transceiver 508, or various combinations thereof or various components thereof may be examples of means for performing various aspects of the present disclosure as described herein. These components may be coupled (e.g., operatively, communicatively, functionally, electronically, electrically) via one or more interfaces.
[0123] The processor 502, the memory 504, the controller 506, or the transceiver 508, or various combinations or components thereof may be implemented in hardware (e.g., circuitry). The hardware may include a processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or other programmable logic device, or any combination thereof configured as or otherwise supporting a means for performing the functions described in the present disclosure.
[0124] The processor 502 may include an intelligent hardware device (e.g., a general-purpose processor, a DSP, a CPU, an ASIC, a field-programmable gate array (FPGA), or any combination thereof). In some implementations, the processor 502 may be configured to operate the memory 504. In some other implementations, the memory 504 may be integrated into the processor 502. The processor 502 may be configured to execute computer-readable instructions stored in the memory 504 to cause the UE 500 to perform various functions of the present disclosure.
[0125] The memory 504 may include volatile or non-volatile memory. The memory 504 may store computer-readable, computer-executable code including instructions when executed by the processor 502 cause the UE 500 to perform various functions described herein. The code may be stored in a non-transitory computer-readable medium such as the memory 504 or another type of memory. Computer-readable media includes both non-transitory computer storage media and Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 47 communication media including any medium facilitates transfer of a computer program from one place to another. A non-transitory storage medium may be any available medium that may be accessed by a general-purpose or special-purpose computer.
[0126] In some implementations, the processor 502 and the memory 504 coupled with the processor 502 may be configured to or operable to cause the UE 500 to perform one or more of the functions described herein (e.g., executing, by the processor 502, instructions stored in the memory 504). For example, the processor 502 may support wireless communication at the UE 500 in accordance with examples as disclosed herein. The UE 500 may be configured to or operable to support a means for transmitting, to a second device, a first message that requests at least one of a first set of one or more parameters associated with respective learning models of a set of learning models or at least one second set of one or more parameters associated with at least one learning model of the set of learning models, where each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and receiving, responsive to the first message and from the second device, a second message including at least one of the first set of one or more parameters or the at least one second set of one or more parameters.
[0127] Additionally, the UE 500 may be configured or operable to support any one or combination of the subset of layers includes at least one layer associated with classification, an additional subset of layers of the respective set of layers includes a set of layers associated with feature extraction, and the set of layers associated with the feature extraction correspond to the first set of one or more parameters. Additionally, or alternatively, the UE 500 may be configured or operable to support obtaining, as output from the set of layers associated with the feature extraction, one or more attributes of a data sample based on providing the data sample as input to the at least one learning model associated with the first set of one or more parameters and the at least one second set of one or more parameters. Additionally, or alternatively, the at least one second set of one or more parameters is determined based on the at least one learning model being trained using a subset of data samples of a set of data samples. Additionally, or alternatively, the UE 500 may be configured or operable to support transmitting a third message that indicates the set of data samples. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 48
[0128] Additionally, or alternatively, the data samples includes a set of subsets of training data samples corresponding to each learning model of the set of learning models, and respective subsets of training data samples of the set of subsets of training data samples are associated with one or more different characteristics. Additionally, or alternatively, the one or more different characteristics include at least one of a mean of the respective subsets of training data samples, a standard deviation of the respective subsets of training data samples, a variance of the respective subsets of training data samples, or a probability distribution of the respective subsets of training data samples. Additionally, or alternatively, the set of subsets of training data samples are associated with at least one of a respective network condition between the first device and the second device, respective parameters associated with the first device, respective parameters associated with the second device, or respective channel characteristics associated with a channel between the first device and the second device. Additionally, or alternatively, the set of data samples includes at least one of a labeled data sample including a data sample for input to a learning model of the set of learning models and an indication of an expected output from the learning model, or an unlabeled data sample including a data sample for input to the learning model of the set of learning models.
[0129] Additionally, or alternatively, the UE 500 may be configured or operable to support generating, based on providing data as input to the at least one learning model, output from the at least one learning model based on the at least one learning model including the first set of one or more parameters and the at least one second set of one or more parameters. Additionally, or alternatively, the UE 500 may be configured or operable to support transmitting, to the second device, a third message that requests an additional second set of one or more parameters associated with an additional learning model of the set of learning models, and receiving, responsive to the third message, a fourth message including the additional second set of one or more parameters, where the additional learning model includes the first set of one or more parameters and the additional second set of one or more parameters. Additionally, or alternatively, transmitting the first message includes determining one or more conditions are satisfied associated with the first message, and where the one or more conditions are associated with at least one of a network condition between the first device and the second device, one or more parameters associated with Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 49 the first device, one or more parameters with the second device, or one or more channel characteristics associated with a channel between the first device and the second device.
[0130] Additionally, or alternatively, the UE 500 may support at least one memory (e.g., the memory 504) and at least one processor (e.g., the processor 502) coupled with the at least one memory and configured to or operable to cause the UE (e.g., the first device) to transmit, to a second device, a first message that requests at least one of a first set of one or more parameters associated with respective learning models of a set of learning models or at least one second set of one or more parameters associated with at least one learning model of the set of learning models, where each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and receive, responsive to the first message and from the second device, a second message including at least one of the first set of one or more parameters or the at least one second set of one or more parameters.
[0131] Additionally, the UE 500 may be configured or operable to support any one or combination of the subset of layers includes at least one layer associated with classification, an additional subset of layers of the respective set of layers includes a set of layers associated with feature extraction, and the set of layers associated with the feature extraction correspond to the first set of one or more parameters. Additionally, or alternatively, the at least one processor is configured to obtain, as output from the set of layers associated with the feature extraction, one or more attributes of a data sample based on providing the data sample as input to the at least one learning model associated with the first set of one or more parameters and the at least one second set of one or more parameters. Additionally, or alternatively, the at least one second set of one or more parameters is determined based on the at least one learning model being trained using a subset of data samples of a set of data samples. Additionally, or alternatively, the at least one processor is configured to transmit a third message that indicates the set of data samples.
[0132] Additionally, or alternatively, the set of data samples includes a set of subsets of training data samples corresponding to each learning model of the set of learning models, and respective subsets of training data samples of the set of subsets of training data samples are associated with one or more different characteristics. Additionally, or alternatively, the one or more different characteristics include at least one of a mean of the respective subsets of training data samples, a Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 50 standard deviation of the respective subsets of data samples, a variance of the respective subsets of training data samples, or a probability distribution of the respective subsets of training data samples. Additionally, or alternatively, the set of subsets of training data samples are associated with at least one of a respective network condition between the first device and the second device, respective parameters associated with the first device, respective parameters associated with the second device, or respective channel characteristics associated with a channel between the first device and the second device. Additionally, or alternatively, the set of data samples includes at least one of a labeled data sample including a data sample for input to a learning model of the set of learning models and an indication of an expected output from the learning model, or an unlabeled data sample including a data sample for input to the learning model of the set of learning models.
[0133] Additionally, or alternatively, the at least one processor is configured to generate, based on providing data as input to the at least one learning model, output from the at least one learning model based on the at least one learning model including the first set of one or more parameters and the at least one second set of one or more parameters. Additionally, or alternatively, the at least one processor is configured to transmit, to the second device, a third message that requests an additional second set of one or more parameters associated with an additional learning model of the set of learning models, and receive, responsive to the third message, a fourth message including the additional second set of one or more parameters, where the additional learning model includes the first set of one or more parameters and the additional second set of one or more parameters. Additionally, or alternatively, to transmit the first message, the at least one processor is configured to determine one or more conditions are satisfied associated with the first message, where the one or more conditions are associated with at least one of a network condition between the first device and the second device, one or more parameters associated with the first device, one or more parameters associated with the second device, or one or more channel characteristics associated with a channel between the first device and the second device.
[0134] The controller 506 may manage input and output signals for the UE 500. The controller 506 may also manage peripherals not integrated into the UE 500. In some implementations, the controller 506 may utilize an operating system such as iOS®, ANDROID®, WINDOWS®, or other Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 51 operating systems. In some implementations, controller 506 may be implemented as part of the processor 502.
[0135] In some implementations, the UE 500 may include at least one transceiver 508. In some other implementations, the UE 500 may have more than one transceiver 508. The transceiver 508 may represent a wireless transceiver. The transceiver 508 may include one or more receiver chains 510, one or more transmitter chains 512, or a combination thereof.
[0136] A receiver chain 510 may be configured to receive signals (e.g., control information, data, packets) over a wireless medium. For example, the receiver chain 510 may include one or more antennas to receive a signal over the air or wireless medium. The receiver chain 510 may include at least one amplifier (e.g., a low-noise amplifier (LNA)) configured to amplify the received signal. The receiver chain 510 may include at least one demodulator configured to demodulate the receive signal and obtain the transmitted data by reversing the modulation technique applied during transmission of the signal. The receiver chain 510 may include at least one decoder for decoding the demodulated signal to receive the transmitted data.
[0137] A transmitter chain 512 may be configured to generate and transmit signals (e.g., control information, data, packets). The transmitter chain 512 may include at least one modulator for modulating data onto a carrier signal, preparing the signal for transmission over a wireless medium. The at least one modulator may be configured or operable to support one or more techniques such as amplitude modulation (AM), frequency modulation (FM), or digital modulation schemes like phase-shift keying (PSK) or quadrature amplitude modulation (QAM). The transmitter chain 512 may also include at least one power amplifier configured to amplify the modulated signal to an appropriate power level suitable for transmission over the wireless medium. The transmitter chain 512 may also include one or more antennas for transmitting the amplified signal into the air or wireless medium.
[0138] Figure 6 illustrates an example of a processor 600 in accordance with aspects of the present disclosure. The processor 600 may be an example of a processor configured to perform various operations in accordance with examples as described herein. The processor 600 may include a controller 602 configured to perform various operations in accordance with examples as described herein. The processor 600 may optionally include at least one memory 604, which may be, for Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 52 example, an L1 / L2 / L3 cache. Additionally, or the processor 600 may optionally include one or more arithmetic-logic units (ALUs) 606. One or more of these components may be in electronic communication or otherwise coupled (e.g., operatively, communicatively, functionally, electronically, electrically) via one or more interfaces (e.g., buses).
[0139] The processor 600 may be a processor chipset and include a protocol stack (e.g., a software stack) executed by the processor chipset to perform various operations (e.g., receiving, obtaining, retrieving, transmitting, outputting, forwarding, storing, determining, identifying, accessing, writing, reading) in accordance with examples as described herein. The processor chipset may include one or more cores, one or more caches (e.g., memory local to or included in the processor chipset (e.g., the processor 600) or other memory (e.g., random access memory (RAM), read-only memory (ROM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), static RAM (SRAM), ferroelectric RAM (FeRAM), magnetic RAM (MRAM), resistive RAM (RRAM), flash memory, phase change memory (PCM), and others).
[0140] The controller 602 may be configured to manage and coordinate various operations (e.g., signaling, receiving, obtaining, retrieving, transmitting, outputting, forwarding, storing, determining, identifying, accessing, writing, reading) of the processor 600 to cause the processor 600 to support various operations in accordance with examples as described herein. For example, the controller 602 may operate as a control unit of the processor 600, generating control signals that manage the operation of various components of the processor 600. These control signals include enabling or disabling functional units, selecting data paths, initiating memory access, and coordinating timing of operations.
[0141] The controller 602 may be configured to fetch (e.g., obtain, retrieve, receive) instructions from the memory 604 and determine subsequent instruction(s) to be executed to cause the processor 600 to support various operations in accordance with examples as described herein. The controller 602 may be configured to track memory addresses of instructions associated with the memory 604. The controller 602 may be configured to decode instructions to determine the operation to be performed and the operands involved. For example, the controller 602 may be configured to interpret the instruction and determine control signals to be output to other components of the processor 600 to cause the processor 600 to support various operations in accordance with examples as described herein. Additionally, or alternatively, the controller 602 may be configured to manage Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 53 flow of data within the processor 600. The 602 may be configured to control transfer of data between registers, ALUs 606, and other functional units of the processor 600.
[0142] The memory 604 may include one or more caches (e.g., memory local to or included in the processor 600 or other memory, such as RAM, ROM, DRAM, SDRAM, SRAM, MRAM, flash memory, etc. In some implementations, the memory 604 may reside within or on a processor chipset (e.g., local to the processor 600). In some other implementations, the memory 604 may reside external to the processor chipset (e.g., remote to the processor 600).
[0143] The memory 604 may store computer-readable, computer-executable code including instructions that, when executed by the processor 600, cause the processor 600 to perform various functions described herein. The code may be stored in a non-transitory computer-readable medium such as system memory or another type of memory. The controller 602 and / or the processor 600 may be configured to execute computer-readable instructions stored in the memory 604 to cause the processor 600 to perform various functions. For example, the processor 600 and / or the controller 602 may be coupled with or to the memory 604, the processor 600, and the controller 602, and may be configured to perform various functions described herein. In some examples, the processor 600 may include multiple processors and the memory 604 may include multiple memories. One or more of the multiple processors may be coupled with one or more of the multiple memories, which may, individually or collectively, be configured to perform various functions herein.
[0144] The one or more ALUs 606 may be configured or operable to support various operations in accordance with examples as described herein. In some implementations, the one or more ALUs 606 may reside within or on a processor chipset (e.g., the processor 600). In some other implementations, the one or more ALUs 606 may reside external to the processor chipset (e.g., the processor 600). One or more ALUs 606 may perform one or more computations such as addition, subtraction, multiplication, and division on data. For example, one or more ALUs 606 may receive input operands and an operation code, which determines an operation to be executed. One or more ALUs 606 may be configured with a variety of logical and arithmetic circuits, including adders, subtractors, shifters, and logic gates, to process and manipulate the data according to the operation. Additionally, or alternatively, the one or more ALUs 606 may support logical operations such as AND, OR, exclusive-OR (XOR), not-OR (NOR), and not-AND (NAND), enabling the one or more ALUs 606 to handle conditional operations, comparisons, and bitwise operations. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 54
[0145] The processor 600 may support communication in accordance with examples as disclosed herein. The processor 600 may be configured to or operable to support at least one controller (e.g., the controller 602) coupled with at least one memory (e.g., the memory 604) and configured to or operable to cause the processor to transmit, to a device, a first message that requests at least one of a first set of one or more parameters associated with respective learning models of a set of learning models or at least one second set of one or more parameters associated with at least one learning model of the set of learning models, where each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and receive, responsive to the first message and from the device, a second message including at least one of the first set of one or more parameters or the at least one second set of one or more parameters.
[0146] Additionally, the processor 600 may be configured to or operable to support any one or combination of the subset of layers includes at least one layer associated with classification, an additional subset of layers of the respective set of layers includes a set of layers associated with feature extraction, and the set of layers associated with the feature extraction correspond to the first set of one or more parameters. Additionally, or alternatively, the at least one controller is configured to or operable to cause the processor to obtain, as output from the set of layers associated with the feature extraction, one or more attributes of a data sample based on providing the data sample as input to the at least one learning model associated with the first set of one or more parameters and the at least one second set of one or more parameters. Additionally, or alternatively, the at least one second set of one or more parameters is determined based on the at least one learning model being trained using a subset of data samples of a set of data samples. Additionally, or alternatively, The at least one controller is configured to or operable to cause the processor to transmit a third message that indicates the set of data samples.
[0147] Additionally, or alternatively, the set of data samples includes a set of subsets of training data samples corresponding to each learning model of the set of learning models, and respective subsets of training data samples of the set of subsets of training data samples are associated with one or more different characteristics. Additionally, or alternatively, the one or more different characteristics include at least one of a mean of the respective subsets of training data samples, a Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 55 standard deviation of the respective subsets of data samples, a variance of the respective subsets of training data samples, or a probability distribution of the respective subsets of training data samples. Additionally, or alternatively, the set of subsets of training data samples are associated with at least one of a respective network condition between the processor and the device, respective parameters associated with the processor, respective parameters associated with the device, or respective channel characteristics associated with a channel between the processor and the device. Additionally, or alternatively, the set of data samples includes at least one of a labeled data sample including a data sample for input to a learning model of the set of learning models and an indication of an expected output from the learning model, or an unlabeled data sample including a data sample for input to the learning model of the set of learning models.
[0148] Additionally, or alternatively, the at least one controller is configured to or operable to cause the processor to generate, based on providing data as input to the at least one learning model, output from the at least one learning model based on the at least one learning model including the first set of one or more parameters and the at least one second set of one or more parameters. Additionally, or alternatively, the at least one controller is configured to or operable to cause the processor to transmit, to the device, a third message that requests an additional second set of one or more parameters associated with an additional learning model of the set of learning models, and receive, responsive to the third message, a fourth message including the additional second set of one or more parameters, where the additional learning model includes the first set of one or more parameters and the additional second set of one or more parameters. Additionally, or alternatively, to transmit the first message, the at least one controller is configured to or operable to cause the processor to determine one or more conditions are satisfied associated with the first message, and where the one or more conditions are associated with at least one of a network condition between the processor and the device, one or more parameters associated with the processor, one or more parameters associated with the device, or one or more channel characteristics associated with a channel between the processor and the device.
[0149] Figure 7 illustrates an example of an NE 700 in accordance with aspects of the present disclosure. The NE 700 may include a processor 702, a memory 704, a controller 706, and a transceiver 708. The processor 702, the memory 704, the controller 706, or the transceiver 708, or various combinations thereof or various components thereof may be examples of means for Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 56 performing various aspects of the present as described herein. These components may be coupled (e.g., operatively, communicatively, functionally, electronically, electrically) via one or more interfaces.
[0150] The processor 702, the memory 704, the controller 706, or the transceiver 708, or various combinations or components thereof may be implemented in hardware (e.g., circuitry). The hardware may include a processor, a DSP, an ASIC, or other programmable logic device, or any combination thereof configured as or otherwise supporting a means for performing the functions described in the present disclosure.
[0151] The processor 702 may include an intelligent hardware device (e.g., a general-purpose processor, a DSP, a CPU, an ASIC, an FPGA, or any combination thereof). In some implementations, the processor 702 may be configured to operate the memory 704. In some other implementations, the memory 704 may be integrated into the processor 702. The processor 702 may be configured to execute computer-readable instructions stored in the memory 704 to cause the NE 700 to perform various functions of the present disclosure.
[0152] The memory 704 may include volatile or non-volatile memory. The memory 704 may store computer-readable, computer-executable code including instructions when executed by the processor 702 cause the NE 700 to perform various functions described herein. The code may be stored in a non-transitory computer-readable medium such as the memory 704 or another type of memory. Computer-readable media includes both non-transitory computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A non-transitory storage medium may be any available medium that may be accessed by a general-purpose or special-purpose computer.
[0153] In some implementations, the processor 702 and the memory 704 coupled with the processor 702 may be configured to or operable to cause the NE 700 to perform one or more of the functions described herein (e.g., executing, by the processor 702, instructions stored in the memory 704). For example, the processor 702 may support wireless communication at the NE 700 in accordance with examples as disclosed herein. The NE 700 may be configured to or operable to support a means for obtaining a set of learning models, where each learning model of the set of learning models is associated with a first set of one or more parameters, each learning model of the Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 57 set of learning models includes a respective set layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and selecting, for executing at the first device, at least one learning model of the set of learning models or transmit, to a second device, a message including the first set of one or more parameters and at least one second set of one or more parameters of the set of second sets of one or more parameters.
[0154] Additionally, the NE 700 may be configured to or operable to support any one or combination of the method further including the subset of layers includes at least one layer associated with classification, an additional subset of layers of the respective set of layers includes a set of layers associated with feature extraction, and the set of layers associated with the feature extraction correspond to the first set of one or more parameters.
[0155] Additionally, or alternatively, the NE 700 may be configured to or operable to support selecting the at least one learning model of the set of learning models, and obtaining, as output from the set of layers associated with the feature extraction, one or more attributes of a data sample based on providing the data sample as input to the at least one learning model of the set of learning models. Additionally, or alternatively, obtaining the set of learning models includes training, by minimizing a loss function using a subset of data samples of a set of data samples, a learning model of the set of learning models to determine the first set of one or more parameters and an initial second set of one or more parameters of the set of second sets of one or more parameters, where the trained learning model is associated with the first set of one or more parameters and the initial second set of one or more parameters, and training, by minimizing the loss function using respective remaining subsets of data samples of the set of data samples, one or more remaining learning models of the set of learning models to determine respective second sets of one or more parameters associated with the one or more remaining learning models, where the one or more remaining learning models are associated with the first set of one or more parameters and the respective second sets of one or more parameters.
[0156] Additionally, or alternatively, the NE 700 may be configured to or operable to support determining the set of data samples independent of the second device. Additionally, or alternatively, the NE 700 may be configured to or operable to support receiving an additional message that indicates the set of data samples. Additionally, or alternatively, the set of data samples includes a set Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 58 of subsets of training data samples to each learning model of the set of learning models, and respective subsets of training data samples of the set of subsets of training data samples are associated with one or more different characteristics. Additionally, or alternatively, the one or more different characteristics include at least one of a mean of the respective subsets of training data samples, a standard deviation of the respective subsets of training data samples, a variance of the respective subsets of training data samples, or a probability distribution of the respective subsets of training data samples. Additionally, or alternatively, the set of subsets of training data samples are associated with at least one of a respective network condition between the first device and the second device, respective parameters associated with the first device, respective parameters associated with the second device, or respective channel characteristics associated with a channel between the first device and the second device. Additionally, or alternatively, the set of data samples includes at least one of a labeled data sample including a data sample for input to respective learning models of the set of learning models and an indication of an expected output from the respective learning models, or an unlabeled data sample including a data sample for input to the respective learning models of the set of learning models.
[0157] Additionally, or alternatively, the NE 700 may be configured to or operable to support selecting the at least one learning model of the set of learning models, where the at least one learning model is associated with the first set of one or more parameters and the at least one second set of one or more parameters, and generating, based on providing data as input to the at least one learning model of the set of learning models, output from the at least one learning model. Additionally, or alternatively, the NE 700 may be configured to or operable to support updating the at least one second set of one or more parameters associated with the at least one learning model to an additional second set of one or more parameters of the set of second sets of one or more parameters, and where the additional second set of one or more parameters is associated with an additional learning model of the set of learning models. Additionally, or alternatively, the NE 700 may be configured to or operable to support transmitting, to the second device, the message including the first set of one or more parameters and the at least one second set of one or more parameters of the set of second sets of one or more parameters, and transmitting, to the second device and after the message is transmitted, an additional message including an additional second set of one or more parameters of the set of second sets of one or more parameters, where the first set Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 59 of one or more parameters and the additional set of one or more parameters are associated with an additional learning model of the set of learning models.
[0158] Additionally, or alternatively, the NE 700 may be configured to or operable to support receiving, from the second device, an additional message that indicates for the first device to transmit the message, and transmitting, responsive to the additional message and to the second device, the message including the first set of one or more parameters and the at least one second set of one or more parameters of the set of second sets of one or more parameters. Additionally, or alternatively, the additional message requests at least one of the first set of one or more parameters or the at least one second set of one or more parameters of the set of second sets of one or more parameters. Additionally, or alternatively, transmitting the message includes determining one or more conditions are satisfied associated with the message, and where the one or more conditions are associated with at least one of a network condition between the first device and the second device, one or more parameters associated with the first device, one or more parameters associated with the second device, or one or more channel characteristics associated with a channel between the first device and the second device.
[0159] Additionally, or alternatively, the NE 700 may support at least one memory (e.g., the memory 704) and at least one processor (e.g., the processor 702) coupled with the at least one memory and configured to or operable to cause the NE to obtain a set of learning models, where each learning model of the set of learning models is associated with a first set of one or more parameters, each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a set of second sets of one or more parameters, and select, for executing at the first device, at least one learning model of the set of learning models or transmit, to a second device, a message including the first set of one or more parameters and at least one second set of one or more parameters of the set of second sets of one or more parameters.
[0160] Additionally, the NE 700 may be configured or operable to support any one or combination of the subset of layers includes at least one layer associated with classification, an additional subset of layers of the respective set of layers includes a set of layers associated with feature extraction, and the set of layers associated with the feature extraction correspond to the first set of one or more parameters. Additionally, or alternatively, the at least one processor is configured Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 60 to or operable to cause the NE to select the at one learning model of the set of learning models, and obtain, as output from the set of layers associated with the feature extraction, one or more attributes of a data sample based on providing the data sample as input to the at least one learning model of the set of learning models. Additionally, or alternatively, to obtain the set of learning models, the at least one processor is configured to or operable to cause the NE to train, by minimizing a loss function using a subset of data samples of a set of data samples, a learning model of the set of learning models to determine the first set of one or more parameters and an initial second set of one or more parameters of the set of second sets of one or more parameters, where the trained learning model is associated with the first set of one or more parameters and the initial second set of one or more parameters, and trains, by minimizing the loss function using respective remaining subsets of data samples of the set of data samples, one or more remaining learning models of the set of learning models to determine respective second sets of one or more parameters associated with the one or more remaining learning models, where the one or more remaining learning models are associated with the first set of one or more parameters and the respective second sets of one or more parameters.
[0161] Additionally, or alternatively, the at least one processor is configured to or operable to cause the NE to determine the set of data samples independent of the second device. Additionally, or alternatively, the at least one processor is configured to or operable to cause the NE to receive an additional message that indicates the set of data samples. Additionally, or alternatively, the set of data samples includes a set of subsets of training data samples corresponding to each learning model of the set of learning models, and respective subsets of training data samples of the set of subsets of training data samples are associated with one or more different characteristics. Additionally, or alternatively, the one or more different characteristics include at least one of a mean of the respective subsets of training data samples, a standard deviation of the respective subsets of training data samples, a variance of the respective subsets of training data samples, or a probability distribution of the respective subsets of training data samples. Additionally, or alternatively, the set of subsets of training data samples are associated with at least one of a respective network condition between the first device and the second device, respective parameters associated with the first device, respective parameters associated with the second device, or respective channel characteristics associated with a channel between the first device and the second device. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 61 Additionally, or alternatively, the set of data includes at least one of a labeled data sample including a data sample for input to respective learning models of the set of learning models and an indication of an expected output from the respective learning models, or an unlabeled data sample including a data sample for input to the respective learning models of the set of learning models.
[0162] Additionally, or alternatively, the at least one processor is configured to or operable to cause the NE to select the at least one learning model of the set of learning models, where the at least one learning model is associated with the first set of one or more parameters and the at least one second set of one or more parameters, and generate, based on providing data as input to the at least one learning model of the set of learning models, output from the at least one learning model. Additionally, or alternatively, the at least one processor is configured to or operable to cause the NE to update the at least one second set of one or more parameters associated with the at least one learning model to an additional second set of one or more parameters of the set of second sets of one or more parameters, where the additional second set of one or more parameters is associated with an additional learning model of the set of learning models. Additionally, or alternatively, the at least one processor is configured to or operable to cause the NE to transmit, to the second device, the message including the first set of one or more parameters and the at least one second set of one or more parameters of the set of second sets of one or more parameters, and transmits, to the second device and after the message is transmitted, an additional message including an additional second set of one or more parameters of the set of second sets of one or more parameters, where the first set of one or more parameters and the additional second set of one or more parameters are associated with an additional learning model of the set of learning models.
[0163] Additionally, or alternatively, the at least one processor is configured to or operable to cause the NE to receive, from the second device, an additional message that indicates for the first device to transmit the message, and transmit, responsive to the additional message and to the second device, the message including the first set of one or more parameters and the at least one second set of one or more parameters of the set of second sets of one or more parameters. Additionally, or alternatively, the additional message requests at least one of the first set of one or more parameters or the at least one second set of one or more parameters of the plurality of second sets of one or more parameters. Additionally, or alternatively, to transmit the message, the at least one processor is configured to or operable to cause the NE to determine one or more conditions are satisfied Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 62 associated with the message, where the one or conditions are associated with at least one of a network condition between the first device and the second device, one or more parameters associated with the first device, one or more parameters associated with the second device, or one or more channel characteristics associated with a channel between the first device and the second device.
[0164] The controller 706 may manage input and output signals for the NE 700. The controller 706 may also manage peripherals not integrated into the NE 700. In some implementations, the controller 706 may utilize an operating system such as iOS®, ANDROID®, WINDOWS®, or other operating systems. In some implementations, the controller 706 may be implemented as part of the processor 702.
[0165] In some implementations, the NE 700 may include at least one transceiver 708. In some other implementations, the NE 700 may have more than one transceiver 708. The transceiver 708 may represent a wireless transceiver. The transceiver 708 may include one or more receiver chains 710, one or more transmitter chains 712, or a combination thereof.
[0166] A receiver chain 710 may be configured to receive signals (e.g., control information, data, packets) over a wireless medium. For example, the receiver chain 710 may include one or more antennas to receive a signal over the air or wireless medium. The receiver chain 710 may include at least one amplifier (e.g., a low-noise amplifier (LNA)) configured to amplify the received signal. The receiver chain 710 may include at least one demodulator configured to demodulate the receive signal and obtain the transmitted data by reversing the modulation technique applied during transmission of the signal. The receiver chain 710 may include at least one decoder for decoding the demodulated signal to receive the transmitted data.
[0167] A transmitter chain 712 may be configured to generate and transmit signals (e.g., control information, data, packets). The transmitter chain 712 may include at least one modulator for modulating data onto a carrier signal, preparing the signal for transmission over a wireless medium. The at least one modulator may be configured or operable to support one or more techniques such as amplitude modulation (AM), frequency modulation (FM), or digital modulation schemes like phase-shift keying (PSK) or quadrature amplitude modulation (QAM). The transmitter chain 712 may also include at least one power amplifier configured to amplify the modulated signal to an Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 63 appropriate power level suitable for over the wireless medium. The transmitter chain 712 may also include one or more antennas for transmitting the amplified signal into the air or wireless medium.
[0168] Figure 8 illustrates a flowchart of a method 800 in accordance with aspects of the present disclosure. The operations of the method may be implemented by a UE as described herein. In some implementations, the UE may execute a set of instructions to control the function elements of the UE to perform the described functions. It should be noted that the method described herein describes a possible implementation, and that the operations and the steps may be rearranged or otherwise modified and that other implementations are possible.
[0169] At 802, the method may include transmit, to a second device, a first message that requests at least one of a first set of one or more parameters associated with respective learning models of a set of learning models or at least one second set of one or more parameters associated with at least one learning model of the set of learning models, where each learning model of the set of learning models includes a respective set of layers and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of multiple second sets of one or more parameters. The operations of 802 may be performed in accordance with examples as described herein. In some implementations, aspects of the operations of 802 may be performed by a UE as described with reference to Figure 5.
[0170] At 804, the method may include receive, responsive to the first message and from the second device, a second message including at least one of the first set of one or more parameters or the at least one second set of one or more parameters. The operations of 804 may be performed in accordance with examples as described herein. In some implementations, aspects of the operations of 804 may be performed by a UE as described with reference to Figure 5.
[0171] Figure 9 illustrates a flowchart of a method 900 in accordance with aspects of the present disclosure. The operations of the method may be implemented by an NE as described herein. In some implementations, the NE may execute a set of instructions to control the function elements of the NE to perform the described functions. It should be noted that the method described herein describes a possible implementation, and that the operations and the steps may be rearranged or otherwise modified and that other implementations are possible. Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 64
[0172] At 902, the method may include a set of learning models, where each learning model of the set of learning models is associated with a first set of one or more parameters, each learning model of the set of learning models includes a respective set of layers, and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of multiple second sets of one or more parameters. The operations of 902 may be performed in accordance with examples as described herein. In some implementations, aspects of the operations of 902 may be performed by an NE as described with reference to Figure 7.
[0173] At 904, the method may include select, for executing at the first device, at least one learning model of the set of learning models or transmit, to a second device, a message including the first set of one or more parameters and at least one second set of one or more parameters of the multiple second sets of one or more parameters. The operations of 904 may be performed in accordance with examples as described herein. In some implementations, aspects of the operations of 904 may be performed by an NE as described with reference to Figure 7.
[0174] The description herein is provided to enable a person having ordinary skill in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to a person having ordinary skill in the art, and the generic principles defined herein may be applied to other variations without departing from the scope of the disclosure. Thus, the disclosure is not limited to the examples and designs described herein but is to be accorded the broadest scope consistent with the principles and novel features disclosed herein. Firm Ref. No. SMM920240006-WO-PCT
Claims
Lenovo Ref. No. SMM920240006-WO-PCT 65 What is claimed is:
1. A first device for wireless communication, comprising: at least one memory; and at least one processor coupled with the at least one memory and operable to cause the first device to: obtain a set of learning models, wherein: each learning model of the set of learning models is associated with a first set of one or more parameters; each learning model of the set of learning models comprises a respective set of layers; and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a plurality of second sets of one or more parameters; and select, for executing at the first device, at least one learning model of the set of learning models or transmit, to a second device, a message comprising the first set of one or more parameters and at least one second set of one or more parameters of the plurality of second sets of one or more parameters.
2. The first device of claim 1, wherein: the subset of layers comprises at least one layer associated with classification; an additional subset of layers of the respective set of layers comprises a plurality of layers associated with feature extraction; and the plurality of layers associated with the feature extraction correspond to the first set of one or more parameters.
3. The first device of claim 2, wherein the at least one processor is further configured to cause the first device to: select the at least one learning model of the set of learning models; and Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 66 obtain, as output from the plurality of associated with the feature extraction, one or more attributes of a data sample based at least in part on providing the data sample as input to the at least one learning model of the set of learning models.
4. The first device of claim 1, wherein to obtain the set of learning models, the at least one processor is configured to cause the first device to: train, by minimizing a loss function using a subset of data samples of a set of data samples, a learning model of the set of learning models to determine the first set of one or more parameters and an initial second set of one or more parameters of the plurality of second sets of one or more parameters, wherein the trained learning model is associated with the first set of one or more parameters and the initial second set of one or more parameters; and train, by minimizing the loss function using respective remaining subsets of data samples of the set of data samples, one or more remaining learning models of the set of learning models to determine respective second sets of one or more parameters associated with the one or more remaining learning models, wherein the one or more remaining learning models are associated with the first set of one or more parameters and the respective second sets of one or more parameters.
5. The first device of claim 4, wherein the at least one processor is further configured to cause the first device to determine the set of data samples independent of the second device.
6. The first device of claim 4, wherein the at least one processor is further configured to cause the first device to receive an additional message that indicates the set of data samples.
7. The first device of claim 4, wherein: the set of data samples comprises a plurality of subsets of training data samples corresponding to each learning model of the set of learning models; and respective subsets of training data samples of the plurality of subsets of training data samples are associated with one or more different characteristics, wherein the one or more different characteristics comprise at least one of a mean of the respective subsets of training data samples, a standard deviation of the respective subsets of training data samples, a variance of the respective subsets of training data samples, or a probability distribution of the respective subsets of training data samples, and wherein the plurality of subsets of training data samples are associated with at least one of a respective network condition between the first device and the second device, Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 67 respective parameters associated with the first respective parameters associated with the second device, or respective channel characteristics associated with a channel between the first device and the second device.
8. The first device of claim 4, wherein the set of data samples comprises at least one of: a labeled data sample comprising a data sample for input to respective learning models of the set of learning models and an indication of an expected output from the respective learning models; or an unlabeled data sample comprising a data sample for input to the respective learning models of the set of learning models.
9. The first device of claim 1, wherein the at least one processor is further configured to cause the first device to: select the at least one learning model of the set of learning models, wherein the at least one learning model is associated with the first set of one or more parameters and the at least one second set of one or more parameters; and generate, based at least in part on providing data as input to the at least one learning model of the set of learning models, output from the at least one learning model.
10. The first device of claim 9, wherein the at least one processor is further configured to cause the first device to update the at least one second set of one or more parameters associated with the at least one learning model to an additional second set of one or more parameters of the plurality of second sets of one or more parameters, and wherein the additional second set of one or more parameters is associated with an additional learning model of the set of learning models.
11. The first device of claim 1, wherein the at least one processor is further configured to cause the first device to: transmit, to the second device, the message comprising the first set of one or more parameters and the at least one second set of one or more parameters of the plurality of second sets of one or more parameters; and transmit, to the second device and after the message is transmitted, an additional message comprising an additional second set of one or more parameters of the plurality of second sets of one or more parameters, wherein the first set of one or more parameters and the additional second set of Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 68 one or more parameters are associated with an learning model of the set of learning models.
12. The first device of claim 1, wherein the at least one processor is further configured to cause the first device to: receive, from the second device, an additional message that indicates for the first device to transmit the message, wherein the additional message requests at least one of the first set of one or more parameters or the at least one second set of one or more parameters of the plurality of second sets of one or more parameters; and transmit, responsive to the additional message and to the second device, the message comprising the first set of one or more parameters and the at least one second set of one or more parameters of the plurality of second sets of one or more parameters.
13. The first device of claim 1, wherein to transmit the message, the at least one processor is configured to cause the first device to determine one or more conditions are satisfied associated with the message, and wherein the one or more conditions are associated with at least one of a network condition between the first device and the second device, one or more parameters associated with the first device, one or more parameters associated with the second device, or one or more channel characteristics associated with a channel between the first device and the second device.
14. A method performed by a first device, the method comprising: obtaining a set of learning models, wherein: each learning model of the set of learning models is associated with a first set of one or more parameters; each learning model of the set of learning models comprises a respective set of layers; and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a plurality of second sets of one or more parameters; and selecting, for executing at the first device, at least one learning model of the set of learning models or transmit, to a second device, a message comprising the first set of one or more parameters Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 69 and at least one second set of one or more of the plurality of second sets of one or more parameters.
15. A first device for wireless communication, comprising: at least one memory; and at least one processor coupled with the at least one memory and operable to cause the first device to: transmit, to a second device, a first message that requests at least one of a first set of one or more parameters associated with respective learning models of a set of learning models or at least one second set of one or more parameters associated with at least one learning model of the set of learning models, wherein: each learning model of the set of learning models comprises a respective set of layers; and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a plurality of second sets of one or more parameters; and receive, responsive to the first message and from the second device, a second message comprising at least one of the first set of one or more parameters or the at least one second set of one or more parameters.
16. The first device of claim 15, wherein: the subset of layers comprises at least one layer associated with classification; an additional subset of layers of the respective set of layers comprises a plurality of layers associated with feature extraction; and the plurality of layers associated with the feature extraction correspond to the first set of one or more parameters.
17. The first device of claim 15, wherein the at least one second set of one or more parameters is determined based at least in part on the at least one learning model being trained using a subset of data samples of a set of data samples.
18. The first device of claim 15, wherein the at least one processor is further configured to cause the first device to generate, based at least in part on providing data as input to the at least Firm Ref. No. SMM920240006-WO-PCTLenovo Ref. No. SMM920240006-WO-PCT 70 one learning model, output from the at least model based at least in part on the at least one learning model comprising the first set of one or more parameters and the at least one second set of one or more parameters.
19. The first device of claim 15, wherein to transmit the first message, the at least one processor is configured to cause the first device to determine one or more conditions are satisfied associated with the first message, and wherein the one or more conditions are associated with at least one of a network condition between the first device and the second device, one or more parameters associated with the first device, one or more parameters associated with the second device, or one or more channel characteristics associated with a channel between the first device and the second device.
20. A method performed by a first device, the method comprising: transmitting, to a second device, a first message that requests at least one of a first set of one or more parameters associated with respective learning models of a set of learning models or at least one second set of one or more parameters associated with at least one learning model of the set of learning models, wherein: each learning model of the set of learning models comprises a respective set of layers; and a subset of layers of the respective set of layers is associated with a second set of one or more parameters of a plurality of second sets of one or more parameters; and receiving, responsive to the first message and from the second device, a second message comprising at least one of the first set of one or more parameters or the at least one second set of one or more parameters. Firm Ref. No. SMM920240006-WO-PCT
Citation Information
Patent Citations
Configuring a user equipment for machine learning
US11916754B2
Customized classifier over common features
US20150324689A1
Managing a wireless device that is operable to connect to a communication network
US20240049003A1
Artificial intelligence and machine learning models management and / or training
WO2024010399A1
US202463574830P
Cited By
Adapting beam prediction neural network models
WO2026133312A1