Apparatus and method for machine learning with low training latency and communication overhead
By determining and transmitting AI/ML model training capabilities, the method reduces training latency and communication overhead in wireless communications systems, enabling faster convergence through selective device participation.
Patent Information
- Application Number
- JP2024554740
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-03-15
AI Technical Summary
Conventional AI training processes in wireless communications systems face high communication overhead and latency due to non-ideal channels and reliance on hybrid automatic repeat request (HARQ) feedback and retransmissions.
A method for determining and transmitting AI/ML model training capability feedback to selectively include or exclude devices based on their processing power, training data volume, and sensing capacity, reducing training latency and communication overhead.
This approach achieves faster training convergence and reduces communication overhead by selectively participating devices in AI/ML model training procedures based on their reported capabilities.
Smart Images

Figure 0007775501000003 
Figure 0007775501000004 
Figure 0007775501000005
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to wireless communications and, in particular embodiments, to methods and apparatus for machine learning with low training latency and low communication overhead. [Background technology]
[0002] Artificial intelligence (AI) techniques may be applied to communications, including AI-based communications in the physical layer and / or in the medium access control (MAC) layer. For example, in the physical layer, AI-based communications may aim to optimize component designs and / or improve algorithm performance. In the case of the MAC layer, AI-based communications may aim to utilize AI capabilities for learning, prediction, and / or decision-making to solve complex optimization problems with better strategies and / or optimal solutions possible to optimize functions in the MAC layer.
[0003] In some implementations, AI architectures in wireless communication networks may involve multiple nodes, which may be organized into one of two modes: centralized mode and distributed mode, both of which may be deployed in access networks, core networks, or edge computing systems or third-party networks. Centralized training and computing architectures are sometimes limited by large communication overhead and strict user data privacy. Distributed training and computing architectures may comprise several frameworks, such as distributed machine learning and federated learning.
[0004] However, communications in wireless communications systems, including communications related to AI training at multiple nodes, generally occur over non-ideal channels. Non-ideal conditions, such as electromagnetic interference, signal degradation, phase delay, fading, and other non-idealities, can attenuate and / or distort communication signals or otherwise interfere with or degrade the communication capabilities of the system.
[0005] Conventional AI training processes generally rely on hybrid automatic repeat request (HARQ) feedback and retransmission processes to ensure that data communicated between devices involved in AI training is successfully received. However, the communication overhead and delays associated with such retransmissions can be problematic. Summary of the Invention
[0006] According to a first broad aspect of the present disclosure, a method for transmitting artificial intelligence or machine learning (AI / ML) data in a wireless communications network is provided herein. The method according to the first broad aspect of the present disclosure may include determining a training capability of an AI / ML model of a first device, where the training capability of the AI / ML model indicates the capability of the first device to contributively participate in a training process of the AI / ML model with at least a second device in the wireless communications network. For example, the training capability of the AI / ML model may be determined based on i) a current processing capability of the first device, ii) a current volume of training data available for the training process of the AI / ML model at the first device, and / or iii) a sensing capacity of the first device to collect training data for the training process of the AI / ML model. The method according to the first broad aspect of the present disclosure may further include transmitting the training capability of the AI / ML model to the second device.
[0007] Providing AI / ML model training capability feedback in accordance with a first broad aspect of the present disclosure can have several advantages. For example, the AI / ML model training capability feedback can be utilized to selectively include or exclude a first device from participating in one or more iterations of an AI / ML model training procedure based on the device's currently reported AI / ML model training capability, which, as described in further detail herein, can potentially reduce training latency, thereby achieving faster training convergence and / or reducing communication overhead associated with the AI / ML model training procedure.
[0008] In some embodiments, determining the training capability of the AI / ML model of the first device includes selecting a type of training capability of the AI / ML model from among a predefined or configured hierarchy of types of training capability of the AI / ML model. In such embodiments, transmitting the training capability of the AI / ML model to the second device may include transmitting an index corresponding to the type of training capability of the AI / ML model selected from among a predefined or configured hierarchy of types of training capability of the AI / ML model.
[0009] In some embodiments, transmitting the AI / ML model training capability to the second device includes determining that the AI / ML model training capability of the first device has changed and, after determining that the AI / ML model training capability of the first device has changed, transmitting the changed AI / ML model training capability to the second device. In such embodiments, determining that the AI / ML model training capability of the first device has changed may include, for example, identifying a change in at least one of: i) a current processing capability of the first device, ii) a current volume of training data available for the AI / ML model training process at the first device, and / or iii) a sensing capacity of the first device for collecting training data for the AI / ML model training process.
[0010] In some embodiments, transmitting the training capability of the AI / ML model to the second device occurs after receiving data or control information from the second device during a training process of the AI / ML model. For example, in some embodiments, the first device may receive control signaling from the second device that identifies physical uplink control channel (PUCCH) resources to be used by the first device and transmit the training capability of the AI / ML model to the second device using the PUCCH resources. In such embodiments, receiving control signaling from the second device that identifies PUCCH resources to be used by the first device may include receiving downlink control information (DCI) from the second device that schedules downlink transmission of data or control information from the second device during the training process of the AI / ML model, the DCI indicating the PUCCH resources to be used by the first device.
[0011] In some embodiments, the data or control information received from the second device includes AI / ML model update information from the second device for an AI / ML model training process.
[0012] In some embodiments, a method according to a first broad aspect of this disclosure further includes receiving control signaling from the second device for an iteration of a training process of the AI / ML model, the control signaling configuring the first device with rules for determining whether the first device will participate in the iteration.
[0013] In some embodiments, the method according to the first broad aspect of this disclosure further includes participating in one or more iterations of a training process of the AI / ML model according to the configured rules.
[0014] In some embodiments, iterations of the training process for the AI / ML model are associated with respective values of iteration identifiers (IDs) such that the respective values of the iteration IDs are incremented by 1 for each subsequent iteration. In such embodiments, control signaling may configure the first device to selectively participate in a given iteration based on the respective values of the iteration IDs associated with the given iteration.
[0015] In some embodiments, a method according to a first broad aspect of this disclosure further includes, for a given iteration of the training process of the AI / ML model, receiving control information from the second device indicating a value of an iteration ID associated with the given iteration.
[0016] In some embodiments, a method according to a first broad aspect of this disclosure further includes, for a given iteration of the AI / ML model training process, receiving control information from the second device indicating a type of training capability of at least one AI / ML model from among a predefined or configured hierarchy of types of training capabilities of the AI / ML models participating in the given iteration.
[0017] In some embodiments, transmission of data or control information for a given iteration of the AI / ML model training process from the second device is scheduled by a first downlink control information (DCI), and a cyclic redundancy check (CRC) value of the first DCI is scrambled using a first radio network temporary identifier (RNTI). In such embodiments, a method according to a first broad aspect of the present disclosure may further include receiving control signaling from the second device to configure the first device to monitor the first DCI according to a first monitoring period.
[0018] In some embodiments, transmission of other data or control information from the second device is scheduled by a second DCI, and a CRC value of the second DCI is scrambled using a second RNTI different from the first RNTI. In such embodiments, a method according to a first broad aspect of this disclosure may further include receiving device-specific control signaling from the second device that configures the first device to monitor the second DCI according to a second monitoring period.
[0019] In some embodiments, the second monitoring period is configured separately from the first monitoring period.
[0020] In some embodiments, the first RNTI is different from the Cell RNTI (C-RNTI).
[0021] In some embodiments, a method according to a first broad aspect of the present disclosure further includes sending local AI / ML model update information to the second device for a given iteration of an AI / ML model training process, the local AI / ML model update information including parameter updates for the AI / ML model based on training of the local AI / ML model at the first device. In such embodiments, the training of the local AI / ML model at the first device may be based on data or control information received from the second device for the given iteration of the AI / ML model training process. For example, the local AI / ML model update information may further include information indicating a value of an iteration ID associated with the given iteration for which the first device received data or control information from the second device.
[0022] In some embodiments, a method according to the first broad aspect of this disclosure further includes sending a request to the second device to participate in a training process of the AI / ML model.
[0023] In some embodiments, an iteration of the AI / ML model training process is associated with a respective value of an iteration identifier (ID), such that the respective value of the ID is incremented by 1 for each subsequent iteration. In such embodiments, a method according to a first broad aspect of the present disclosure may further include receiving a transmission from the second device indicating a value of the iteration ID associated with a current iteration of the AI / ML model training process, and transmitting a request to participate in the AI / ML model training process is based on the value of the iteration ID associated with the current iteration of the AI / ML model training process.
[0024] In some embodiments, sending a request to the second device to participate in the training process of the AI / ML model includes sending a request to the second device to participate in the training process of the AI / ML model after determining that the training capability of the AI / ML model of the first device has changed.
[0025] In some embodiments, a method according to a first broad aspect of this disclosure further includes receiving control signaling from the second device that configures the first device to train a partial AI / ML model that includes a partial subset of parameters of the local AI / ML model at the first device.
[0026] In some embodiments, a method according to a first broad aspect of this disclosure further includes transmitting local AI / ML model updates to the second device, the local AI / ML model updates including parameter updates of the AI / ML model for a partial subset of parameters based on training of the partial AI / ML model.
[0027] In some embodiments, a method according to a first broad aspect of this disclosure further includes receiving control signaling from the second device that configures the first device not to participate in the training process of the AI / ML model.
[0028] In some embodiments, the method according to the first broad aspect of this disclosure further includes sending a request to the second device that the first device does not want to participate in the training process of the AI / ML model.
[0029] In some embodiments, a method according to a first broad aspect of this disclosure further includes, for a given iteration of the AI / ML model training process, sending a local AI / ML model update to the second device. In some such embodiments, the local AI / ML model update may include AI / ML model parameter updates for only a partial subset of the AI / ML model parameters that characterize the local AI / ML model at the first device based on training of the local AI / ML model at the first device.
[0030] In some embodiments, the local AI / ML model update information includes value information including AI / ML model parameter update values for a partial subset of the AI / ML model parameters that characterize the local AI / ML model at the first device, and assignment information that maps the AI / ML model parameter update values to corresponding AI / ML model parameters of the local AI / ML model.
[0031] In some embodiments, the parameter update values of the AI / ML model are arranged in a predefined or configured order in the value information.
[0032] In some embodiments, the assignment information includes a bitmap that maps parameter groups (PGs) of the local AI / ML model to parameter update values of the corresponding AI / ML model in the value information, where each PG is a set of consecutive parameters of the local AI / ML model.
[0033] In some embodiments, the size of each PG is predefined or configured by the second device.
[0034] In some embodiments, the allocation information indicates a set of parameters for consecutive AI / ML models of the local AI / ML model, and the allocation information includes starting locations of the parameters for the AI / ML models in the set and parameters for some of the AI / ML models in the set.
[0035] In some embodiments, the allocation information indicates multiple sets of parameters for successive AI / ML models of the local AI / ML model, and for each set, the allocation information includes a starting location of the parameters for the AI / ML models in the set and the parameters for some of the AI / ML models in the set.
[0036] In some embodiments, the local AI / ML model has a multi-layer structure, and the allocation information indicates one or more sets of parameters of the AI / ML model between two layers of the local AI / ML model.
[0037] In some embodiments, the value information comprises sets of one or more AI / ML model parameter update values, and for each set of AI / ML model parameter update values, the value information indicates a representation of each value represented as a bit string for each AI / ML model parameter update value in the set of AI / ML model parameter update values and a range ID value associated with the set of AI / ML model parameter update values. For example, in some embodiments, the range ID value may be represented as one or more bits and selected from a plurality of range ID values, each range ID value of the plurality of range ID values mapping to a different respective value range. In such embodiments, the respective value ranges mapped to the range ID value associated with the set of AI / ML model parameter update values may determine the range and bit meaning of the bit string for the AI / ML model parameter update value in the set of AI / ML model parameter update values.
[0038] In some embodiments, the value information includes at least a set of parameter update values for a first AI / ML model and a set of parameter update values for a second AI / ML model. In such embodiments, the value information for the set of parameter update values for the first AI / ML model may indicate a representation of each value, represented as a bit string, for each parameter update value of the AI / ML model in the set of parameter update values for the first AI / ML model, and a first range ID value associated with the set of parameter update values for the first AI / ML model, the first range ID value being represented as one or more bits and selected from a plurality of range ID values, the first range ID value mapping to the first range of values. Additionally or alternatively, in such embodiments, the value information for the set of parameter update values for the second AI / ML model may indicate a representation of each value, represented as a bit string, for each parameter update value of the AI / ML model in the set of parameter update values for the second AI / ML model, and a second range ID value associated with the set of parameter update values for the second AI / ML model, the second range ID value being represented as one or more bits and selected from a plurality of range ID values. For example, in some embodiments, the second range ID value may be different from the first range ID value and may be mapped to a second range of values that is different from the first range of values.
[0039] In some embodiments, the mapping between range IDs and respective value ranges is predefined or configured by the second device.
[0040] According to a second broad aspect of the present disclosure, another method for transmitting artificial intelligence or machine learning (AI / ML) data in a wireless communications network is provided herein. The method according to the second broad aspect of the present disclosure may include receiving a training capability of an AI / ML model from a first device, the training capability of the AI / ML model from the first device indicating a capability of the first device to contribute to a training process of the AI / ML model with at least a second device in the wireless communications network; and transmitting, for each of at least one iteration of the training process of the AI / ML model, information that enables the first device to determine whether the first device will participate in the iteration based on the training capability of the AI / ML model received from the first device.
[0041] Providing information that enables a first device to decide whether to participate in one or more iterations of an AI / ML model training procedure in accordance with a second broad aspect of the present disclosure, and basing that information on feedback of the AI / ML model's training capability from the device, can have several advantages. For example, selectively including or excluding a device from participating in one or more iterations of an AI / ML model training procedure based on the device's currently reported AI / ML model training capability can potentially reduce training latency, thereby achieving faster training convergence and / or reducing communication overhead associated with the AI / ML model training procedure, as described in further detail herein.
[0042] In some embodiments, the AI / ML model training capability includes i) the current processing power of the first device, ii) the current volume of training data available for the iterative AI / ML model training process at the first device, and / or iii) the sensing capacity of the first device to collect training data for the iterative AI / ML model training process.
[0043] In some embodiments, the AI / ML model training capabilities of the first device include a type of AI / ML model training capability selected from among a predefined or configured hierarchy of types of AI / ML model training capabilities.
[0044] In some embodiments, receiving the AI / ML model training capability from the first device includes receiving an index corresponding to a type of AI / ML model training capability selected from among a predefined or configured hierarchy of types of AI / ML model training capabilities.
[0045] In some embodiments, the method according to the second broad aspect of this disclosure further includes transmitting control signaling scheduling transmission of data or control information for a given iteration of a training process of the AI / ML model. In such embodiments, receiving the training capability of the AI / ML model from the first device may include receiving the training capability of the AI / ML model from the first device after the control signaling scheduling transmission of data or control information for the given iteration is transmitted.
[0046] In some embodiments, the control signaling scheduling transmission of data or control information for a given iteration includes control information identifying physical uplink control channel (PUCCH) resources to be used by the first device for the given iteration. In such embodiments, receiving the training capability of the AI / ML model from the first device includes receiving the training capability of the AI / ML model for the given iteration on PUCCH resources.
[0047] In some embodiments, the control signaling scheduling the transmission of data or control information for a given repetition includes downlink control information (DCI), which indicates a PUCCH resource to be used by the first device for the given repetition.
[0048] In some embodiments, the data or control information includes updates to the AI / ML model from the second device that include parameter updates to the AI / ML model based on training of the AI / ML model on the second device.
[0049] In some embodiments, receiving the training capabilities of the AI / ML model from the first device includes receiving the training capabilities of the respective AI / ML model from each device of a plurality of devices including the first device. In such embodiments, the second device may transmit, for each of at least one iteration of the training process of the AI / ML model, information that enables the device to determine whether each device of the plurality of devices will participate in the iteration based on the training capabilities of the respective AI / ML model received from each device of the plurality of devices.
[0050] In some embodiments, for each of at least one iteration of the training process of the AI / ML model, transmitting information that enables each device to determine whether it will participate in the iteration includes transmitting control signaling for each device of the plurality of devices to configure each device of the plurality of devices with device-specific rules for determining, for each iteration of the training process of the AI / ML model, whether it will participate in the iteration.
[0051] In some embodiments, iterations of the AI / ML model training process are associated with respective values of iteration identifiers (IDs) such that the respective values of the iteration IDs are incremented by 1 for each subsequent iteration. In such embodiments, device-specific rules for configuring a plurality of devices may configure each device of the plurality of devices to selectively participate in a given iteration based on the respective values of the iteration IDs associated with the given iteration.
[0052] In some embodiments, a method according to a second broad aspect of this disclosure further includes transmitting control information indicating, for an iteration of a training process for an AI / ML model, a value of an iteration ID associated with a given iteration.
[0053] In some embodiments, for each of at least one iteration of the AI / ML model training process, transmitting information that enables the first device to determine whether the first device will participate in the iteration includes transmitting, for the iteration of the AI / ML model training process, control information that indicates a type of training capability of at least one AI / ML model from among a predefined or configured hierarchy of types of training capabilities of the AI / ML models that will participate in the iteration.
[0054] In some embodiments, for each of at least one iteration of the AI / ML model training process, transmitting information that enables the first device to determine whether to participate in the iteration includes transmitting first downlink control information (DCI) including first scheduling information for scheduling transmission of data or control information for a given iteration of the AI / ML model training process. In such embodiments, a cyclic redundancy check (CRC) value of the first DCI may be scrambled with a first radio network temporary identifier (RNTI), and the second device may transmit control signaling to configure the first device to monitor the first DCI according to a first monitoring period.
[0055] In some embodiments, a method according to a second broad aspect of the present disclosure further includes transmitting a second DCI including second scheduling information for scheduling transmission of other data or control information. In such embodiments, a CRC value of the second DCI may be scrambled with a second RNTI different from the first RNTI, and the second device may transmit device-specific control signaling to configure the first device to monitor the second DCI according to a second monitoring period.
[0056] In some embodiments, the second monitoring period is configured separately from the first monitoring period.
[0057] In some embodiments, for each of at least one iteration of the training process of the AI / ML model, transmitting information that enables the first device to determine whether to participate in the iteration further includes transmitting control signaling to configure the third device to monitor the first DCI according to a third monitoring period that is different from the first monitoring period.
[0058] In some embodiments, the first RNTI is different from the Cell RNTI (C-RNTI).
[0059] In some embodiments, a method according to a second broad aspect of the present disclosure further includes receiving a local AI / ML model update from the first device for a given iteration of an AI / ML model training process, the local AI / ML model update from the first device including parameter updates for the AI / ML model based on training of the local AI / ML model at the first device, and the training of the local AI / ML model at the first device based on data or control information transmitted from the second device for the given iteration of the AI / ML model training process. In such embodiments, the local AI / ML model update from the first device may further include information indicating a value of an iteration ID associated with the given iteration for which the first device received the data or control information from the second device.
[0060] In some embodiments, a method according to a second broad aspect of this disclosure further includes receiving a request from the first device to participate in a training process of the AI / ML model. For example, the request from the first device may include information indicating a value of an iteration ID associated with a given iteration of the training process of the AI / ML model.
[0061] In some embodiments, a method according to a second broad aspect of the disclosure further includes transmitting control signaling for the first device to configure the first device to train a partial AI / ML model that includes a partial subset of parameters of the local AI / ML model at the first device.
[0062] In some embodiments, a method according to a second broad aspect of the present disclosure further includes receiving local AI / ML model updates from the first device, where the local AI / ML model updates from the first device include parameter updates of the AI / ML model for a partial subset of parameters based on training of the partial AI / ML model at the first device.
[0063] In some embodiments, for each of at least one iteration of the training process of the AI / ML model, transmitting information that enables the first device to determine whether the first device will participate in the iteration includes transmitting control signaling for the first device to configure the first device not to participate in the training process of the AI / ML model based on the training capability of the AI / ML model received from the first device.
[0064] In some embodiments, the method according to the second broad aspect of this disclosure further includes receiving a request from the first device that it does not want to participate in the training process of the AI / ML model.
[0065] Corresponding apparatus and devices for carrying out the method are disclosed.
[0066] For example, according to another aspect of the present disclosure, there is provided a device including a processor and a memory storing processor-executable instructions that, when executed, cause the processor to perform a method according to the first broad aspect of the present disclosure described above.
[0067] As another example, according to another aspect of the present disclosure, there is provided a device including a processor and a memory storing processor-executable instructions that, when executed, cause the processor to perform a method according to the second broad aspect of the present disclosure described above.
[0068] According to another aspect of the present disclosure, there is provided an apparatus including one or more units for implementing any of the method aspects disclosed in this disclosure. The term "unit" is used broadly and may be referred to by any of a variety of names including, for example, module, component, element, means, etc. A unit may be implemented using hardware, software, firmware, or any combination thereof.
[0069] Corresponding apparatus and devices for carrying out the method are disclosed.
[0070] For example, according to another aspect of the present disclosure, there is provided a device including a processor and a memory storing processor-executable instructions that, when executed, cause the processor to perform a method according to the first broad aspect of the present disclosure described above.
[0071] According to another aspect of the present disclosure, there is provided an apparatus including one or more units for implementing any of the method aspects disclosed in this disclosure. The term "unit" is used broadly and may be referred to by any of a variety of names including, for example, module, component, element, means, etc. A unit may be implemented using hardware, software, firmware, or any combination thereof. [Brief explanation of the drawings]
[0072] Reference will now be made, by way of example, to the accompanying drawings which illustrate exemplary embodiments of the present application.
[0073] [Figure 1] 1 is a simplified schematic diagram of a communication system, according to an example; [Figure 2] FIG. 1 illustrates another example of a communication system. [Figure 3] FIG. 1 illustrates an example of an electronic device (ED), a terrestrial transmission / reception point (T-TRP), and a non-terrestrial transmission / reception point (NT-TRP). [Figure 4] FIG. 1 illustrates an exemplary unit or module in a device. [Figure 5] FIG. 1 illustrates four EDs communicating with network devices in a communication system, according to one embodiment. [Figure 6A] FIG. 1 illustrates an example of a neural network with multiple layers of neurons, according to one embodiment. [Figure 6B] FIG. 1 illustrates an example of a neuron that may be used as a building block for a neural network, according to one embodiment. [Figure 7] FIG. 1 illustrates a timeline of actions performed by four EDs for one iteration of a synchronous associative learning procedure. [Figure 8] FIG. 1 illustrates a timeline of actions performed by four EDs across multiple iterations of an asynchronous associative learning procedure. [Figure 9] FIG. 1 illustrates a timeline of actions performed by four EDs over multiple iterations of a semi-synchronous federated learning procedure, according to one embodiment. [Figure 10] FIG. 1 illustrates an example of a flowchart for semi-synchronous federated learning, according to one embodiment. [Figure 11] FIG. 1 illustrates a timeline of actions performed by four EDs over multiple iterations of an asynchronous federated learning procedure, according to one embodiment. [Figure 12] FIG. 1 illustrates an example of a flowchart for asynchronous federated learning, according to one embodiment.
[0074] Similar reference numbers may be used in different figures to indicate similar components. DETAILED DESCRIPTION OF THE INVENTION
[0075] For purposes of explanation, certain exemplary embodiments will now be described in further detail below in conjunction with the figures. Exemplary Communication Systems and Devices
[0076] Referring to FIG. 1 , a simplified schematic diagram of a communication system is provided as an illustrative, non-limiting example. The communication system 100 comprises a radio access network 120. The radio access network 120 may be a next-generation (e.g., sixth-generation (6G) or later) radio access network or a legacy (e.g., 5G, 4G, 3G, or 2G) radio access network. One or more communication electrical devices (EDs) 110a-120j (commonly referred to as 110) may be interconnected to each other or to one or more network nodes (170a, 170b, commonly referred to as 170) in the radio access network 120. A core network 130 may be part of the communication system and may or may not be related to the radio access technology used in the communication system 100. The communication system 100 also comprises a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160.
[0077] FIG. 2 illustrates an exemplary communication system 100. Generally, the communication system 100 enables multiple wireless or wired elements to communicate data and other content. The purpose of the communication system 100 may be to provide content such as voice, data, video, and / or text via broadcast, multicast, unicast, and the like. The communication system 100 may operate by sharing resources, such as carrier spectrum bandwidth, among its components. The communication system 100 may include terrestrial and / or non-terrestrial communication systems. The communication system 100 may provide a wide range of communication services and applications (such as Earth observation, remote sensing, passive detection and positioning, navigation and tracking, autonomous delivery, and mobility). The communication system 100 may provide a high degree of availability and robustness through cooperation between the terrestrial and non-terrestrial communication systems. For example, integrating a non-terrestrial communication system (or components thereof) into a terrestrial communication system may result in what may be considered a heterogeneous network comprising multiple layers. Compared to traditional communication networks, heterogeneous networks can achieve better overall performance through efficient multi-link cooperation between terrestrial and non-terrestrial networks, more flexible function sharing, and faster physical layer link switching.
[0078] The terrestrial and non-terrestrial communication systems may be considered subsystems of a communication system. In the illustrated example, communication system 100 includes electronic devices (EDs) 110a-110d (commonly referred to as EDs 110), radio access networks (RANs) 120a-120b, a non-terrestrial communication network 120c, a core network 130, a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160. RANs 120a-120b include respective base stations (BSs) 170a-170b, which may be commonly referred to as terrestrial transmission / reception points (T-TRPs) 170a-170b. Non-terrestrial communication network 120c includes access nodes 120c, which may be commonly referred to as non-terrestrial transmission / reception points (NT-TRPs) 172.
[0079] Any ED 110 may alternatively or additionally be configured to interface with, access, or communicate with any other T-TRPs 170a-170b and NT-TRPs 172, the Internet 150, the core network 130, the PSTN 140, other networks 160, or any combination of the above. In some examples, the ED 110a may communicate uplink and / or downlink transmissions via an interface 190a with the T-TRP 170a. In some examples, the EDs 110a, 110b, and 110d may also communicate directly with each other via one or more sidelink air interfaces 190b. In some examples, the ED 110d may communicate uplink and / or downlink transmissions via an interface 190c with the NT-TRP 172.
[0080] Air interfaces 190a and 190b may use similar communication technologies, such as any suitable radio access technology. For example, communication system 100 may implement one or more channel access methods in air interfaces 190a and 190b, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), or single-carrier FDMA (SC-FDMA). Air interfaces 190a and 190b may utilize other higher-dimensional signal spaces, which may involve combinations of orthogonal and / or non-orthogonal dimensions.
[0081] The air interface 190c may enable communication between the EDs 110d and one or more NT-TRPs 172 via a wireless link or simply a link. In some examples, the link is a dedicated connection for unicast transmission, a connection for broadcast transmission, or a connection between a group of EDs and one or more NT-TRPs for multicast transmission.
[0082] The RANs 120a and 120b are in communication with the core network 130 to provide various services, such as voice, data, and other services, to the EDs 110a, 110b, and 110c. The RANs 120a and 120b and / or the core network 130 may be in direct or indirect communication with one or more other RANs (not shown), which may or may not be served directly by the core network 130 and which may or may not employ the same radio access technology as the RAN 120a, RAN 120b, or both. The core network 130 may also serve as a gateway access between (i) the RANs 120a and 120b or the EDs 110a, 110b, and 110c, or both, and (ii) other networks (such as the PSTN 140, the Internet 150, and other networks 160). Additionally, some or all of the EDs 110a, 110b, and 110c may include functionality for communicating with different wireless networks over different wireless links using different wireless technologies and / or protocols. Instead of (or in addition to) wireless communication, the EDs 110a, 110b, and 110c may communicate via wired communication channels to a service provider or switch (not shown) and to the Internet 150. The PSTN 140 may include a circuit-switched telephone network for providing plain old telephone service (POTS). The Internet 150 may include computer networks and subnets (intranets), or both, and may incorporate protocols such as Internet Protocol (IP), Transmission Control Protocol (TCP), and User Datagram Protocol (UDP). The EDs 110a, 110b, and 110c may be multimode devices enabling operation with multiple radio access technologies and may incorporate multiple transceivers necessary to support such.
[0083] 3 shows another example of the ED 110 and the base stations 170a, 170b, and / or 170c. The ED 110 is used to connect people, objects, machines, etc. The ED 110 can be widely used in various scenarios, such as cellular communication, device-to-device (D2D), vehicle-to-everything (V2X), peer-to-peer (P2P), machine-to-machine (M2M), machine-type communication (MTC), Internet of Things (IOT), virtual reality (VR), augmented reality (AR), industrial control, autonomous driving, remote medical care, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drone, robot, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery, and mobility.
[0084] Each ED 110 represents any suitable end-user device for wireless operation and may include (or be referred to as) a device such as a user equipment / device (UE), a wireless transmit / receive unit (WTRU), a mobile station, a fixed or mobile subscriber unit, a cellular telephone, a station (STA), a machine-type communication (MTC) device, a personal digital assistant (PDA), a smartphone, a laptop, a computer, a tablet, a wireless sensor, a consumer electronics device, a smart book, a vehicle, a car, a truck, a bus, a train, or an IoT device, an industrial device, or an apparatus (e.g., a communication module, a modem, or a chip) among the aforementioned devices, among other possibilities. Future generation EDs 110 may be referred to using other terminology. Base stations 170a and 170b are T-TRPs and are hereinafter referred to as T-TRP 170. Also shown in FIG. 3, an NT-TRP is hereinafter referred to as NT-TRP 172. Each ED110 connected to a T-TRP170 and / or NT-TRP172 may be dynamically or semi-statically turned on (i.e., established, activated, or enabled), turned off (i.e., released, deactivated, or disabled), and / or configured in response to one or more of connection availability and connection need.
[0085] The ED 110 includes a transmitter 201 and a receiver 203 coupled to one or more antennas 204. Only one antenna 204 is shown. One, some, or all of the antennas may alternatively be panels. The transmitter 201 and receiver 203 may be integrated, for example, as a transceiver. The transceiver is configured to modulate data or other content for transmission by at least one antenna 204 or a network interface controller (NIC). The transceiver is also configured to demodulate data or other content received by at least one antenna 204. Each transceiver includes any suitable structure for generating signals for wireless or wired transmission and / or for processing signals received wirelessly or by wire. Each antenna 204 includes any suitable structure for transmitting and / or receiving wireless or wired signals.
[0086] The ED 110 includes at least one memory 208. The memory 208 stores instructions and data used, generated, or collected by the ED 110. For example, the memory 208 may store software instructions or modules configured to implement some or all of the functions and / or embodiments described herein and performed by the processing unit 210. Each memory 208 includes any suitable volatile and / or non-volatile storage and retrieval device. Any suitable type of memory may be used, such as random access memory (RAM), read-only memory (ROM), hard disk, optical disk, subscriber identity module (SIM) card, memory stick, secure digital (SD) memory card, on-processor cache, etc.
[0087] ED 110 may further include one or more input / output devices (not shown) or interfaces (such as a wired interface to Internet 150 in FIG. 1). The input / output devices enable interaction with a user or other devices in a network. Each input / output device includes any suitable structure for providing information to or receiving information from a user, such as a speaker, microphone, keypad, keyboard, display, or touch screen, including network interface communications.
[0088] The ED 110 further includes a processor 210 for performing operations including operations related to preparing a transmission for uplink transmission to the NT-TRP 172 and / or the T-TRP 170, operations related to processing a downlink transmission received from the NT-TRP 172 and / or the T-TRP 170, and operations related to processing a sidelink transmission to or from another ED 110. The processing operations related to preparing a transmission for uplink transmission may include operations such as encoding, modulating, transmit beamforming, and generating symbols for transmission. The processing operations related to processing a downlink transmission may include operations such as receive beamforming, demodulating, and decoding received symbols. Depending on the embodiment, downlink transmissions may in some cases be received by the receiver 203 using receive beamforming, and the processor 210 may extract the signaling from the downlink transmission (e.g., by detecting and / or decoding the signaling). An example of signaling may be a reference signal transmitted by the NT-TRP 172 and / or the T-TRP 170. In some embodiments, the processor 276 implements transmit beamforming and / or receive beamforming based on an indication of beam direction, e.g., beam angle information (BAI), received from the T-TRP 170. In some embodiments, the processor 210 may perform operations related to network access (e.g., initial access) and / or downlink synchronization, such as operations related to detecting synchronization sequences, decoding and obtaining system information, etc. In some embodiments, the processor 210 may perform channel estimation using, for example, reference signals received from the NT-TRP 172 and / or the T-TRP 170.
[0089] Although not shown, the processor 210 may form part of the transmitter 201 and / or the receiver 203. Although not shown, the memory 208 may form part of the processor 210.
[0090] The processor 210 and the processing components of the transmitter 201 and receiver 203 may each be implemented by the same or different one or more processors configured to execute instructions stored in a memory (e.g., in memory 208). Alternatively, some or all of the processor 210 and the processing components of the transmitter 201 and receiver 203 may be implemented using special purpose circuitry, such as a programmed field programmable gate array (FPGA), a graphics processing unit (GPU), or an application specific integrated circuit (ASIC).
[0091] In some implementations, the T-TRP 170 may be known by other names such as a base station, base transceiver station (BTS), radio base station, network node, network device, network-side device, transmitting / receiving node, Node B, evolved Node B (eNodeB or eNB), Home eNodeB, next-generation Node B (gNB), transmission point (TP), site controller, access point (AP), or wireless router, relay station, remote radio head, terrestrial node, terrestrial network device or terrestrial base station, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), positioning node, etc. The T-TRP 170 may be a macro BS, pico BS, relay node, donor node, etc., or a combination thereof. The T-TRP 170 may refer to any of the aforementioned devices or an apparatus in any of the aforementioned devices (e.g., a communication module, a modem, or a chip).
[0092] In some embodiments, parts of the T-TRP 170 may be distributed. For example, some of the modules of the T-TRP 170 may be located remotely from the equipment housing the antenna of the T-TRP 170 and may be coupled to the equipment housing the antenna via a communications link (not shown), sometimes known as a fronthaul, such as a Common Public Radio Interface (CPRI). Thus, in some embodiments, the term T-TRP 170 may also refer to network-side modules that perform processing operations such as determining the location of the ED 110, resource allocation (scheduling), message generation, and encoding / decoding, and that are not necessarily part of the equipment housing the antenna of the T-TRP 170. Modules may also be coupled to other T-TRPs. In some embodiments, the T-TRP 170 may actually be multiple T-TRPs operating together to serve the ED 110, for example, through coordinated multipoint transmission.
[0093] The T-TRP 170 includes at least one transmitter 252 and at least one receiver 254 coupled to one or more antennas 256. Only one antenna 256 is shown. One, some, or all of the antennas may alternatively be panels. The transmitter 252 and receiver 254 may be integrated as a transceiver. The T-TRP 170 further includes a processor 260 for performing operations including operations related to preparing a transmission for downlink transmission to the ED 110, processing uplink transmissions received from the ED 110, preparing a transmission for backhaul transmission to the NT-TRP 172, and processing transmissions received via the backhaul from the NT-TRP 172. Processing operations related to preparing a transmission for downlink or backhaul transmission may include operations such as encoding, modulating, precoding (e.g., MIMO precoding), transmit beamforming, and generating symbols for transmission. Processing operations related to processing transmissions received in the uplink or over the backhaul may include operations such as receive beamforming, demodulating, and decoding received symbols. The processor 260 may also perform operations related to network access (e.g., initial access) and / or downlink synchronization, such as generating synchronization signal block (SSB) content and generating system information. In some embodiments, the processor 260 also generates an indication of a beam direction, e.g., a BAI, that may be scheduled for transmission by the scheduler 253. The processor 260 performs other network-side processing operations described herein, such as determining the location of the ED 110 and determining where to deploy the NT-TRP 172. In some embodiments, the processor 260 may generate signaling to configure, for example, one or more parameters of the ED 110 and / or one or more parameters of the NT-TRP 172. Any signaling generated by the processor 260 is sent by the transmitter 252.Note that "signaling" as used herein may alternatively be referred to as control signaling. Dynamic signaling may be transmitted in a control channel, e.g., a Physical Downlink Control Channel (PDCCH), and static or semi-static upper layer signaling may be included in packets transmitted in a data channel, e.g., a Physical Downlink Shared Channel (PDSCH).
[0094] The scheduler 253 may be coupled to the processor 260. The scheduler 253 may be included within the T-TRP 170 or operated separately therefrom, which may schedule uplink, downlink, and / or backhaul transmissions, including issuing scheduling grants and / or configuring resources without scheduling (“grant configured”). The T-TRP 170 further includes a memory 258 for storing information and data. The memory 258 stores instructions and data used, generated, or collected by the T-TRP 170. For example, the memory 258 may store software instructions or modules configured to implement some or all of the functions and / or embodiments described herein and performed by the processor 260.
[0095] Although not shown, the processor 260 may form part of the transmitter 252 and / or the receiver 254. Also, although not shown, the processor 260 may implement the scheduler 253. Although not shown, the memory 258 may form part of the processor 260.
[0096] The processor 260, the scheduler 253, and the processing components of the transmitter 252 and the receiver 254 may each be implemented by the same or different one or more processors configured to execute instructions stored in a memory, for example, in the memory 258. Alternatively, some or all of the processor 260, the scheduler 253, and the processing components of the transmitter 252 and the receiver 254 may be implemented using dedicated circuitry, such as an FPGA, a GPU, or an ASIC.
[0097] Although the NT-TRP 172 is illustrated as a drone merely as an example, the NT-TRP 172 may be implemented in any suitable non-terrestrial form. The NT-TRP 172 may also be known by other names, such as a non-terrestrial node, a non-terrestrial network device, or a non-terrestrial base station, in some implementations. The NT-TRP 172 includes a transmitter 272 and a receiver 274 coupled to one or more antennas 280. Only one antenna 280 is shown. One, some, or all of the antennas may alternatively be panels. The transmitter 272 and the receiver 274 may be integrated as a transceiver. The NT-TRP 172 further includes a processor 276 for performing operations, including operations related to preparing a transmission for downlink transmission to the ED 110, processing an uplink transmission received from the ED 110, preparing a transmission for backhaul transmission to the T-TRP 170, and processing a transmission received via the backhaul from the T-TRP 170. Processing operations related to preparing a transmission for downlink or backhaul transmission may include operations such as encoding, modulating, precoding (e.g., MIMO precoding), transmit beamforming, and generating symbols for transmission. Processing operations related to processing a transmission received in the uplink or over the backhaul may include operations such as receive beamforming, demodulating, and decoding received symbols. In some embodiments, the processor 276 implements transmit beamforming and / or receive beamforming based on beam direction information (e.g., BAI) received from the T-TRP 170. In some embodiments, the processor 276 may generate signaling, for example, to configure one or more parameters of the ED 110. In some embodiments, the NT-TRP 172 implements physical layer processing but does not implement higher layer functions, such as functions at the medium access control (MAC) or radio link control (RLC) layer. As this is just one example, more generally, the NT-TRP 172 may implement higher layer functions in addition to physical layer processing.
[0098] The NT-TRP 172 further includes a memory 278 for storing information and data. Although not shown, the processor 276 may form part of the transmitter 272 and / or the receiver 274. Although not shown, the memory 278 may form part of the processor 276.
[0099] The processor 276 and the processing components of the transmitter 272 and receiver 274 may each be implemented by the same or different processor(s) configured to execute instructions stored in memory, e.g., in memory 278. Alternatively, some or all of the processor 276 and the processing components of the transmitter 272 and receiver 274 may be implemented using dedicated circuitry, such as a programmed FPGA, GPU, or ASIC. In some embodiments, the NT-TRP 172 may actually be multiple NT-TRPs operating together to service the ED 110, e.g., through coordinated multipoint transmission.
[0100] It should be noted that "TRP" as used herein can refer to T-TRP or NT-TRP.
[0101] T-TRP170, NT-TRP172, and / or ED110 may contain other components, which have been omitted for clarity.
[0102] One or more steps of the embodiment methods provided herein may be implemented by corresponding units or modules according to FIG. 4. FIG. 4 illustrates units or modules in a device, such as the ED 110, the T-TRP 170, or the NT-TRP 172. For example, a signal may be transmitted by a transmitting unit or a transmitting module. For example, a signal may be transmitted by a transmitting unit or a transmitting module. A signal may be received by a receiving unit or a receiving module. A signal may be processed by a processing unit or a processing module. Other steps may be performed by an artificial intelligence (AI) or machine learning (ML) module. Each unit or module may be implemented using hardware, one or more components or devices executing software, or a combination thereof. For example, one or more of the units or modules may be an integrated circuit, such as a programmed FPGA, a GPU, or an ASIC. It will be appreciated that where modules are implemented using software, for example, for execution by a processor, they may be retrieved by the processor in single or multiple instances, individually or together for processing, in whole or in part, as needed, and that the modules themselves may include instructions for further deployment and instantiation.
[0103] Further details regarding ED110, T-TRP170, and NT-TRP172 are known to those skilled in the art, and therefore these details are omitted here.
[0104] Control signaling is described herein in some embodiments. Control signaling may sometimes be alternatively referred to as signaling, or control information, or setting information, or configuration. In some cases, control signaling may be dynamically indicated in the physical layer, for example, in a control channel. An example of dynamically indicated control signaling is information sent in physical layer control signaling, for example, downlink control information (DCI). Control signaling may sometimes alternatively be semi-statically indicated, for example, in RRC signaling or in a MAC control element (CE). A dynamic indication may be an indication in a lower layer, for example, physical layer / Layer 1 signaling (e.g., in a DCI), rather than in a higher layer (e.g., rather than in RRC signaling or in a MAC CE). A semi-static indication may be an indication in semi-static signaling. As used herein, semi-static signaling may refer to signaling that is not dynamic, e.g., higher layer signaling, RRC signaling, and / or MAC CE. As used herein, dynamic signaling may refer to signaling that is dynamic, e.g., physical layer control signaling sent in the physical layer, such as DCI.
[0105] An air interface generally includes several components and associated parameters that collectively specify how transmissions are sent and / or received over a wireless communication link between two or more communication devices. For example, an air interface may include one or more components that define a waveform, frame structure, multiple access scheme, protocol, coding scheme, and / or modulation scheme for carrying information (e.g., data) over the wireless communication link. The wireless communication link may support a link between a radio access network and user equipment (e.g., a “Uu” link), and / or the wireless communication link may support a device-to-device link (e.g., a “sidelink”), such as between two user equipments, and / or the wireless communication link may support a link between a non-terrestrial (NT) communication network and user equipment (UE). The following are some examples of the above components: The waveform component may specify the shape and form of the signal being transmitted. Waveform options may include orthogonal multiple access waveforms and non-orthogonal multiple access waveforms. Non-limiting examples of such waveform options include Orthogonal Frequency Division Multiplexing (OFDM), Filtered OFDM (f-OFDM), Time Windowing OFDM, Filter Bank Multicarrier (FBMC), Universal Filter Multicarrier (UFMC), Generalized Frequency Division Multiplexing (GFDM), Wavelet Packet Modulation (WPM), Faster Than Nyquist (FTN) waveforms, and Low Peak-to-Average Power Ratio (Low PAPR) waveforms. The frame structure component may specify the configuration of a frame or group of frames. The frame structure component may indicate one or more of the time, frequency, pilot signature, code, or other parameters of the frame or group of frames. Further details of the frame structure are described below. The multiple access scheme component may specify multiple access technology options, including technologies that define how communication devices share a common physical channel, such as time division multiple access (TDMA), frequency division multiple access (FDMA), code division multiple access (CDMA), single-carrier frequency division multiple access (SC-FDMA), low-density signature multi-carrier code division multiple access (LDS-MC-CDMA), non-orthogonal multiple access (NOMA), pattern division multiple access (PDMA), lattice division multiple access (LPMA), resource spreading multiple access (RSMA), and sparse code multiple access (SCMA). Additionally, multiple access technology options may include scheduled access versus unscheduled access, also known as permissionless access, e.g., non-orthogonal multiple access versus orthogonal multiple access (no sharing between multiple communication devices) via dedicated channel resources, contention-based shared channel resources versus non-contention-based shared channel resources, and cognitive radio-based access. Hybrid Automatic Repeat Request (HARQ) protocol components may specify how transmissions and / or retransmissions are performed. Non-limiting examples of transmission and / or retransmission mechanism options include those specifying scheduled data pipe sizes, signaling mechanisms for transmissions and / or retransmissions, and retransmission mechanisms. The coding and modulation components may specify how the information being transmitted can be coded / decoded and modulated / demodulated for transmission / reception purposes. Coding may refer to methods of error detection and forward error correction. Non-limiting examples of coding options include turbo trellis codes, turbo product codes, fountain codes, low-density parity-check codes, and polar codes. Modulation may simply refer to the constellation (e.g., including the modulation technique and order), or more specifically, to various types of advanced modulation methods such as hierarchical modulation and low-PAPR modulation.
[0106] In some embodiments, the air interface may be a “one-size-fits-all concept.” For example, once the air interface is defined, components within the air interface may not be changed or adapted. In some implementations, only limited parameters or modes of the air interface, such as the cyclic prefix (CP) length or multiple-input multiple-output (MIMO) mode, may be configured. In some embodiments, the air interface design may provide a unified or flexible framework to support frequency (e.g., mmWave) bands below and above 6 GHz for both licensed and unlicensed access. As an example, the flexibility of the configurable air interface afforded by scalable numerology and symbol duration may enable optimization of transmission parameters for different spectrum bands and for different services / devices. As another example, the unified air interface may be self-contained in the frequency domain, and a self-contained design in the frequency domain may support more flexible radio access network (RAN) slicing through the sharing of channel resources between different services in both frequency and time.
[0107] Frame structure The frame structure is a feature of the physical layer of wireless communications that defines a time-domain signal transmission structure to enable, for example, timing reference and timing alignment of basic time-domain transmission units. Wireless communications between communicating devices may occur over time-frequency resources governed by the frame structure. The frame structure is sometimes alternatively referred to as a radio frame structure.
[0108] Depending on the frame structure and / or configuration of frames in the frame structure, frequency division duplex (FDD) and / or time division duplex (TDD) and / or full duplex (FD) communication may be possible. FDD communication is when transmissions in different directions (e.g., uplink vs. downlink) occur on different frequency bands. TDD communication is when transmissions in different directions (e.g., uplink vs. downlink) occur at different time periods. FD communication is when transmission and reception occur on the same time-frequency resource, i.e., a device can both transmit and receive on the same frequency resource simultaneously in time.
[0109] An example of a frame structure is that in Long Term Evolution (LTE), where each frame is 10 ms in duration, each frame has 10 subframes, each of 1 ms in duration, each subframe includes two slots, each of 0.5 ms in duration, each slot is for the transmission of seven OFDM symbols (assuming a normal CP), each OFDM symbol has a symbol period and a specific bandwidth (or partial bandwidth or bandwidth partition) in terms of the number of subcarriers and the spacing of the subcarriers, the frame structure is based on OFDM waveform parameters such as subcarrier spacing and CP length (where the CP has fixed or limited length options), and the switching gap between uplink and downlink in TDD must be an integer number of OFDM symbol periods.
[0110] Another example of a frame structure is the frame structure in New Radio (NR), which supports multiple subcarrier spacings, each corresponding to a different numerology. The frame structure depends on the numerology, but in all cases, the frame length is set to 10 ms, consisting of 10 subframes of 1 ms each. A slot is defined as 14 OFDM symbols, and the slot length depends on the numerology. For example, the NR frame structure for 15 kHz subcarrier spacing with normal CP ("numerology 1") differs from the NR frame structure for 30 kHz subcarrier spacing with normal CP ("numerology 2"). For 15 kHz subcarrier spacing, the slot length is 1 ms, and for 30 kHz subcarrier spacing, the slot length is 0.5 ms. The NR frame structure may have more flexibility than the LTE frame structure.
[0111] Another example of a frame structure is an exemplary flexible frame structure, e.g., for use in 6G networks and beyond. In a flexible frame structure, a symbol block may be defined as the smallest period that can be scheduled during the flexible frame structure. A symbol block may be a unit of transmission having an optional redundant portion (e.g., a CP portion) and an information (e.g., data) portion. An OFDM symbol is an example of a symbol block. A symbol block may alternatively be referred to as a symbol. Embodiments of a flexible frame structure include different parameters that may be configurable, such as, for example, a frame length, a subframe length, a symbol block length, etc. A non-exhaustive list of configurable parameters possible in some embodiments of a flexible frame structure includes the following: (1) Frame: The frame length does not need to be limited to 10 ms and may be configurable and changed over time. In some embodiments, each frame includes one or more downlink synchronization channels and / or one or more downlink broadcast channels, and each synchronization channel and / or broadcast channel may be transmitted in a different direction by different beamforming. The frame length is two or more possible values and may be configured based on the application scenario. For example, an autonomous vehicle may require relatively fast initial access, in which case the frame length may be set as 5 ms for an autonomous vehicle application. As another example, a residential smart meter may not require fast initial access, in which case the frame length may be set as 20 ms for a smart meter application. (2) Subframe Duration: Subframes may or may not be defined in a flexible frame structure depending on the implementation. For example, a frame may be defined to include slots but not subframes. In a frame where subframes are defined, for example, for time domain alignment, then the duration of the subframe may be configurable. For example, a subframe may be configured to have a length of 0.1 ms, 0.2 ms, 0.5 ms, 1 ms, 2 ms, 5 ms, etc. In some embodiments, if a subframe is not required in a particular scenario, the subframe length may be defined to be the same as the frame length or may not be defined. (3) Slot Configuration: Slots may or may not be defined in a flexible frame structure depending on the implementation. In a frame in which slots are defined, then the definition of the slots (e.g., over a period of time and / or within some symbol blocks) may be configurable. In one embodiment, the slot configuration is common to all UEs or a group of UEs. In this case, slot configuration information may be transmitted to the UE in a broadcast channel or a common control channel. In other embodiments, the slot configuration may be UE-specific, in which case the slot configuration information may be transmitted in a UE-specific control channel. In some embodiments, slot configuration signaling may be transmitted together with frame configuration signaling and / or subframe configuration signaling. In other embodiments, the slot configuration may be transmitted independently of frame configuration signaling and / or subframe configuration signaling. In general, the slot configuration may be common to the system, common to a base station, common to a group of UEs, or specific to a UE. (4) Subcarrier Spacing (SCS): SCS is one parameter of a scalable numerology that may allow the SCS to range from 15 KHz to 480 KHz in some cases. The SCS may vary with the frequency of the spectrum and / or the maximum UE speed to minimize the effects of Doppler shift and phase noise. In some examples, there may be separate transmit and receive frames, and the SCS of the symbols in the receive frame structure may be configured independently from the SCS of the symbols in the transmit frame structure. The SCS in the receive frames may differ from the SCS in the transmit frames. In some examples, the SCS of each transmit frame may be half the SCS of each receive frame. If the SCS between the receive and transmit frames differs, for example, if a more flexible symbol period is implemented using an inverse discrete Fourier transform (IDFT) instead of a fast Fourier transform (FFT), the difference does not necessarily need to be scaled by a factor of two. Additional example frame structures may be used with different SCSs. (5) Flexible Transmission Duration of Basic Transmission Unit: The basic transmission unit may be a symbol block (alternatively called a symbol), which generally includes a redundant portion (called a CP) and an information (e.g., data) portion, although in some embodiments, the CP may be omitted from the symbol block. The CP length may be flexible and configurable. The CP length may be fixed within a frame or flexible within a frame, and the CP length may possibly be changed per frame, per group of frames, per subframe, per slot, or dynamically per scheduling. The information (e.g., data) portion may be flexible and configurable. Another possible parameter for a symbol block that may be defined is the ratio of the CP duration to the information (e.g., data) duration. In some embodiments, the symbol block length may be adjusted according to channel conditions (e.g., multipath delay, Doppler), and / or latency requirements, and / or available duration. As another example, the symbol block length may be adjusted to fit the available duration during a frame. (6) Flexible switching gap: A frame may include both a downlink portion for downlink transmission from the base station and an uplink portion for uplink transmission from the UE. A gap may exist between each uplink and downlink portion, which is called a switching gap. The switching gap length (duration) may be configurable. The switching gap duration may be fixed within a frame or flexible within a frame, and the switching gap duration may possibly be changed per frame, per group of frames, per subframe, per slot, or dynamically per scheduling.
[0112] Cell / Carrier / Bandwidth Portion (BWP) / Occupied Bandwidth A device such as a base station may provide coverage beyond a cell. Wireless communication with a device may occur over one or more carrier frequencies. A carrier frequency is referred to as a carrier. A carrier may alternatively be referred to as a component carrier (CC). A carrier may be characterized by its bandwidth and a reference frequency, e.g., the center or lowest or highest frequency of the carrier. A carrier may be on a licensed or unlicensed spectrum. Wireless communication with a device may also or alternatively occur over one or more bandwidth portions (BWPs). For example, a carrier may have one or more BWPs. More generally, wireless communication with a device may occur over a spectrum. A spectrum may comprise one or more carriers and / or one or more BWPs.
[0113] A cell may include one or more downlink resources and, optionally, one or more uplink resources, or a cell may include one or more uplink resources and, optionally, one or more downlink resources, or a cell may include both one or more downlink resources and one or more uplink resources. As an example, a cell may include only one downlink carrier / BWP, or only one uplink carrier / BWP, or multiple downlink carriers / BWPs, or multiple uplink carriers / BWPs, or one downlink carrier / BWP and one uplink carrier / BWP, or one downlink carrier / BWP and multiple uplink carriers / BWPs, or multiple downlink carriers / BWPs and one uplink carrier / BWP, or multiple downlink carriers / BWPs and multiple uplink carriers / BWPs. In some embodiments, a cell may alternatively or additionally include one or more sidelink resources, including sidelink transmission and reception resources.
[0114] A BWP is a set of contiguous or non-contiguous frequency subcarriers on a carrier, or a set of contiguous or non-contiguous frequency subcarriers on multiple carriers, or a set of non-contiguous or contiguous frequency subcarriers that may have one or more carriers.
[0115] In some embodiments, a carrier may have one or more BWPs; for example, a carrier may have a 20 MHz bandwidth and consist of one BWP, or a carrier may have an 80 MHz bandwidth and consist of two adjacent, contiguous BWPs, etc. In other embodiments, a BWP may have one or more carriers; for example, a BWP may have a 40 MHz bandwidth and consist of two adjacent, contiguous carriers, where each carrier has a 20 MHz bandwidth. In some embodiments, a BWP may comprise discontinuous spectral resources consisting of discontinuous carriers, where a first carrier of the discontinuous carriers may be in the mmW band, a second carrier may be in the low band (such as the 2 GHz band), a third carrier (if present) may be in the THz band, and a fourth carrier (if present) may be in the visible light band. The resources in one carrier belonging to a BWP may be contiguous or discontinuous. In some embodiments, a BWP has discontinuous spectral resources on one carrier.
[0116] Wireless communication may occur over an occupied bandwidth, which may be defined as the width of a frequency band below a lower frequency limit and above an upper frequency limit, each of which emits an average power equal to a specified fraction □ / 2 of the total average transmitted power, where, for example, the value of □ / 2 is taken as 0.5%.
[0117] The carrier, BWP, or occupied bandwidth may be signaled by a network device (e.g., a base station) dynamically, e.g., in physical layer control signaling such as Downlink Control Information (DCI), or semi-statically, e.g., in Radio Resource Control (RRC) signaling or at the Medium Access Control (MAC) layer, or may be pre-defined based on an application scenario or determined by the UE depending on other parameters known by the UE, or may be fixed, e.g., by a standard.
[0118] Artificial Intelligence (AI) and / or Machine Learning (ML) The number of new devices in future wireless networks is expected to increase exponentially, and device capabilities are expected to become increasingly diverse. Many new applications and use cases are also expected to emerge with more diverse quality of service requirements than those of 5G applications / use cases. These will become new key performance indicators (KPIs) for future wireless networks (e.g., 6G networks), which may be extremely challenging. AI techniques, such as ML techniques (e.g., deep learning), have been introduced into telecommunications applications with the goal of improving system performance and efficiency.
[0119] Additionally, advances in antenna and bandwidth capabilities continue to be made, potentially enabling more and / or better communications over wireless links. Furthermore, advances continue in the areas of computer architecture and computing power, for example, with the introduction of general-purpose graphics processing units (GP-GPUs). Future generations of communication devices may have more computing and / or communication capabilities than previous generations, which may enable the adoption of AI to implement air interface components. Future generations of networks may also have access to more accurate and / or new information (compared to previous networks), such as the physical speed / velocity at which the device is traveling, the device's link budget, the device's channel conditions, one or more device capabilities and / or service types to be supported, sensing information, and / or positioning information, which may form the basis of input to AI models. To obtain sensing information, the TRP may transmit a signal to a target object (e.g., a suspect UE), and based on the signal's reflection, the TRP or another network device calculates the angle (for beamforming for the device), the device's distance from the TRP, and / or Doppler shift information. Positioning information, sometimes referred to as location determination, may be obtained in various ways, such as positioning reports from the UE (such as reporting the UE's GPS coordinates), using positioning reference signals (PRS), using the sensing described above, tracking and / or predicting the device's location, etc.
[0120] AI techniques (including ML techniques) may be applied to communications, including AI-based communications in the physical layer and / or MAC layer. In the case of the physical layer, AI communications may aim to optimize component designs and / or improve algorithm performance. For example, AI may be applied with respect to implementing channel coding, channel modeling, channel estimation, channel decoding, modulation, demodulation, MIMO, waveforms, multiple access, physical layer element parameter optimization and updates, beamforming, tracking, sensing, and / or positioning, etc. In the case of the MAC layer, AI communications may aim to utilize AI capabilities for learning, prediction, and / or decision-making to solve complex optimization problems with possible better strategies and / or optimal solutions for optimizing functions in the MAC layer. For example, AI may be applied to implement intelligent TRP management, intelligent beam management, intelligent channel resource allocation, intelligent power control, intelligent spectrum utilization, intelligent MCS, intelligent HARQ strategies, and / or intelligent transmit / receive mode adaptation, etc.
[0121] In some embodiments, the AI architecture may involve multiple nodes, which may be organized into one of two modes: centralized mode and distributed mode, both of which may be deployed in access networks, core networks, edge computing systems, or third-party networks. Centralized training and computing architectures are sometimes limited by high communication overhead and strict user data privacy. Distributed training and computing architectures may comprise several frameworks, e.g., distributed machine learning and federated learning. In some embodiments, the AI architecture may comprise an intelligent controller that can be implemented as a single agent or multiple agents based on joint or individual optimization. New protocols and signaling mechanisms are desired such that the corresponding interface links can be personalized with customized parameters to meet specific requirements while minimizing signaling overhead and maximizing system-wide spectral efficiency through personalized AI techniques.
[0122] In some embodiments herein, new protocols and signaling mechanisms are provided for operating within and switching between different operating modes for AI training and for measurements and feedback, including between training and normal operating modes to accommodate different possible measurements and information that may need to be fed back, depending on the implementation.
[0123] AI Training 1 and 2, embodiments of the present disclosure may be used to implement AI training involving two or more communication devices in communication system 100. For example, FIG. 5 illustrates four EDs communicating with network device 452 in communication system 100, according to one embodiment. Each of the four EDs is shown as a respective different UE, hereinafter referred to as UEs 402, 404, 406, and 408. However, an ED does not necessarily have to be a UE.
[0124] The network device 452 is part of a network (e.g., the radio access network 120). The network device 452 may be deployed in an access network, a core network, or an edge computing system or a third-party network, depending on the implementation. The network device 452 may be (or may be part of) a T-TRP or a server. In one example, the network device 452 may be (or may be implemented in) the T-TRP 170 or the NT-TRP 172. In another example, the network device 452 may be a T-TRP controller and / or an NT-TRP controller that can manage the T-TRP 170 or the NT-TRP 172. In some embodiments, the components of the network device 452 may be distributed. For example, if the network device 452 is part of a T-TRP that serves the UEs 402, 404, 406, and 408, the UEs 402, 404, 406, and 408 may communicate directly with the network device 452. Alternatively, the UEs 402, 404, 406, and 408 may communicate with the network device 352 via one or more intermediate components, such as, for example, via a T-TRP and / or an NT-TRP. For example, the network device 452 may send and / or receive information (e.g., control signaling, data, training sequences, etc.) to and from one or more of the UEs 402, 404, 406, and 408 via backhaul links and wireless channels interposed between the network device 452 and the UEs 402, 404, 406, and 408.
[0125] Each UE 402, 404, 406, and 408 includes a respective processor 210, memory 208, transmitter 201, receiver 203, and one or more antennas 204 (or alternatively, a panel), as described above. Only the processor 210, memory 208, transmitter 201, receiver 203, and antenna 204 for UE 402 are shown for simplicity, although the other UEs 404, 406, and 408 also include the same respective components.
[0126] For each UE 402, 404, 406, and 408, the communication link between that UE and a respective TRP in the network is an air interface. The air interface generally includes several components and associated parameters that collectively specify how transmissions are sent and / or received over the wireless medium.
[0127] The processor 210 of the UE in FIG. 5 implements one or more air interface components on the UE side. The air interface components configure and / or implement transmission and / or reception over the air interface. Examples of air interface components are described herein. The air interface components, such as a channel encoder (or decoder) that implements the coding components of the air interface for the UE, and / or a modulator (or demodulator) that implements the modulation components of the air interface for the UE, and / or a waveform generator that implements the waveform components of the air interface for the UE, may be in the physical layer. The air interface components, such as a module that implements channel estimation / tracking and / or a module that implements a retransmission protocol (e.g., that implements the HARQ protocol components of the air interface for the UE), may be in or part of a higher layer, such as the MAC layer. The processor 210 also directly implements (or controls the UE to implement) the UE-side operations described herein.
[0128] The network device 452 includes a processor 454, a memory 456, and an input / output device 458. The processor 454 implements or instructs other network devices (e.g., a T-TRP) to implement one or more of the air interface components on the network side. The air interface components may be implemented differently on the network side for one UE compared to another UE. The processor 454 directly implements (or controls network components to implement) the network-side operations described herein.
[0129] Processor 454 may be implemented by the same or different one or more processors configured to execute instructions stored in a memory (e.g., in memory 456). Alternatively, some or all of processor 454 may be implemented using dedicated circuitry such as a programmed FPGA, GPU, or ASIC. Memory 456 may be implemented by volatile and / or non-volatile storage. Any suitable type of memory may be used, such as RAM, ROM, hard disk, optical disk, on-processor cache, etc.
[0130] The input / output devices 458 enable interaction with other devices by receiving (input) and sending (output) information. In some embodiments, the input / output devices 458 may be implemented by a transmitter and / or receiver (or transceiver) and / or one or more interfaces (e.g., a wired interface to an internal network, the Internet, etc.). In some implementations, the input / output devices 458 may be implemented by a network interface, which may be implemented as a network interface card (NIC), and / or a computer port (e.g., a physical outlet to which a plug or cable connects), and / or a network socket, etc., depending on the implementation.
[0131] Network device 452 and UE 402 are capable of implementing one or more AI-enabled processes. In particular, in the embodiment of FIG. 5, network device 452 and UE 402 include ML modules 410 and 460, respectively. ML module 410 is implemented by processor 210 of UE 402, and ML module 460 is implemented by processor 454 of network device 452; thus, ML module 410 is shown within processor 210, and ML module 460 is shown with processor 454 in FIG. 5. ML modules 410 and 460 execute one or more AI / ML algorithms, for example, to perform one or more AI-enabled processes, e.g., AI-enabled link adaptation to optimize the communications link between the network and UE 402.
[0132] The ML modules 410 and 460 may be implemented using an AI model. The term AI model may refer to a computer algorithm configured to accept defined input data and output defined inference data, and the algorithm's parameters (e.g., weights) may be updated and optimized through training (e.g., using a training dataset or using real-world collected data). The AI model may use one or more neural networks (e.g., including deep neural networks (DNNs), recurrent neural networks (RNNs), convolutional neural networks (CNNs), and combinations thereof) and may be implemented using various neural network architectures (e.g., autoencoders, generative adversarial networks, etc.). Various techniques may be used to train an AI model to update and optimize its parameters. For example, backpropagation is a common technique for training DNNs, in which a loss function is calculated between the inference data generated by the DNN and some target output (e.g., ground truth data). The gradient of the loss function is computed with respect to the parameters of the DNN, and the computed gradient is used to update the parameters (e.g., using a gradient descent algorithm) with the goal of minimizing the loss function.
[0133] In some embodiments, the AI model encompasses a neural network used in machine learning. A neural network consists of multiple computational units (sometimes called neurons), which are organized into one or more layers. The process of receiving input at an input layer and generating output at an output layer is sometimes called forward propagation. In forward propagation, each layer receives input (which may have any suitable data format, such as a vector, matrix, or multidimensional array) and performs a computation to generate output (which may have different dimensions from the input). The computation performed by a layer generally involves applying (e.g., multiplying) the input by a set of weights (also called coefficients). With the exception of the first layer (i.e., input layer) of a neural network, the input to each layer is the output of the previous layer. A neural network may include one or more layers between the first layer (i.e., input layer) and the last layer (i.e., output layer), which may be called inner layers or hidden layers. For example, FIG. 6A shows an example of a neural network 600 including an input layer, an output layer, and two hidden layers. In this example, it can be seen that the outputs of each of the three neurons in the input layer of neural network 600 are included in the input vectors to each of the three neurons in the first hidden layer. Similarly, the outputs of each of the three neurons in the first hidden layer are included in the input vectors to each of the three neurons in the second hidden layer, and the outputs of each of the three neurons in the second hidden layer are included in the input vectors to each of the two neurons in the output layer. As noted above, the basic computational unit in a neural network is the neuron, as shown at 650 in FIG. 6A. FIG. 6B shows an example of a neuron 650 that can be used as a building block for neural network 600. As shown in FIG. 6B, in this example, neuron 650 takes a vector x as input and performs a dot product with an associated vector of weights w. The neuron's final output, z, is the result of the activation function f() on the dot product.Different neural networks can be designed with different architectures (eg, different numbers of layers with different functions performed by each layer).
[0134] Neural networks are trained to optimize the neural network's parameters (e.g., weights). This optimization is performed in an automated manner and is sometimes referred to as machine learning. Training a neural network involves forward propagating input data samples to generate output values (also called predicted or estimated output values) and comparing the generated output values with known or desired target values (e.g., ground truth values). A loss function is defined to quantitatively represent the difference between the generated output values and the target values, and the goal of training a neural network is to minimize the loss function. Backpropagation is an algorithm for training neural networks. Backpropagation is used to adjust (also called update) the values of parameters (e.g., weights) in a neural network so that the calculated loss function becomes smaller. Backpropagation involves calculating the gradient of the loss function with respect to the parameters to be optimized, and a gradient algorithm (e.g., gradient descent) is used to update the parameters to reduce the loss function. Backpropagation is performed iteratively, so that the loss function converges or is minimized over several iterations. After a training condition is satisfied (e.g., the loss function converges or a predefined number of training iterations are performed), the neural network is considered trained. The trained neural network may be deployed (or run) to generate output data inferred from input data. In some embodiments, training of the neural network may be ongoing even after the neural network is deployed, so that the parameters of the neural network may be iteratively updated with the latest training data.
[0135] Referring again to FIG. 5, in some embodiments, the UE 402 and the network device 452 may exchange information for training. The information exchanged between the UE 402 and the network device 452 is implementation-specific and may not have a human-understandable meaning (e.g., it may be intermediate data generated during the execution of an ML algorithm). Also or alternatively, the exchanged information may not be predefined by a standard; for example, bits may be exchanged, but the bits may not be associated with a predefined meaning. In some embodiments, the network device 452 may provide or indicate to the UE 402 one or more parameters used in the ML module 410 implemented in the UE 402. As an example, the network device 452 may send or indicate updated neural network weights implemented in the neural network executed by the ML module 410 on the UE side to attempt to optimize one or more aspects of the modulation and / or coding used for communications between the UE 402 and the T-TRP or NT-TRP.
[0136] In some embodiments, the UE 402 may implement the AI itself, e.g., perform learning, while in other embodiments, the UE 402 may not perform the learning itself, but may be able to operate in conjunction with an AI implementation on the network side, e.g., by receiving configurations for an AI model (such as a neural network or other ML algorithm) implemented by the ML module 410 from the network and / or by assisting other devices (such as network devices or other AI-enabled UEs) in training an AI model (such as a neural network or other ML algorithm) by providing requested measurements or observations. For example, in some embodiments, the UE 402 may not itself implement learning or training, but the UE 402 may receive trained configuration information for an ML model determined by the network device 452 and execute the model.
[0137] 5 assumes network-side AI / ML capabilities, there may be cases where the network itself does not perform the training / learning; instead, the UE may perform the learning / training itself, possibly using dedicated training signals sent from the network. In other embodiments, end-to-end (E2E) learning may be implemented by the UE and the network device 452.
[0138] For example, as described above, various processes, such as link adaptation, can be AI-enabled using AI by implementing AI models. Below are described some examples of possible AI / ML training processes and over-the-air information exchange procedures between devices during the training phase to facilitate AI-enabled processes according to embodiments of the present disclosure.
[0139] Referring again to FIG. 5, in the case of wireless federated learning (FL), the network device 452 may initialize a global AI / ML model implemented by the ML module 500, sample a group of UEs, such as the four UEs 402, 404, 406, and 408 shown in FIG. 5, and broadcast the parameters of the global AI / ML model to the UEs. Each of the UEs 402, 404, 406, and 408 may then initialize its local AI / ML model using the parameters of the global AI / ML model and update (train) its local AI / ML model using its own data. Each of the UEs 402, 404, 406, and 408 may then report the parameters of its updated local AI / ML model to the network device 452. The network device 452 may then aggregate the updated parameters reported from the UEs 402, 404, 406, and 408 and update the global AI / ML model. The above procedure is one iteration of the FL-based AI / ML model training procedure. The network device 452 and the UEs 402, 404, 406, and 408 perform multiple iterations until the AI / ML model has converged sufficiently to satisfy one or more training objectives / criteria and the AI / ML model is finalized.
[0140] There are two types of conventional FL processing: synchronous FL and asynchronous FL.
[0141] In conventional synchronous FL, which is an iterative training process, for each training iteration, a network device such as a BS updates a global model (e.g., aggregation and average) after receiving updates from all UEs participating in the synchronous FL training process. For example, FIG. 7 shows a timeline of acts 700 performed by four UEs (UE1, UE2, UE3, and UE4) for one iteration of the synchronous FL training process. In particular, FIG. 7 shows acts performed by the four UEs for the Nth iteration, where N≧1. In FIG. 7, the communication delay between the BS and each of the four UEs for the Nth iteration, including transmission delay, retransmission delay, signal processing delay, etc., is shown as 702 for DL communication between the BS and each UE. 1,N , 702 2,N , 702 3,N and 702 4,N 706 for UL communication between each UE and the BS. 1,N , 706 2,N , 706 3,N and 706 4,N Further, the AI / ML processing delay at each of the four UEs, including, for example, the delay for determining the update of the local AI / ML model for the Nth iteration, is 704 for each of the four UEs, respectively. 1,N , 704 2,N , 704 3,N and 704 4,N Further, the AI / ML processing delay at the BS side, including, for example, the delay for updating the global model with the update parameters of the local AI / ML model received from the UE, is shown at 708.
[0142] However, one significant problem with conventional synchronous FL training processes is the long training delay that can be caused by lagging UEs, e.g., UEs experiencing poor channel quality and / or insufficient computational power. This is because the BS does not begin the next iteration for training, e.g., the (N+1)th iteration shown at 710 in FIG. 7, until all UEs have successfully decoded the DL transmission, updated their local models, and reported the local model parameters to the BS. That is, the training delay is dominated by the worst-case UE with the longest communication and computation delays, resulting in a long delay for training the AI model. An example of this is shown in FIG. 7, where UE 4 is in the process of updating its local model parameters at 702. 4,N and 706 4,N 704 has significantly longer DL and UL communication delays as shown in 4,N , which delays the (N+1)th iteration relative to the time that UE1, UE2 and UE3 successfully report their local model parameters to the BS.
[0143] In asynchronous FL, the BS immediately updates the global AI / ML model whenever it receives an update from a UE. For example, FIG. 8 shows a timeline of acts 800 performed by four UEs (UE1, UE2, UE3, and UE4) over multiple iterations of the training process for asynchronous FL. In FIG. 8, the communication delay between the BS and each of the four UEs for the first DL transmission of the global AI / ML model is 802, respectively. 1,1 , 802 2,1 , 802 3,1 and 802 4,1 and the computation delay for the first update of those local models is 804 1,1 , 804 2,1 , 804 3,1 and 804 4,1 , and the communication delay between each of the UEs and the BS for the first UL transmission to report their first local AI / ML model parameter update is 806 1,1 , 806 2,1 , 8063,1 and 806 4,1 is shown in.
[0144] In this example, UE3 has the shortest combined communication and computation delay, which means that the BS first updates the global AI / ML model after receiving parameter updates for the local AI / ML model from UE3, as shown at 8081 in Figure 8, and then the BS starts the second iteration with UE3.
[0145] UE1 has the second shortest combined communication and computation delay, which means that the BS then updates the global AI / ML model as shown in 8082 after receiving the parameter update of the first local AI / ML model from UE1, and then the BS starts the second iteration with UE1.
[0146] UE2 has the second shortest combined communication and computation delay, which means that the BS then updates the global AI / ML model as shown in 8083 after receiving the parameter update of the first local AI / ML model from UE2, and then the BS starts the second iteration with UE2.
[0147] BS is 802 for the second iteration for UE3 3,2 , 804 3,2 and 806 3,2 Following the communication and computation delay shown in , once it receives the updated local AI / ML parameters from UE3, it updates the global AI / ML model for the fourth iteration as shown in 8084, after which the BS starts the third iteration with UE3.
[0148] The BS sends 802 for the second iteration for UE1. 1,2 , 804 1,2 and 806 1,2Upon receiving the updated local AI / ML parameters from UE1 following the communication and computation delay shown in , the BS updates the global AI / ML model for the fifth iteration as shown at 8085. After updating the global AI / ML model for the fifth iteration as shown at 8085, the BS may begin a third iteration with UE1 (not shown).
[0149] BS is 802 for the third iteration for UE3 3,3 , 804 3,3 and 806 3,3 When it receives the updated local AI / ML parameters from UE3 following the communication and computation delay shown in 8086, it updates the global AI / ML model for the sixth iteration as shown at 8086. After updating the global AI / ML model for the sixth iteration as shown at 8086, the BS may begin a fourth iteration with UE3 (not shown).
[0150] UE4 has the longest combined communication and computation delays, which in this example are so long that the BS does not receive the first local AI / ML model parameter update from UE4 until it completes its sixth update of the global AI / ML model. Therefore, the BS then performs 802 for the first iteration for UE4. 4,1 , 804 4,1 and 806 4,1 Once it receives the updated local AI / ML parameters from UE4, following the communication and computation delay shown in , it updates the global AI / ML model for the 7th iteration as shown in 8087.
[0151] Because the waiting delays that typically plague synchronous FLs are avoided, asynchronous FL training processes generally have lower training latency than synchronous FL training processes. However, there are two major drawbacks to conventional asynchronous FL training processes. The first drawback is the large communication overhead due to asynchronous DL transmissions. The second drawback is that parameter updates to the local AI / ML model from a lagging UE may be outdated, which can have a negative impact on the accuracy of the global AI / ML model if an outdated update from the lagging UE is used by the BS to update the global AI / ML model. For example, as shown in FIG. 8, the BS does not receive the first local AI / ML model update from UE4 until it has already received and incorporated multiple updates from other UEs. Therefore, updating the global AI / ML model based on the parameters of the local AI / ML model provided by UE4 can adversely affect the accuracy of the global AI / ML model.
[0152] From the above, it can be seen that both synchronous and asynchronous FL have their own drawbacks.
[0153] In addition to FL, large communication overhead and large learning delays also exist in other learning methods. For example, in distributed learning, the UE and network devices cooperatively train an AI model in a manner similar to FL. The main difference between FL and distributed learning is that in FL, DL transmission is via broadcast or groupcast transmission, while in distributed learning, unicast transmission is used for DL.
[0154] Another drawback of existing AI / ML model training procedures relates to the payload size of the exchanged data, which is typically very large. For example, the exchanged data often includes hundreds or thousands of AI / ML model parameters, such as gradients, connection weights, biases, etc. Therefore, due to the often unreliable nature of transmissions in wireless communications and the typically large data volumes for data exchanged between devices for AI training, the air interface resource overhead required for training an AI / ML model can be significant. Therefore, techniques that reduce the overhead and delays associated with online AI / ML model training are highly desirable.
[0155] The present disclosure describes example AI / ML model training procedures that avoid or at least mitigate one or more of the above-mentioned problems associated with conventional AI / ML model training procedures. For example, as described in further detail below, in some embodiments described herein, different techniques are used to configure UEs to selectively participate in an iterative AI / ML model training procedure. For example, a first aspect of the present disclosure provides a semi-synchronous federated learning process in which a group of UEs is configured to participate in a given iteration. Such embodiments are based on a novel feedback signaling mechanism in which UEs report their current processing delay and / or training data volume to a BS, which then uses that information to determine the group of UEs that will participate in a given iteration of the AI / ML model training process. A second aspect of the present disclosure provides an asynchronous number of training iterations for different UEs during the AI / ML model training process. In such embodiments, the AI / ML model training process supports dynamic joining, suspension, or dropping of individual UEs from the training process for one or more iterations. Furthermore, in some embodiments, UEs that dynamically join the training process may be configured to train only a partial model (a subset of parameters) to reduce overhead.
[0156] It should be noted that while many of the following examples are described in the context of federated learning-based or distributed learning-based training procedures for AI / ML models, the techniques described herein may also be applied to AI training with other learning methods, e.g., convergent learning, autoencoders, DNNs (deep neural networks), CNNs (convolutional neural networks), etc.
[0157] Semi-synchronous federated learning As described above, one aspect of the present disclosure provides a quasi-synchronous federated learning process for training an AI / ML model, where the BS can statically and / or dynamically determine the group of UEs to participate in each learning iteration based on semi-static and / or dynamic feedback from the UEs. For example, FIG. 9 shows a timeline of acts 900 performed by four UEs (UE1, UE2, UE3, and UE4) over multiple iterations of a quasi-synchronous federated learning procedure according to one embodiment. In particular, FIG. 9 shows acts performed by the four UEs for the Nth and (N+1)th iterations, where N≧1. The semi-static and / or dynamic feedback from the UEs may include at least one of the UEs' current processing capabilities, the current volume of training data available for the iterative AI / ML model training process at the UEs, or the UEs' sensing capacity to collect training data for the iterative AI / ML model training process.
[0158] In FIG. 9, the delay of communication between the BS and each of the four UEs for the Nth iteration, including transmission delay, retransmission delay, signal processing delay, etc., is shown as 902 for DL communication between the BS and each UE. 1,N , 902 2,N , 902 3,N and 902 4,N 906 for UL communication between each UE and the BS. 1,N , 906 2,N , 906 3,N and 906 4,NFurther, the AI / ML processing delay at each of the four UEs, including, for example, the delay for determining the update of the local AI / ML model for the Nth iteration, is 904 for each of the four UEs, respectively. 1,N , 904 2,N , 904 3,N and 904 4,N Furthermore, the AI / ML processing delay at the BS side for the Nth iteration, including the delay for updating the global model with the update parameters of the local AI / ML model received from the UE, is 908 N is shown in.
[0159] For illustrative purposes, in this example, UE1 and UE2 have high processing capabilities, low communication and AI / ML processing latency, and a large training data volume. In contrast, UE3 has low communication and AI / ML processing latency but only a small volume of training, and UE4 has high communication and processing latency.
[0160] For UEs with low communication and AI / ML processing latency and large data volumes (UE1 and UE2 in the example shown in FIG. 9), the BS may indicate / configure (e.g., through control signaling) these UEs to implement a larger number of learning iterations for other UEs to achieve faster training convergence. For example, as shown in FIG. 9, the BS may indicate / configure (e.g., through control signaling) these UEs to implement a larger number of learning iterations for other UEs to achieve faster training convergence. N After receiving the parameter updates of the first local AI / ML model from UE1 and UE2 as shown in FIG. 1, the BS updates the global AI / ML model, and then starts the second iteration with UE1 and UE2. The communication delay between the BS and UE1 and UE2 for the (N+1)th iteration is 902 for DL communication between the BS and the UEs. 1,N+1 and 902 2,N+1 906 for UL communication between the UE and the BS. 1,N+1 and 906 2,N+1 Furthermore, the AI / ML processing delays at UE1 and UE2 for the (N+1)th iteration are 904, respectively.1,N+1 and 904 2,N is shown in.
[0161] For those UEs with high communication and AI / ML processing latency (UE4 in the example shown in FIG. 9), the BS may indicate to these UEs to implement a smaller number of learning iterations to reduce the waiting delay caused by these lagging UEs. Similarly, for those UEs with small data volumes (UE3 in the example shown in FIG. 9), the BS may indicate to such UEs to implement a smaller number of learning iterations to reduce the communication overhead of UL reporting for parameter updates of local AI / ML models, since the contribution to the global AI / ML model based on such updates from such UEs may be negligible due to the small volume of training data on which the updates are based. For example, as shown in FIG. 9, the BS may indicate to such UEs that 908 in FIG. 9 N+1 As shown in Figure 1, the BS may receive parameter updates for the second local AI / ML model from UE1 and UE2 and update the global AI / ML model after receiving parameter updates for the first local AI / ML model from UE3 and UE4, so that UE1 and UE2 participate in two learning iterations in the time that UE3 and UE4 participate in only one learning iteration. After updating the global AI / ML model based on the updates from UE1 and UE2 for the (N+1)th iteration and the updates from UE3 and UE4 for the Nth iteration, the BS may start the (N+2)th iteration with UE1, UE2, UE3, and UE4.
[0162] A semi-synchronous FL, such as the example shown in FIG. 9, has several potential benefits over traditional synchronous and asynchronous FL. For example, relative to synchronous FL, the semi-synchronous FL disclosed herein can potentially reduce training latency resulting from lagging UEs during synchronous FL and achieve faster training convergence. For example, rather than having to wait for the communication and computation delays of lagging UE4 before updating the global AI / ML model as shown in FIG. 9, the BS can update the global AI / ML model based on updates received from UE1 and UE2 for the Nth iteration, and then update the global AI / ML model again based on updates from UE1 and UE2 for the (N+1)th iteration and updates from UE3 and UE4 for the Nth iteration. Furthermore, relative to asynchronous FL, the semi-synchronous FL disclosed herein can potentially reduce communication overhead associated with DL transmission of a representation of the global AI / ML model during asynchronous FL.
[0163] To assist the BS in determining the participating UE group for a given iteration, each UE sends some assistance information to the BS. For example, as described above, the assistance information may include information indicating the amount of training data available at the UE and / or the AI / ML processing capabilities of the UE. The amount of training data available at different UEs is often unbalanced, and different UEs often have different AI / ML processing capabilities. Furthermore, for a particular UE, the training data volume generally changes over time, and the AI / ML processing capabilities of the UE can change dynamically.
[0164] For example, after a UE performs sensing, there may be a large amount of training data available for training an AI / ML model. However, if the UE periodically measures only some basic channel information at relatively long intervals, there may only be a small amount of training data available at the UE.
[0165] Similarly, the AI / ML processing capabilities of the UE may vary depending on the processing resources utilized by other tasks / services and / or depending on the power saving state of the UE. For example, when the UE is in a power saving mode and / or when the UE is processing other computing tasks, such as sensing processing, the AI / ML processing capabilities of the UE may be lower.
[0166] Thus, in some embodiments, after a UE first accesses a BS, the UE may report to the BS its AI / ML model training capability and its training data acquisition capability. For example, the data acquisition capability may include sensing capability and / or data buffering capability for training data. Furthermore, the UE may semi-statically or dynamically send its current learning capability to the BS. The learning capability may be based on and / or include the UE's current AI / ML processing capability or current training data volume, or a combination of the UE's current AI / ML processing capability and training data volume. Based on the AI / ML model training capability feedback provided by each UE, the BS may then semi-statically or dynamically determine the participating UE group for one training iteration.
[0167] In some embodiments, a UE may semi-statically report the training capability of its AI / ML model in response to an event trigger, such as a change in the UE's learning capability, including, for example, AI / ML processing latency and / or training data volume and / or communication channel quality. For example, when the UE enters a power saving mode, it may have less power available for AI learning and therefore may report a higher processing latency, i.e., a lower AI / ML processing capability. As another example, when the UE has recently collected a large amount of training data, for example by sensing, the UE may report a larger training data volume.
[0168] In some embodiments, the UE may also or instead dynamically provide feedback of the AI / ML model's training capability for each DL reception during the iterative AI / ML model training process. For example, in some embodiments, the UE may use the PUCCH resource indicated in the DCI scheduling the DL transmission for dynamic learning capability feedback, or the UE may use a specific PUCCH resource configured by the BS via RRC / MAC-CE or other DCI for the feedback. As previously described, the content of such dynamic feedback may include the current AI / ML processing delay (e.g., the processing delay for the local model update for this iteration) and / or the current training data volume for this iteration (e.g., the amount of remaining training data available at the UE or the amount of training data the UE expects to be able to acquire).
[0169] In some embodiments, the learning capability may be conveyed as an AI / ML model training capability type selected from among a predefined or configured hierarchy of AI / ML model training capability types. For example, Table 1 includes an example hierarchy of AI / ML model training capability types that includes four levels (i.e., Level 1 through Level 4) of increasing or decreasing AI / ML model training capability.
[0170] [Table 1]
[0171] In Table 1, the level of training capability of each AI / ML model is associated with a corresponding level ID. A UE that has a training capability of a given current AI / ML model that corresponds to one of the predefined or configured levels / types may indicate the corresponding level ID to the BS to advise the BS about the UE's current AI / ML model training capability.
[0172] Referring again to Figure 9, in quasi-synchronous FL, for a given training iteration, only some UEs may participate; for example, in Figure 9, it can be seen that only UE1 and UE2 participate in the (N+1)th iteration. Three example mechanisms for achieving this selective participation in a given iteration are described below. Note that these examples are non-limiting and are provided for illustrative purposes only.
[0173] For example, as a first option, the BS may configure a group of UEs to participate in a given iteration by individually configuring the UEs with rules that allow each UE to decide whether to participate in a given iteration. For example, in such an embodiment, for each DL transmission to indicate updated global AI / ML model parameters, the BS may indicate a value for the iteration ID that will be incremented by 1 in the next iteration. The BS also configures each UE with rules indicating which iteration IDs the UE must ignore or respond to. For example, referring again to FIG. 9, UE1 and UE2 may be configured with a rule indicating that all iteration IDs must not be ignored, while UE3 and UE4 may be configured with a rule indicating that iterations with even-numbered iteration IDs (e.g., the (N+1)th iteration, assuming N is odd) must be ignored.
[0174] As another option, the BS may also or instead indicate the types of UEs participating in the iteration via a DCI (e.g., a DCI scheduling DL transmissions for model updates or other dedicated DCIs that are not scheduling DCIs used to schedule DL transmissions for model updates) or RRC or MAC-CE. For example, as shown in Table 1 above, a UE may report its own training capability by sending a corresponding training level ID to the BS. For a given iteration, the BS may indicate to all UEs or UEs with a training capability greater than a threshold (e.g., level 2) to participate in the iteration. 9, UE1 and UE2 may report to the BS a level ID of 3 (indicating that UE1 and UE2 currently have a level 4 AI / ML model training capability, e.g., high AI / ML model training capability and a large amount of training data), while UE3 may report a level ID of 1 (indicating that UE3 currently has a level 2 AI / ML model training capability (e.g., high AI / ML model training capability but a small amount of training data)), and UE4 may report a level ID of 0 (indicating that UE4 currently has a level 1 AI / ML model training capability, e.g., low AI / ML model training capability and a small amount of training data). In this scenario, for the (N+1)th iteration, the BS may indicate all UEs with a learning capability greater than level 2 (i.e., only UE1 and UE2 in this example scenario) to participate in the (N+1)th iteration.
[0175] As a third option, the BS may also or instead configure the UE with UE-specific monitoring opportunities for the DCI that schedules DL transmissions of global AI / ML model updates for each iteration. For example, to schedule DL transmissions of global AI / ML model updates during the training process, the cyclic redundancy check (CRC) of the DCI used to schedule the DL transmission may be scrambled with a new radio network temporary identifier (RNTI), e.g., a new RNTI, that is different from the cell RNTI (C-RNTI). For example, the new RNTI may be an RNTI used specifically for AI / ML training-related communications in the wireless communication network. The DCI monitoring opportunities (e.g., monitoring symbols, monitoring periodicity) may be configured on a UE-specific basis. For example, referring again to FIG. 9, UE1 and UE2 may each be configured with a shorter DCI monitoring period for the DCI that schedules DL transmissions of global AI / ML model updates, while UE3 and UE4 may each be configured with a longer DCI monitoring period. For example, after all four UEs monitor a DCI scheduling a DL transmission for the Nth iteration, only UE1 and UE2 monitor a DCI scheduling a DL transmission for the (N+1)th iteration, and all four UEs may again monitor a DCI scheduling a DL transmission for the (N+2)th iteration.
[0176] For other DL scheduling, the CRC of the DCI may be scrambled with another RNTI, for example, C-RNTI. Also, monitoring opportunities for this type of DCI and DCI whose CRC is scrambled with a new RNTI may be configured separately. For example, the BS may configure the same monitoring opportunity for DCI scrambled with C-RNTI for UE1, UE2, UE3, and UE4.
[0177] As previously explained, if a lagging UE provides an outdated update to the BS due to, for example, long communication and / or computation delays, utilizing an outdated update can be detrimental to the convergence of the global AI / ML model. To avoid or at least mitigate this possible drawback that may result from asynchronous updates from some UEs, in some embodiments, when a UE reports its updated local AI / ML model parameters to inform the BS about the iteration on which the update is based, the UE is configured to include an indication of the corresponding iteration ID in the UL report. For example, referring again to FIG. 9, 906 1,N , 906 1,N+1 , 906 2,N , 906 2,N+1 , 906 3,N and 906 4,N Each of the UL transmissions shown in indicates a corresponding iteration ID on which the updates to the local AI / ML model included in the UL transmission are based. 1,N , 910 1,N+1 , 910 2,N , 910 2,N+1 , 910 3,N and 910 4,N Therefore, based on the reported iteration ID, the BS may identify whether the local update is out of date and decide whether to discard or utilize the reported data.
[0178] FIG. 10 illustrates an example of a flowchart for semi-synchronous federated learning, according to one embodiment.
[0179] In block 1002, the UE reports its AI / ML model training capability to the BS. As previously described, the AI / ML model training capability may be based on or include the AI / ML processing capability and / or training data volume for the UE and may be conveyed by sending an AI / ML model training capability level ID, as described above with respect to the hierarchy of AI / ML model training capabilities listed in Table 1.
[0180] In block 1004, the BS selects UEs to participate in the training process based on the training capabilities of the AI / ML model reported by the UEs.
[0181] In block 1006, the BS indicates to the UEs that it will participate in training and indicates one or more rules to enable the UEs to decide whether to participate in a given iteration of the training process. For example, as described above, this may involve the BS individually configuring each UE with rules that enable it to decide whether to participate in a given iteration. Alternatively or alternatively, this may involve indicating via RRC or MAC-CE a DCI (e.g., a DCI that schedules DL transmissions for update models) or the type of UE that will participate in the iteration. As another option, the BS may also or instead configure the UEs with UE-specific monitoring opportunities for a DCI that schedules DL transmissions of global AI / ML model updates for each iteration.
[0182] In block 1008, the UE determines whether the UE's AI / ML model training capability has changed (increased or decreased). If not, in block 1010, the UE selectively participates in a given training iteration in accordance with the training rules configured in block 1006. If the training rules indicate that the UE should participate in the current training iteration, in block 1012, the UE receives updated global AI / ML model parameters for the current training iteration from the BS, trains its local AI / ML model based on the updated global AI / ML model parameters from the BS, and reports the updated parameters of the local AI / ML model to the BS, after which the UE returns to block 1008.
[0183] If the UE instead determines in block 1008 that the training capability of its AI / ML model has changed, then in block 1014 the UE reports its updated AI / ML model training capability to the BS, and in response, the BS may optionally indicate updated AI / ML model training rules to the UE, as shown in block 1016, and the method proceeds to block 1012, where the UE selectively participates in the training process in accordance with the updated training rules.
[0184] As previously discussed, a semi-synchronous FL training process, such as the example represented by the flowchart shown in FIG. 10, has several potential benefits over traditional synchronous and asynchronous FL, such as reduced per-iteration training latency, thereby potentially achieving faster training convergence, and / or potentially reducing communication overhead associated with DL transmission of global AI / ML model parameters and / or UL transmission of local AI / ML model parameters.
[0185] Learning asynchronous numbers for different UEs A second aspect of the present disclosure provides a mechanism for asynchronously numbering training iterations for different UEs during an AI / ML model training process. In such embodiments, the AI / ML model training process supports dynamic joining, suspending, or dropping of individual UEs from the training process for one or more iterations. Even in the case of synchronous FL, the number of training iterations for different UEs can vary. In particular, this aspect of the present disclosure enables UEs to be dynamically suspended / dropped from training and to join training based, for example, on dynamic reporting of the UE's AI / ML model training capability. For example, a UE with low AI / ML or training capability may join the training process at a later stage of training to reduce training latency at the beginning of the training process. Furthermore, in some embodiments, a UE that dynamically joins the training process may be configured to train only a partial model (a subset of parameters) to reduce overhead. Meanwhile, for a UE with high AI / ML model training capability, the UE may dynamically suspend or join the training process based on the availability of training data to reduce communication and computing overhead.
[0186] 11 illustrates a timeline of actions 1100 performed by four UEs (UE1, UE2, UE3, and UE4) over multiple iterations of an asynchronous federated learning procedure, according to one embodiment. In particular, FIG. 11 illustrates actions performed by the four UEs for the first, second, Nth, (N+1)th, and (N+2)th iterations of the asynchronous federated learning procedure, where N≧3. The semi-static and / or dynamic feedback from the UEs may include at least one of the UE's current processing capabilities, the current volume of training data available for the iterative AI / ML model training process at the UE, or the UE's sensing capacity to collect training data for the iterative AI / ML model training process.
[0187] In FIG. 11, prior to the first iteration, UE1, UE2, UE3, and UE4 are each 1,1 , 1102 2,1 , 1102 3,1 and 1102 4,1 UEs report their current AI / ML model training capabilities to the BS, as shown in Figure 1. As previously described, the AI / ML model training capabilities reported by the UE may include at least one of the UE's current processing capability, the UE's current volume of training data available for the iterative AI / ML model training process, or the UE's sensing capacity to collect training data for the iterative AI / ML model training process.
[0188] For illustration purposes, in this example, 1,1 , 1102 2,1 and 1102 3,1 The AI / ML model training capabilities reported by UE1, UE2, and UE3 in 1102 show that UE1, UE2, and UE3 have high processing capabilities and large training data volumes. 4,1 The AI / ML model training capabilities reported by UE4 in [https: / / www.microsoft.com / downloads / details.aspx?id=1070707] indicate that UE4 has low processing power (high processing latency) and / or a small training data volume.
[0189] For those UEs with low communication and AI / ML processing latency and large data volumes (UE1, UE2, and UE3 in the example shown in FIG. 11), the BS may indicate / configure (e.g., through control signaling) these UEs to join the training process and participate in the first iteration. On the other hand, for a UE with high AI / ML processing latency and / or small training data volume (UE4 in the example shown in FIG. 11), the BS may instead suspend / drop the UE from the training process, which in this example means that UE4 does not participate in the first iteration with other UEs.
[0190] The participation of UE1, UE2 and UE3 in the first iteration of the training process is 1104, respectively. 1,1 , 1104 2,1 and 1104 3,1 In this example, UE1, UE2, and UE3 continue to participate in the training process from the second iteration through the (N-1)th iteration, as shown in 1104, respectively, as shown in UE1, UE2, and UE3's participation in the first iteration of the training process. 1,2...N-1 , 1104 2,2...N-1 and 1104 3,2...N-1 However, prior to the Nth iteration, UE3 executes 1104 3,N , which indicates that UE 3 now has high processing latency and / or low training data volume. Alternatively, as described in further detail below, rather than providing its updated AI / ML model training capability to the BS, UE 3 may also or instead send an explicit request to the BS to change its learning status.
[0191] 1102 3,N In response to the feedback given by UE3 at 1104, UE3 is dynamically suspended / dropped from the training process for the Nth iteration. In contrast, UE1 and UE2 are dynamically suspended / dropped from the training process for the Nth iteration at 1104, respectively. 1,N and 1104 2,N The UE continues to participate in the Nth iteration as shown in Figure 1. When the UE pauses training, the UE generally does not perform local AI / ML model updates and does not report its local model parameters. However, in some cases, the paused UE may receive DL global AI / ML model updates and store the latest DL AI / ML model updates for future training.
[0192] Prior to the (N+1)th iteration, UE3 may indicate that UE3 currently has low processing delay and / or high training data volume, 1104 3,N+111 , UE4 may report further changes in its AI / ML model training capability as shown in (N+1). For example, UE3 may have collected enough training data to satisfy a minimum threshold for participating in the training process and / or may have exited a power saving mode. Alternatively, as described in further detail below, rather than providing its updated AI / ML model training capability to the BS, UE3 may also or instead send an explicit request to the BS to change its learning status to rejoin the training process. For example, as shown in FIG. 11 , UE4's AI / ML model training capability has not changed (i.e., UE4 still has high processing latency and / or a small training data volume), but UE4 may nevertheless send a request to the UE to dynamically join the training process for the (N+1)th iteration in response to determining that the training process is in a slow stage. For example, the UE may be configured to request dynamic joining in the training process once a threshold number of training iterations have occurred (e.g., in this case, the threshold number may be N). As previously described, the BS may notify the UE of the current iteration ID so that the UE is aware of the current stage of training. When requesting to dynamically join the training process, the UE may inform the BS of the number of the iteration (i.e., iteration ID) to which the UE is joining, allowing the BS to recognize the UEs participating in that iteration.
[0193] 1102 respectively 3,N+1 and 1102 4,N+1 In response to the feedback provided by UE3 and UE4 at 1104, UE3 and UE4 dynamically join the training process with UE1 and UE2 for the (N+1)th iteration. The participation of UE1, UE2, UE3, and UE4 in the (N+1)th iteration is shown in 1104, respectively. 1,N+1 , 1104 2,N+1 , 1104 3,N+1 and 1104 4,N+1As previously described, a UE that dynamically joins the training process late may be configured to train only a partial model (a subset of its local AI / ML model parameters) to reduce overhead. For example, in FIG. 11, UE4 trains only a partial model after dynamically joining the (N+1)th iteration of the training process. For example, UE4 may train fewer gradients / weights of its local AI / ML model that are not stable with the global AI / ML model at the BS.
[0194] FIG. 12 illustrates an example of a flowchart for providing an asynchronous number of learning iterations to different devices, according to one embodiment.
[0195] In block 1202, the UE reports the training capability of its AI / ML model to the BS. As previously described, the training capability of the AI / ML model may be based on or include the AI / ML processing capability and / or training data volume for the UE, and may be conveyed by sending the training capability level ID of the AI / ML model, as described above with respect to the hierarchy of training capabilities of the AI / ML models listed in Table 1.
[0196] In block 1204, the BS selects UEs to participate in the training process based on the training capabilities of the AI / ML models reported by the UEs.
[0197] In block 1206, the BS indicates the training status (e.g., joined or suspended / dropped) of the current AI / ML model to each UE participating in the training.
[0198] In block 1208, the UE determines whether its current AI / ML training status is suspended / dropped, indicating that the UE should not participate in the current iteration of the training process. If the UE's current AI / ML training status is not suspended / dropped, the method proceeds to block 1210, where the UE checks whether its AI / ML training capability has changed. If the UE's AI / ML training capability has not changed, in block 1212, the UE implements the training procedure for the current iteration of the training process, e.g., receiving a DL transmission of an updated global AI / ML model, training its local AI / ML model, and reporting the local AI / ML model update to the BS, after which the method returns to block 1208 to check whether the current AI / ML training status is suspended / dropped.
[0199] If the UE determines that its AI / ML training capability has changed in block 1210, the UE reports its changed AI / ML model training capability to the BS and / or sends a request to the BS to change its AI / ML training status in block 1214. In block 1216, based on the feedback provided by the UE in block 1214, the BS indicates the updated AI / ML training status to the UE, and the method returns from block 1216 to block 1208 to check whether the current AI / ML training status is suspended / dropped.
[0200] If the UE determines in block 1208 that its current learning status is suspended / dropped, the method proceeds to block 1218, where the UE does not update its local AI / ML model and does not report updates to its local AI / ML model to the BS; the UE then checks whether its AI / ML model training capability has changed or whether the training process is in a slow stage in block 1220. If not, the method returns from block 1220 to block 1208. On the other hand, if the UE determines in block 1220 that its AI / ML model training capability has changed and / or the training process is in a slow stage, in some embodiments, the UE may optionally report its changed AI / ML model training capability to the BS and / or send a request to the BS to change its AI / ML training status in block 1222. In this scenario, in block 1216, the BS optionally indicates the updated AI / ML training status to the UE based on the feedback provided by the UE in block 1222, and the method returns from block 1216 to block 1208 to check whether the current AI / ML training status is suspended / dropped.
[0201] As previously explained, allowing dynamic joining and / or pausing / dropping of training so that different UEs may asynchronously participate in different numbers of learning iterations may have several potential benefits, such as reducing UL communication overhead for some UEs, reducing training latency during FL, etc.
[0202] Compressed Feedback Parameters In conventional FL, distributed learning, and other learning processes, the communication overhead utilized for UE model parameter reporting can be quite high due to the large number of parameters (e.g., gradients, weights, biases) that are typically reported, e.g., due to the use of large neural networks containing many neuron nodes and connections of neuron nodes. However, most of the exchanged parameters are redundant.
[0203] Thus, the UE can potentially save communication overhead by compressing the reported AI / ML parameters, e.g., reporting only a few important parameters, without adversely affecting the training process. Another aspect of the present disclosure provides a mechanism for indicating which parameters need to be reported and their values.
[0204] For example, referring again to FIG. 6A , which illustrates an example of a generic AI / ML model, for parameter reporting (e.g., gradients, weights, biases), the UE may send indication signaling to inform the BS about the reported parameters for the AI / ML model. For a type of parameter, e.g., weights, the UE may configure the parameters in a predefined order or an order configured by the BS, e.g., w1, w2, ..., wn. The UE then determines which parameters will be reported and sends assignment information for those parameters to the BS. Four example mechanisms for determining which parameters will be reported and sending assignment information to the BS are described below. Note that these examples are non-limiting and are provided for illustrative purposes only. Parameter Group (PG) based: As a first option, the allocation information may include a bitmap indicating the PG to be reported, where a PG is a set of consecutive parameters. The size of the PG may be configured by the BS or predefined. For example, if the corresponding bit value in the bitmap is 1, the PG may be reported to the BS, but if not, the PG is not reported, or vice versa. Consecutive Parameter Reporting: In this option, the allocation information indicates a set of parameters that are reported consecutively, and the allocation information includes the starting location of the parameters and the number of parameters to be reported. Multiple Clusters of Consecutive Parameters: In this option, the allocation information indicates multiple sets of parameters to be reported consecutively. For each set, the allocation information indicates the starting location of the parameters and the number of parameters to be reported. Parameters between several layers: In this option, the UE reports one or more sets of parameters between several layers of the AI / ML model, for example, a set between layer N and layer M, where 1≦N and M≦the total number of layers of the AI / ML model.
[0205] Furthermore, the UE needs to signal the values of the reported parameters. To reduce communication overhead, in some embodiments, the UE may indicate the range of values to be reported by indicating a range ID based on a predefined or configured mapping between range IDs and value ranges. In such embodiments, a parameter set is associated with one range ID, i.e., parameters within a set have the same value range, and an indication of the individual exact values is given for each parameter in the set. Furthermore, multiple sets of parameters reported in a report may be mapped to different value ranges, where the sets are associated with one range ID.
[0206] A benefit of utilizing range IDs is reduced bit overhead. For example, without using range IDs, 3 bits are needed per parameter to indicate a value from -2 to 2, which means that for N parameters, 3N bits are needed. In contrast, by using range IDs as shown in Table 2 for parameter sets, 1 bit is used to indicate the range ID, and 2 bits are used for each parameter in the set, which means that for N parameters, a total of 2N+1 bits are needed. Thus, for large N, the overhead reduction can be significant.
[0207] [Table 2]
[0208] With reference to the exemplary mapping between Range IDs and value ranges shown in Table 2, it can be seen that by associating Range ID values (represented as one or more bits) with AI / ML model parameter sets, each value range mapped to a Range ID value associated with an AI / ML model parameter set determines the range of the bit string and the meaning of the bit for the AI / ML model's parameter value in the AI / ML model's parameter set. For example, if a given AI / ML model parameter value is represented as the bit string "10" and the Range ID value is "0," then the bit string "10" maps to a decimal value of 1, but if the Range ID value is "1," then the same bit string "10" maps to a decimal value of 2.
[0209] As explained above, the header in the UE's report may signal the indication method and / or the parameter and / or range ID for the set of parameters being reported.
[0210] Furthermore, it should be noted that the use of range ID and allocation information described above in the context of UL parameter transmission is also applicable to DL parameter transmission to reduce communication overhead between the UE and the BS for transmission of AI / ML parameters in DL and UL.
[0211] By implementing the methods disclosed herein, the air interface resource overhead and latency associated with training online AI / ML models may be reduced, providing a trade-off between reduced overhead and training performance.
[0212] Examples of devices (e.g., EDs or UEs and TRPs or network devices) that implement the various methods described herein are also disclosed.
[0213] For example, a first device may include a memory that stores processor-executable instructions and a processor that executes the processor-executable instructions. When the processor executes the processor-executable instructions, the processor may be caused to perform one or more method steps of the devices described herein, e.g., with respect to Figures 9-12. For example, the processor may cause the device to implement operations consistent with an operating mode, e.g., perform necessary measurements and generate content from these measurements as configured for the operating mode, prepare uplink transmissions, process (e.g., encode, decode, etc.) downlink transmissions, and communicate over the air interface in that operating mode by configuring and / or commanding transmit / receive on RF chains and antennas.
[0214] Note that the phrase "at least one of A or B" as used herein is interchangeable with the phrase "A and / or B." It refers to a list from which one can select A or B, or both A and B. Similarly, "at least one of A, B, or C" as used herein is interchangeable with "A and / or B and / or C" or "A, B, and / or C." It refers to a list from which one can select A or B or C, or both A and B, or both A and C, or both B and C, or all of A, B, and C. The same principle applies to longer lists having the same format.
[0215] While the present invention has been described with reference to specific features and embodiments thereof, various modifications and combinations may be made thereto without departing from the invention. The description and drawings should therefore be considered merely illustrative of some embodiments of the invention as defined by the appended claims, and any and all modifications, variations, combinations, or equivalents falling within the scope of the invention are intended to cover. Thus, while the invention and its advantages have been described in detail, various changes, substitutions, and alterations may be made therein without departing from the invention as defined by the appended claims. Moreover, the scope of this application is not limited to the particular embodiments of the processes, machines, manufacture, compositions, means, methods, and steps described herein. As those skilled in the art will readily appreciate from this disclosure of the invention, any now-existing or later-developed processes, machines, manufactures, compositions, means, methods, or steps that perform substantially the same function or achieve substantially the same results as the corresponding embodiments described herein can be utilized in accordance with the present invention. Accordingly, the appended claims are intended to include within their scope such processes, machines, manufactures, compositions, means, methods, or steps.
[0216] Additionally, any module, component, or device illustrated herein that executes instructions may include or otherwise have access to one or more non-transitory computer / processor-readable storage media or media for storage of information such as computer / processor-readable instructions, data structures, program modules, and / or other data. A non-exhaustive list of examples of non-transitory computer / processor-readable storage media includes magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, compact disc read-only memories (CD-ROMs), digital video discs or digital versatile discs (DVDs), optical discs such as Blu-ray Discs™ or other optical storage, volatile and non-volatile removable and non-removable media implemented in any manner or technology, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technology. Any such non-transitory computer / processor storage media may be part of or accessible to or connectable to the device. Any application or module described herein may be implemented using computer / processor readable / executable instructions, which may be stored or otherwise retained by such non-transitory computer / processor readable storage media. Abbreviations LTE Long Term Evolution NR new radio BWP Bandwidth Portion BS base station CA Carrier Aggregation CC Component Carrier CG Cell Group CSI Channel State Information CSI-RS Channel State Information Reference Signal DC Dual Connectivity DCI Downlink Control Information DL Downlink DL-SCH Downlink Shared Channel E-UTRA NR dual connectivity with MCG using EN-DC E-UTRA and with SCG using NR gNB Next Generation (or 5G) Base Station HARQ-ACK Hybrid Automatic Repeat Request Acknowledgment MCG Master Cell Group MCS modulation and coding scheme MAC-CE Medium Access Control - Control Element PBCH Physical Broadcast Channel PCell Primary Cell PDCCH Physical Downlink Control Channel PDSCH Physical Downlink Shared Channel PRACH Physical Random Access Channel PRG Physical Resource Block Group PSCell Primary SCG cell PSS primary synchronization signal PUCCH Physical Uplink Control Channel PUSCH Physical Uplink Shared Channel RACH Random Access Channel RAPID Random Access Preamble Identification RB Resource Block RE Resource Element RRM Radio Resource Management RMSI Residual System Information RS reference signal RSRP reference signal received power RRC Radio Resource Control SCG Secondary Cell Group SFN System Frame Number SL Side Link SCell Secondary Cell SPS Semi-persistent scheduling SR Scheduling Request SRI SRS Resource Indicator SRS Sounding Reference Signal SSS secondary synchronization signal SSB sync signal block SUL Auxiliary Uplink TA Timing Advance TAG Timing Advance Group TUE Target UE UCI Uplink Control Information UE User Equipment UL Uplink UL-SCH Uplink Shared Channel
Claims
1. 1. A method for transmission of artificial intelligence or machine learning (AI / ML) data in a wireless communications network, comprising: determining a training capability of an AI / ML model of a first device, the training capability of the AI / ML model indicating an ability of the first device to participate in a training process of an AI / ML model with at least a second device in the wireless communication network, the training capability of the AI / ML model comprising: i) the current processing power of the first device; and ii) the current volume of training data available for the training process of the AI / ML model at the first device; and iii) a sensing capacity of the first device that collects training data for the training process of the AI / ML model; and and determining the time based on at least one of: transmitting the training capabilities of the AI / ML model to the second device; receiving control signaling from the second device for configuring the first device with rules for determining whether the first device will participate in an iteration of a training process of the AI / ML model; A method for providing
2. The step of transmitting the training capabilities of the AI / ML model to the second device comprises: determining that the training capability of the AI / ML model of the first device has changed; after determining that the training capability of the AI / ML model of the first device has changed, transmitting the changed training capability of the AI / ML model to the second device; The method of claim 1 , comprising:
3. The step of determining that the training capability of the AI / ML model of the first device has changed includes: i) the current processing capacity of the first device; and ii) the current volume of training data available for the training process of the AI / ML model at the first device; and iii) the sensing capacitance of the first device, which collects training data for a training process of the AI / ML model; The method of claim 2 , comprising identifying a change in at least one of:
4. an iteration of the AI / ML model training process is associated with a respective value of an iteration identifier (ID) such that the respective value of the ID is incremented by one for each subsequent iteration; The method of claim 1 , wherein the control signaling configures the first device to selectively participate in a given iteration based on the respective value of the iteration ID associated with the given iteration.
5. Transmission of data or control information from the second device for a given iteration of a training process of the AI / ML model is scheduled by a first downlink control information (DCI), and a cyclic redundancy check (CRC) value of the first DCI is scrambled with a first radio network temporary identifier (RNTI), and the method includes: receiving control signaling from the second device to configure the first device to monitor the first DCI according to a first monitoring period; The method of claim 1 , further comprising:
6. a second DCI for scheduling transmission of other data or control information from the second device, the second DCI having a CRC value scrambled with a second RNTI different from the first RNTI; receiving, from the second device, device-specific control signaling that configures the first device to monitor the second DCI according to a second monitoring period; The method of claim 5 further comprising:
7. The method of claim 6 , wherein the second monitoring period is configured separately from the first monitoring period.
8. sending local AI / ML model update information to the second device for a given iteration of the AI / ML model training process, the local AI / ML model update information comprising AI / ML model parameter updates based on training of the local AI / ML model at the first device, the training of the local AI / ML model at the first device being based on data or control information received from the second device for the given iteration of the AI / ML model training process, the local AI / ML model update information further comprising information indicating a value of an iteration ID associated with the given iteration for which the first device received the data or control information from the second device. The method of any one of claims 1 to 7, further comprising:
9. sending a request to the second device to participate in the training process of the AI / ML model; The method of any one of claims 1 to 8, further comprising:
10. receiving control signaling from the second device that configures the first device to train a partial AI / ML model comprising a partial subset of parameters of a local AI / ML model at the first device; The method of claim 9 further comprising:
11. receiving control signaling from the second device that configures the first device not to participate in the training process of the AI / ML model; The method of any one of claims 1 to 10, further comprising:
12. 1. A method for transmission of artificial intelligence or machine learning (AI / ML) data in a wireless communications network, the method comprising: receiving a training capability of an AI / ML model from a first device, the training capability of the AI / ML model from the first device indicating an ability of the first device to participate in a training process of an AI / ML model with at least a second device in the wireless communication network; transmitting, for each of at least one iteration of the training process of the AI / ML model, information that enables the first device to decide whether to participate in the iteration based on the training capability of the AI / ML model received from the first device; A method for providing
13. The training ability of the AI / ML model is i) the current processing power of the first device; and ii) the current volume of training data available for the iterative AI / ML model training process at the first device; and iii) a sensing capacity of the first device that collects training data for the iterative AI / ML model training process; and The method of claim 12 , comprising at least one of:
14. receiving training capabilities of an AI / ML model from the first device comprises receiving training capabilities of a respective AI / ML model from each device of a plurality of devices including the first device; 14. The method of claim 12 or 13, wherein transmitting information that enables the first device to determine, for each of at least one iteration of the training process of the AI / ML model, whether to participate in the iteration based on the training capability of the AI / ML model received from the first device, comprises transmitting information that enables each device of the plurality of devices to determine, for each of at least one iteration of the training process of the AI / ML model, whether to participate in the iteration based on the training capability of the respective AI / ML model received from each device of the plurality of devices.
15. For each of at least one iteration of the training process of the AI / ML model, transmitting information that enables each device to decide whether to participate in the iteration includes: For each iteration of the AI / ML model training process, sending control signaling for each device of the plurality of devices to configure each device of the plurality of devices with device-specific rules for determining whether each device of the plurality of devices will participate in the iteration. The method of claim 14, comprising:
16. an iteration of the AI / ML model training process is associated with a respective value of an iteration identifier (ID) such that the respective value of the ID is incremented by one for each subsequent iteration; 16. The method of claim 15, wherein the device-specific rules by which the plurality of devices are configured configure each device of the plurality of devices to selectively participate in the given iteration based on the respective value of the iteration ID associated with the given iteration.
17. For each of at least one iteration of the training process of the AI / ML model, transmitting information that enables the first device to decide whether to participate in the iteration includes: transmitting first downlink control information (DCI) including first scheduling information for scheduling transmission of data or control information for a given iteration of a training process of the AI / ML model, wherein a cyclic redundancy check (CRC) value of the first DCI is scrambled using a first radio network temporary identifier (RNTI); transmitting control signaling to configure the first device to monitor the first DCI according to a first monitoring period; The method of claim 12 or 13, comprising:
18. transmitting a second DCI including second scheduling information for scheduling transmission of other data or control information, wherein a CRC value of the second DCI is scrambled with a second RNTI different from the first RNTI; The method further comprises: sending device-specific control signaling to configure the first device to monitor the second DCI according to a second monitoring period; 20. The method of claim 17, further comprising:
19. 20. The method of claim 18, wherein the second monitoring period is configured separately from the first monitoring period.
20. receiving local AI / ML model update information from the first device for a given iteration of the AI / ML model training process, the local AI / ML model update information from the first device comprising AI / ML model parameter updates based on training of the local AI / ML model at the first device, the training of the local AI / ML model at the first device being based on data or control information transmitted from the second device for the given iteration of the AI / ML model training process, the local AI / ML model update information from the first device further comprising information indicating a value of an iteration ID associated with the given iteration for which the first device received the data or control information from the second device.
16. The method of any one of claims 12 to 15, further comprising:
21. receiving a request from the first device that it does not want to participate in the training process of the AI / ML model; 21. The method of any one of claims 12 to 20, further comprising:
22. Apparatus comprising one or more units for carrying out the method according to any one of claims 1 to 11.
23. Apparatus comprising one or more units for carrying out the method according to any one of claims 12 to 21.
24. 22. A processor readable storage medium storing instructions that, when executed by a processor, cause an apparatus to perform the method of any one of claims 1 to 21.
Citation Information
Patent Citations
Terminal and base station
JP2020191623A
Distributed machine learning in an information centric network
US20200027022A1
Information reporting method, apparatus and device, and storage medium
WO2021142609A1