Model inference time and state determinations in wireless communications
By coordinating computing power through nominal and actual inference time management for subsets of activated models, wireless communication systems effectively manage resource constraints, optimizing model inference and state determinations.
Patent Information
- Application Number
- PCT/CN2024/089027
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-10-30
AI Technical Summary
Wireless communication systems face challenges in coordinating computing power for AI/ML model operations due to insufficient resources at a given time, leading to inefficiencies in model inference and state determinations.
User devices and base stations coordinate by transmitting and receiving nominal and actual inference times for subsets of activated models, determining model states, and adjusting operations based on available computing power and additional time units.
This approach optimizes computing power usage, ensuring efficient model inference and state management, even when resources are limited, by dynamically adjusting operations based on available capacity.
Smart Images

Figure CN2024089027_30102025_PF_FP_ABST
Abstract
Description
MODEL INFERENCE TIME AND STATE DETERMINATIONS IN WIRELESS COMMUNICATIONSTECHNICAL FIELD
[0001] This document is directed generally to model inference time and state determinations in wireless communications.BACKGROUND
[0002] In some wireless communication systems, a base station and / or a user device may utilize artificial intelligence (AI) and / or machine learning (ML) to enhance performance. For example, a wireless communication system may employ one or more AI / ML models to perform one or more functions or operations. However, use of AI / ML models to perform operations consumes computing power. In any of various situations, the amount of computing power available to a device to use a model to perform a function at a given point in time may not be sufficient. As such, ways to coordinate computing power in wireless communication systems may be desirable.SUMMARY
[0003] This document relates to methods, systems, apparatuses and devices for wireless communication. In some implementations, a method for wireless communication includes: transmitting, by a user device, a nominal inference time of an M number of models to a base station, wherein M is an integer of two or more; receiving, by the user device, an indication of a subset models from among the M number of models activated by the base station, wherein the subset comprises an N number of models activated by the base station, wherein N is an integer of one or more, and N is less than M; determining, by the user device an actual inference time of the subset based on the subset of models that are activated and the nominal inference time.
[0004] In some other implementations, a method for wireless communication includes: receiving, by a base station, a nominal inference time of an M number of models from a user device, wherein M is an integer of two or more; activating, by the base station, a subset of models from among the M number of models, wherein the subset comprises an N number of models activated by the base station, wherein N is an integer of one or more, and N is less than M; and transmitting, by the base station, an indication of the subset of models that is activated to a user device.
[0005] In some other implementations, a method for wireless communication includes: determining, by a user device, a state of a plurality of states for a model; and determining, by the user device, whether or not to perform at least one operation associated with the model according to the state, wherein the at least one operation comprises at least one of: model inference or model monitoring.
[0006] In some other implementations, a device, such as a network device, is disclosed. The device may include one or more processors and one or more memories, wherein the one or more processors are configured to read computer code from the one or more memories to implement any of the methods above.
[0007] In yet some other implementations, a computer program product is disclosed. The computer program product may include a non-transitory computer-readable program medium with computer code stored thereupon, the computer code, when executed by one or more processors, causing the one or more processors to implement any of the methods above.
[0008] The above and other aspects and their implementations are described in greater detail in the drawings, the descriptions, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] FIG. 1 shows a block diagram of an example of a wireless communication system.
[0010] FIG. 2 shows a flow chart of a method for wireless communication.
[0011] FIG. 3 shows a flow chart of another method for wireless communication.
[0012] FIG. 4 shows a flow chart of another method for wireless communication.
[0013] Fig. 5 shows a timing diagram of an example implementation of a model pattern.
[0014] Fig. 6 shows a schematic diagram of an overall state switching of model states.DETAILED DESCRIPTION
[0015] The example headings for the various sections below are used to facilitate the understanding of the disclosed subject matter and do not limit the scope of the claimed subject matter in any way. Accordingly, one or more features of one example section can be combined with one or more features of another example section. Furthermore, 5G terminology is used for the sake of clarity of explanation, but the techniques disclosed in the present document are not limited to 5G technology only, and may be used in wireless systems that implemented other protocols, e.g., 5G- Advanced (5G-A) , 6G or beyond.
[0016] The present description describes various embodiments of systems, apparatuses, devices, and methods for wireless communications related to model inference time and / or state determinations.
[0017] Fig. 1 shows a diagram of an example wireless communication system 100 (also called herein a mobile communication system) including a plurality of communication nodes (or just nodes) that are configured to wirelessly communicate with each other. In general, the communication nodes include at least one user device 102 and at least one network device or base station 104. The example wireless communication system 100 in Fig. 1 is shown as including two user devices 102, including a first user device 102(1) and a second user device 102(2) , and a network device 104. However, various other examples of the wireless communication system 100 that include any of various combinations of one or more user devices 102 and / or one or more network devices 104 may be possible.
[0018] In general, a user device as described herein, such as the user device 102, may include a single electronic device or apparatus, or multiple (e.g., a network of) electronic devices or apparatuses, capable of communicating wirelessly over a network. A user device may comprise or otherwise be referred to as a user terminal, a user terminal device, or a user equipment (UE) . Additionally, a user device may be or include, but not limited to, a mobile device (such as a mobile phone, a smart phone, a smart watch, a tablet, a laptop computer, vehicle or other vessel (human, motor, or engine-powered, such as an automobile, a plane, a train, a ship, or a bicycle as non-limiting examples) or a fixed or stationary device, (such as a desktop computer or other computing device that is not ordinarily moved for long periods of time, such as appliances, other relatively heavy devices including Internet of things (IoT) , or computing devices used in commercial or industrial environments, as non-limiting examples) . In various embodiments, a user device 102 may include transceiver circuitry 106 coupled to an antenna 108 to effect wireless communication with the network device 104. The transceiver circuitry 106 may also be coupled to a processor 110, which may also be coupled to a memory 112 or other storage device. The memory 112 may store therein instructions or code that, when read and executed by the processor 110, cause the processor 110 to implement various ones of the methods described herein.
[0019] Additionally, in general, a network device as described herein, such as the network device 104, may include a single electronic device or apparatus, or multiple (e.g., a network of) electronic devices or apparatuses, and may comprise one or more wireless access nodes, base stations, or other wireless network access points capable of communicating wirelessly over a network with one or more user devices and / or with one or more other network devices 104. For example, the network device 104 may comprise a 4G LTE base station, a 5G NR base station, a 5G central-unit base station, a 5G distributed-unit base station, a next generation Node B (gNB) , an enhanced Node B (eNB) , or other similar or next-generation (e.g., 6G) base stations, in various embodiments. A network device 104 may include transceiver circuitry 114 coupled to an antenna 116, which may include an antenna tower 118 in various approaches, to effect wireless communication with the user device 102 or another network device 104. The transceiver circuitry 114 may also be coupled to one or more processors 120, which may also be coupled to a memory 122 or other storage device. The memory 122 may store therein instructions or code that, when read and executed by the processor 120, cause the processor 120 to implement one or more of the methods described herein.
[0020] In various embodiments, two communication nodes in the wireless system 100-such as a user device 102 and a network device 104, two user devices 102 without a network device 104, or two network devices 104 without a user device 102-may be configured to wirelessly communicate with each other in or over a mobile network and / or a wireless access network according to one or more standards and / or specifications. In general, the standards and / or specifications may define the rules or procedures under which the communication nodes can wirelessly communicate, which, in various embodiments, may include those for communicating in millimeter (mm) -Wave bands, and / or with multi-antenna schemes and beamforming functions. In addition or alternatively, the standards and / or specifications are those that define a radio access technology and / or a cellular technology, such as Fourth Generation (4G) Long Term Evolution (LTE) , Fifth Generation (5G) New Radio (NR) , or New Radio Unlicensed (NR-U) , as non-limiting examples.
[0021] Additionally, in the wireless system 100, the communication nodes are configured to wirelessly communicate signals between each other. In general, a communication in the wireless system 100 between two communication nodes can be or include a transmission or a reception, and is generally both simultaneously, depending on the perspective of a particular node in the communication. For example, for a given communication between a first node and a second node where the first node is transmitting a signal to the second node and the second node is receiving the signal from the first node, the first node may be referred to as a source or transmitting node or device, the second node may be referred to as a destination or receiving node or device, and the communication may be considered a transmission for the first node and a reception for the second node. Of course, since communication nodes in a wireless system 100 can both send and receive signals, a single communication node may be both a transmitting / source node and a receiving / destination node simultaneously or switch between being a source / transmitting node and a destination / receiving node.
[0022] Also, particular signals can be characterized or defined as either an uplink (UL) signal, a downlink (DL) signal, or a sidelink (SL) signal. An uplink signal is a signal transmitted from a user device 102 to a network device 104. A downlink signal is a signal transmitted from a network device 104 to a user device 102. A sidelink signal is a signal transmitted from one user device 102 to another user device 102, or a signal transmitted from one network device 104 to another network device 104. Also, for sidelink transmissions, a first / source user device 102 directly transmits a sidelink signal to a second / destination user device 102 without any forwarding of the sidelink signal to a network device 104. Similarly, a first / source network device 104 directly transmits a sidelink signal to a second / destination network device 104 without any forwarding of the sidelink signal to a user device 102.
[0023] Additionally, at least some signals communicated between communication nodes in the system 100 may be characterized or defined as a data signal or a control signal. In general, a data signal is a signal that includes or carries data, such multimedia data (e.g., voice and / or image data) , and a control signal is a signal that carries control information that configures the communication nodes in certain ways in order to communicate with each other, or otherwise controls how the communication nodes communicate data signals with each other. Also, certain signals may be defined or characterized by combinations of data / control and uplink / downlink / sidelink, including uplink control signals, uplink data signals, downlink control signals, downlink data signals, sidelink control signals, and sidelink data signals.
[0024] For at least some specifications, such as 5G NR, data and control signals are transmitted and / or carried on physical channels. Generally, a physical channel corresponds to a set of time-frequency resources used for transmission of a signal. Different types of physical channels may be used to transmit different types of signals. For example, physical data channels (or just data channels) , also herein called traffic channels, are used to transmit data signals, and physical control channels (or just control channels) are used to transmit control signals. Example types of traffic channels (or physical data channels) include, but are not limited to, a physical downlink shared channel (PDSCH) used to communicate downlink data signals, a physical uplink shared channel (PUSCH) used to communicate uplink data signals, and a physical sidelink shared channel (PSSCH) used to communicate sidelink data signals. In addition, example types of physical control channels include, but are not limited to, a physical downlink control channel (PDCCH) used to communicate downlink control signals, a physical uplink control channel (PUCCH) used to communicate uplink control signals, and a physical sidelink control channel (PSCCH) used to communicate sidelink control signals. As used herein for simplicity, unless specified otherwise, a particular type of physical channel is also used to refer to a signal that is transmitted on that particular type of physical channel, and / or a transmission on that particular type of transmission. As an example illustration, a PDSCH refers to the physical downlink shared channel itself, a downlink data signal transmitted on the PDSCH, or a downlink data transmission. Accordingly, a communication node transmitting or receiving a PDSCH means that the communication node is transmitting or receiving a signal on a PDSCH.
[0025] Additionally, for at least some specifications, such as 5G NR, and / or for at least some types of control signals, a control signal that a communication node transmits may include control information comprising the information necessary to enable transmission of one or more data signals between communication nodes, and / or to schedule one or more data channels (or one or more transmissions on data channels) . For example, such control information may include the information necessary for proper reception, decoding, and demodulation of a data signals received on physical data channels during a data transmission, and / or for uplink scheduling grants that inform the user device about the resources and transport format to use for uplink data transmissions. In some embodiments, the control information includes downlink control information (DCI) that is transmitted in the downlink direction from a network device 104 to a user device 102. In other embodiments, the control information includes uplink control information (UCI) that is transmitted in the uplink direction from a user device 102 to a network device 104, or sidelink control information (SCI) that is transmitted in the sidelink direction from one user device 102(1) to another user device 102(2) , or from one network device 104(1) to another network device 104(2) .
[0026] Fig. 2 is a flow chart of an example method 200 for wireless communication related to model inference time. At block 202, a user device 102 transmits a nominal inference time of an M-number of models to a base station 104, wherein M is an integer of two or more. At block 204, the user device 102 receives an indication of a subset models from among the M-number of models activated by the base station 104. The subset includes an N number of models activated by the base station 104, where N is an integer of one or more, and N is less than M. At block 206, the user device 102 determines an actual inference time of the subset based on the subset of models that are activated and the nominal inference time.
[0027] Fig. 3 is a flow chart of another example method 300 for wireless communication related to model inference time. At block 302, a base station 104 receives a nominal inference time of an M number of models from a user device, wherein M is an integer of two or more. At block 304, the base station 104 activates a subset of models from among the M-number of models. The subset includes an N number of models activated by the base station, where N is an integer of one or more, and N is less than M. At block 306, the base station 104 transmits an indication of the subset of models that is activated to a user device 102.
[0028] Fig. 4 is a flow chart of an example method 400 for wireless communication related to model states. At block 402, a user device 102 determines a state of a plurality of states for a model. At block 404, the user device 102 determines whether or not to perform at least one operation associated with the model according to the state, where the at least one operation includes at least one of: model inference or model monitoring.
[0029] In some implementations of the method 200, 300, and / or 400, the actual inference time is determined based on the nominal inference time and the N number of models of the subset activated by the base station.
[0030] In some implementations of the method 200, 300, and / or 400, the actual inference time is determined based on the nominal inference time, the N number of models of the subset activated by the base station, and an additional time unit.
[0031] In some implementations of the method 200, 300, and / or 400, the actual inference time includes an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to:
[0032] where Ti is a nominal inference time of the i-th model, N is the number of models of the subset, TΔ is the additional time unit, and i is an integer where 1≤i≤N.
[0033] In some implementations of the method 200, 300, and / or 400, the actual inference is determined based on the nominal inference time, the number of models of the subset, a total artificial intelligence processing unit (APU) for the user device 102, and an APU requirement of each model of the subset.
[0034] In some implementations of the method 200, 300, and / or 400, the actual interference time includes an i-th actual inference time of an i-th model, and the i-th actual inference time is determined according to:
[0035] or
[0036] where Ti is a nominal inference time of the i-th model, Atotal is the total APU for the user device, Ai is an APU requirement of the i-th model, is a total APU requirement of the subset of models, and i is an integer where 1≤i≤N.
[0037] In some implementations of the method 200, 300, and / or 400, the total APU for the user device 102 and the APU requirement for the i-th model are reported by the user device 102, indicated by the base station 104, or predefined.
[0038] In some implementations of the method 200, 300, and / or 400, the actual inference time is determined based on the nominal inference time, the N number of models activated by the base station 104, and a nominal number of activated models.
[0039] In some implementations of the method 200, 300, and / or 400, the actual inference time is determined further based on an additional time unit.
[0040] In some implementations of the method 200, 300, and / or 400, the actual inference time includes an i-th actual inference time of an i-th model, and the i-th actual inference time is determined according to:
[0041] where Ti is a nominal inference time of the i-th model, N is the number of models of the subset, TΔ is the additional time unit, i is an integer where 1≤i≤N, and P is the nominal number of activated models.
[0042] In some implementations of the method 200, 300, and / or 400, the actual inference time is determined based on the nominal inference time and a number of layers.
[0043] In some implementations of the method 200, 300, and / or 400, the actual inference time includes an i-th actual inference time of an i-th model, and the i-th actual inference time is determined according to:
[0044] where Ti is a nominal inference time of the i-th model, L is the number of layers, and i is an integer where 1≤i≤N.
[0045] In some implementations of the method 200, 300, and / or 400, the actual inference time is determined based a maximum number of layers.
[0046] In some implementations of the method 200, 300, and / or 400, the actual inference time includes an i-th actual inference time of an i-th model, and the i-th actual inference time is determined according to:
[0047] where Ti is a nominal inference time of the i-th model, L is the number of layers, Lmax is the maximum number of layers, and i is an integer where 1≤i≤N.
[0048] In some implementations of the method 200, 300, and / or 400, the actual inference time is determined based on the nominal inference time and a number of time instances of a model input.
[0049] In some implementations of the method 200, 300, and / or 400, the actual inference time includes an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to:
[0050] where Ti is a nominal inference time of the i-th model, I is the number of time instances, and i is an integer where 1≤i≤N.
[0051] In some implementations of the method 200, 300, and / or 400, the actual inference time is determined based on a maximum number of time instances of a model input.
[0052] In some implementations of the method 200, 300, and / or 400, the actual inference time includes an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to:
[0053] where Ti is a nominal inference time of the i-th model, I is the number of time instances, Imax is the maximum number of time instances, and i is an integer where 1≤i≤N.
[0054] In some implementations of the method 200, 300, and / or 400, the user device 102 receives, and / or the base station 104 transmits, a command from the base station 104. The command indicates to activate the subset of models. Additionally, the user device 102 sends, and / or the base station 104 receives, a report of the actual inference time of the subset. In any of various implementations, the command is the same as, or different or separate from, the indication that the user device 102 receives at block 204, and / or that the base station 104 transmits at block 304.
[0055] In some implementations of the method 200, 300, and / or 400, the actual inference time is valid for a valid time duration Tvalid. In some of these implementations, the valid time duration Tvalid is indicated by the user device 102 or configured by the base station 104.
[0056] In some implementations of the method 200, 300, and / or 400, the command indicates an expected inference time of the subset of models.
[0057] In some implementations of the method 200, 300, and / or 400, the user device 102 transmits, and / or the base station 104 receives, an assessment report that indicates whether or not the expected inference time can be satisfied.
[0058] In some implementations of the method 200, 300, and / or 400, the user device 102 sends the report of the actual inference time of the subset according to a reporting period indicated by the base station 104. In addition or alternatively, the base station 104 indicates, to the user device 102, the reporting period according to which the user device 102 is to report the actual inference time.
[0059] In some implementations of the method 200, 300, and / or 400, the reporting period is equal to a valid time duration Tvalid that the actual inference time is valid.
[0060] In some implementations of the method 200, 300, and / or 400, the user device 102 transmits, and / or the base station 104 receives, the valid time duration with the actual inference time according to the reporting period.
[0061] In some implementations of the method 200, 300, and / or 400, the plurality of states for the model includes: an activated state in which the user device 102 performs model inference for an activated state inference period and model monitoring for an activated state monitoring period a deactivated state in which the user device 102 does not perform model inference and does not perform model monitoring, or in which the user device performs model monitoring for a deactivated state monitoring period that is longer than the activated state monitoring period and an inactive state in which the user device 102 does not perform model inference and performs model monitoring, where the model monitoring is performed for an inactive state monitoring period that is longer or shorter than the activated state monitoring period or performs model inference and performs model monitoring, where the model inference is performed for an inactive state inference period that is longer than the activated state inference period
[0062] In some implementations of the method 200, 300, and / or 400, the user device 102 receives a model pattern, where the model pattern includes a pattern period and an active time duration, and where the activate time duration is smaller than the pattern period
[0063] In some implementations of the method 200, 300, and / or 400, the user device 102 receives the model pattern from the base station 104, and / or the base station 104 transmits the model pattern to the user device 102.
[0064] In some implementations of the method 200, 300, and / or 400, the active time duration starts from a beginning of each pattern period or from an offset relative to the beginning of each pattern period.
[0065] In some implementations of the method 200, 300, and / or 400, the model is in an activated state during the active time duration, and the model is in an inactive state during an inactive time duration.
[0066] In some implementations of the method 200, 300, and / or 400, wherein a timer controls the model. In some of these implementations, the user device 102 restarts the timer in response to a timer restart condition being satisfied. In addition or alternatively, in some of these implementations, the timer is configured by the base station.
[0067] In some implementations of the method 200, 300, and / or 400, the timer restart condition includes at least one of: the user device receives a command to activate the model; the user device receives a command to perform model monitoring; the user device reports a model monitoring result; the user device reports an inference time; the user device reports a model activation report; or the user device reports an inference result.
[0068] In some implementations of the method 200, 300, and / or 400, in response to the timer restart condition not being satisfied during a single time unit, the user device decrements the timer at an end of the single time unit. In some of these implementations, the single time unit includes: one frame, one subframe, one slot, one sub-slot, one second, one millisecond, or one period.
[0069] In some implementations of the method 200, 300, and / or 400, the user device 102 switches the model to an inactivate state when the timer expires.
[0070] In some implementations, wherein the user device 102 indicates a first set of one or more models and a second set of one or more models to the base station 104. In any of various of these implementations, the user device 102 indicates the first set and the second set, in combination with, or independent of, actions performed related to a plurality of states, such as performed at block 402 and / or block 404.
[0071] In some implementations of the method 200, 300, and / or 400, the set is available at the user device 102, and the second set is available at a server of the user device 102.
[0072] In some implementations of the method 200, 300, and / or 400, the first set requires a first time delay for the user device 102 to activate the one or more models of the first set, and the second set requires the first time delay plus a second time delay for the user device 102 to activate the one or more models of the second set, where the first time delay is for the user device 102 to prepare for model activation, and the second time delay is for the user device 102 to download the one or more models of the second set.
[0073] In some implementations of the method 200, 300, and / or 400, the user device 102 receives an indication of a first set of one or more models and a second set of one or more models from the base station 104. In some of these implementations, the first set is configured by the base station 104, the second set is indicated by downlink control information (DCI) or a medium access control (MAC) -control element (CE) , the second set is selected from the first set of models, and / or the base station 104 indicates to the user device that at least one of the one or more models of the second set is activated.
[0074] Other methods are possible, including but not limited to those that combine one or more blocks or actions from each of two or more of the methods 200, 300, 400, and / or those that include fewer than all of actions performed for an above recited implementation of the methods 200, 300, 400. For example, some other methods may communicate indications of first and second sets of one or more models independent from, or without necessarily, determining one or more states of the models. Various methods or combinations of methods based on the numerous implementations of the methods 200, 300, 400 are possible.
[0075] Further details of actions performed by communication nodes in the wireless communication system 100, any or all of which may be implemented in any of various implementations of the method 200, the method 300, the method 400, and / or other methods, are now described.
[0076] Some implementations of the wireless communication system 100 may utilize artificial intelligence (AI) and / or machine learning (ML) (herein collectively referred to as AI / ML) to enhance performance and / or operation. For example, implementation of AI / ML technology into the wireless communication system may improve the system operating efficiency, such as by reducing overhead of reference signals via AI / ML inference and prediction.
[0077] Additionally, some implementations where the wireless communication system 100 implements AI / ML technology, the wireless communication system 100 includes or adopts an AI / ML model that performs one or more functions or operations, such as inference for example. Generally, a model may refer to a functionality, function, functionality module, function module, processing method, information processing method, implementation, feature, feature group, configuration, configuration set, dataset (e.g., for model training) , or data-driven algorithms. Generally, these models are performed, calculated, or processed by a user device 102 (e.g., User Equipment (UE) ) , but may also be performed at a base station 104 in any of various implementations.
[0078] Additionally, in any various implementations, a model (as used herein, unless expressly described otherwise, the term model is used to refer to, and / or be used interchangeably with, an AI model, an ML model, or an AI / ML model) may be a data driven algorithm that applies AI / ML techniques to generate a set of outputs based on a set of inputs. Alternatively, a model can be linear or non-linear algorithms or combination of both algorithms. In addition, functionality may refer to a feature enabled by the AI / ML model. Alternatively, functionality may refer to a set of parameters or configurations for one feature. For example, a user device 102 may include, adopt, utilize, or implement a convolutional neural network (CNN) model to predict the beams for the communication, and the CNN model is the model and the beam prediction is the functionality. Different models and / or functionalities may be associated with different configurations (e.g., a Radio Resource Control (RRC) configuration) . Additionally, as used herein, model activation may refer to activating a corresponding configuration for the user device 102. Similarly, model deactivation, switching, and fallback may refer to deactivating the corresponding configuration, switching the configuration, and falling back to a configuration without the model, respectively.
[0079] Additionally, in some implementations, an AI model may perform one or more functions, such as model training and model inference. Performance of such functions consumes computing power. Described herein include ways for communication nodes in the wireless communication system 100 to coordinate computing power, including in situations where computing power in one or more communication nodes is not sufficient.
[0080] In some implementations or scenarios, different user devices 102 may be equipped with different computing power. In addition or alternatively, different models with different complexity in the same user device 102 may require different computing power. However, reporting the computing power of each user device 102 may undesirably include the disclosure of the implementation design detail of the user device 102. To avoid such reporting, the user device 102 may report the inference time of each model to the base station.
[0081] In some implementations, a user device 102 may report a nominal inference time of an M number of models to the network device or base station 104. In response, the base station 104 may activate a subset of the M models to the user device 102. As used herein unless expressly described otherwise, the subset is described as including N models, where M and N are both integers, M and N are each one or more, and N is not larger than M. In some of these implementations, M is two or more, and M is larger than N. The user device 102 may determine an actual inference time based on the activated models of the subset and the nominal inference time.
[0082] Additionally, in some implementations, the nominal inference time refers to an inference time when all of the computing power of a given user device 102 is allocated to a model. Thus, the more models that are activated, the longer the inference time required by the user device. In any of various implementations, the actual inference time is determined according to the following schemes.
[0083] In a first scheme, the actual inference time is determined based on the nominal inference time and the number N of activated models. In some implementations of the first scheme, suppose a nominal inference time of an i-th model denoted as Ti, the actual inference time for the i-th model is denoted as and N models are activated. In some implementations, the actual inference time for the i-th model is determined according to the following mathematical formula: where i is an integer number such that 1≤i≤N. For example, if a user device 102 reports the inference time as 1 millisecond (ms) for one model, such reporting may indicate or refer to that the user device 102 requires 1 ms to process the model inference if all of the computing power of the user device 102 is allocated to the model. Staying with this example, if two models are activated (i.e., N=2) , the computing power is allocated to both of the two models, and in turn, the user device 102 requires double the amount of processing time (i.e., 2 ms in this is case) to process the inference for the model.
[0084] In a second scheme, the actual inference time is determined based on the nominal inference time, the number N of activated models, and an additional time unit. In some implementations of the second scheme, where N models are activated, the actual inference time for the i-th model is determined according to the following mathematical formula: where Ti is the nominal inference time of the i-th model, N is the number of activated models,TΔ is the additional time unit, and i is an integer number satisfying 1≤i≤N. Additionally, in some implementations of the second scheme, the additional time unit may be configured by the base station 104, may be reported by the user device 102, or may be predefined.
[0085] In a third scheme, the actual inference time is determined based on the nominal inference time, the number N of activated models, and an additional time unit. In some implementations of the third scheme, where N models are activated, the actual inference time for an i-th model is determined according to the following mathematical formula: where Ti is the nominal inference time of the i-th model, N is the number of activated models, TΔ is the additional time unit, and i is an integer number satisfying 1≤i≤N. To illustrate, suppose the user device 102 reports an inference time of 1 ms for one model and reports an additional time unit of 0.5 ms. Such reporting indicates or refers to that if all of the computing power of the UE is allocated to a model, the user device 102 requires 1 ms to process the model inference. Additionally, if two models are activated (i.e., N=2) , the computing power is allocated to both of the two models. Correspondingly, the user device 102 requires longer inference time 1.5ms (i.e., ) to process the inference. Additionally, in some implementations of the third scheme, the additional time unit may be configured by the base station 104, may be reported by the user device 102, or may be predefined.
[0086] In a fourth scheme, the actual inference time is determined based on the nominal inference time, the number N of activated models, a total AI processing unit (APU) , and an APU requirement of each model. In some implementations of the fourth scheme, the actual inference time for an i-th model is determined according to the following mathematical formula: If the nominal inference time of the ith model is Ti and N models are activated, then the actual inference time of this model is where Ti is the nominal inference time of the i-th model, N is the number of activated models, Atotal is the total APU of the user device 102, Ai is the APU requirement of the i-th activated model, is the total APU requirement of the N activated model (s) , and i is integer number satisfying 1≤i≤N. In some of these implementations, the APU requirement of the i-th model Ai is used to represent the hardware and / or software resource (e.g., computing power) requirement of the i-th model. Additionally, the total APU Atotal represents the user device’s 102 hardware and / or software resource (e.g., computing power) capability. Additionally, in any of various implementations, the total APU Atotal and / or the APU requirement Ai of the i-th model may be reported by the user device 102, indicated by the base station 104, or predefined.
[0087] Additionally, in some implementations, the nominal inference time refers to an inference time when the computing power of the user device 102 is sufficient to power all of a certain number of activated models. However, if one or more additional models are to be activated beyond the certain number, then the computing power requirement of the total number of activated models, including the certain number and the one or more additional models, may exceed the capability of the user device 102. In turn, the user device 102 may require a longer inference time.
[0088] The following describes additional schemes that may be used to determine the actual inference time, including in situations where a number of activated models requires computing power beyond the capability of the user device 102.
[0089] In a fourth scheme, the actual inference time is determined based on the nominal inference time, the number N of activated models, and a nominal number of activated models. In some implementations of the fourth scheme, the actual inference time for an i-th model is determined according to the following mathematical formula:
[0090] where Ti is the nominal inference time of the i-th model, N is the number of activated models, P is the nominal number of activated models, and i is an integer number satisfying 1≤i≤N. Effectively, if the number N of activated models exceeds the nominal number P of models, more inference time is needed. To illustrate, suppose the user device 102 reports the inference time as 1 ms for one model, and reports a nominal number of activated models of 2. Correspondingly, if only one model is activated, then the inference time for the one activated model is 1 ms. However, if, for example, three models are activated, then 1.5 ms inference time is used.
[0091] In a fifth scheme, the actual inference time is determined based on the nominal inference time, the number N of activated models, the nominal number of activated models, and an additional time unit. In some implementations of the fifth scheme, the actual inference time for an i-th model is determined according to the following mathematical formula:
[0092] where Ti is the nominal inference time of the i-th model, N is the number of activated models, P is the nominal number of activated models, TΔ is the additional time unit, and i is an integer number satisfying 1≤i≤N. Effectively, if the number of activated models exceeds the nominal number of models, more inference time is needed. To illustrate, suppose the user device 102 reports the inference time as 1 ms for one model, and reports a nominal number of activated models of 2. Correspondingly, if only one model is activated, then the inference time for the one model is 1 ms. However, if, for example, 3 models are activated, then 1.5 ms inference time is used, assuming TΔ is 0.5 ms.
[0093] In a sixth scheme, the actual inference time is determined based on the nominal inference time, number N of activated models, the total APU, and the APU requirement of each model. In some implementations of the sixth scheme, the actual inference time for an i-th model is determined according to the following mathematical formula:
[0094] where Ti is the nominal inference time of the i-th model, N is the number of activated models, Atotal is the total APU of the user device 102, Ai is the APU requirement of the i-th activated model, is the total APU requirement of the N activated model (s) , and i is integer number satisfying 1≤i≤N.
[0095] Additionally, in some implementations, the user device 102 may determine the actual inference time based on an input dimension. For example, if the number of inputs is relatively large, the user device 102 may require a longer inference time.
[0096] The following describes additional schemes that may be used to determine the actual inference time.
[0097] In a seventh scheme, the actual inference time is determined based on the nominal inference time and a number layers. In some implementations of the seventh scheme, the actual inference time for an i-th model is determined according to the following mathematical formula: where Ti is the nominal inference time of the i-th model, L is the number of layers, and i is an integer number satisfying 1≤i≤N. In some of these implementations, the nominal inference time is determined based on one layer. For example, the layer may be or refer to a multiple input multiple output (MIMO) layer or a model layer. In addition or alternatively, in some implementations, the number of layers Lis configured by the base station 104.
[0098] In an eighth scheme, the actual inference time is determined based on the nominal inference time, a number layers, and a maximum number of layers. In some implementations of the eighth scheme, the actual inference time for an i-th model is determined according to the following mathematical formula: where Ti is the nominal inference time of the i-th model, L is the number of layers, Lmax is the maximum number of layers, and i is an integer number satisfying 1≤i≤N. Under the eighth scheme, effectively, the nominal inference time is determined based on the maximum number of layers. In some implementations, the layer for the eighth scheme is or refers to the MIMO layer or model layer.
[0099] In a ninth scheme, the actual inference time is determined based on the nominal inference time, and a number of time instances of model input. In some implementations of the ninth scheme, the actual inference time for an i-th model is determined according to the following mathematical formula: where Ti is the nominal inference time of the i-th model, I is the number of time instances of model input, and i is an integer number satisfying 1≤i≤N. Under the ninth scheme, effectively, the nominal inference time may be determined based on one time instance. To illustrate using channel state information (CSI) prediction as an example, if the user device 102 uses or applies five historical CSI as model input in the time domain, then I is equal to 5. Additionally, in some implementations of the ninth scheme, the number of time instances I may be indicated or configured by the base station 104.
[0100] In a tenth scheme, the actual inference time is determined based on the nominal inference time, a number of time instances of model input, and a maximum number of time instances of model input. In some implementations of the tenth scheme, the actual inference time for an i-th model is determined according to the following mathematical formula: where Ti is the nominal inference time of the i-th model, I is the number of time instances of model input, Imax is the maximum number of time instances, and i is an integer number satisfying 1≤i≤N. Under the tenth scheme, effectively, the nominal inference time may be determined based on the maximum number of time instances of model input.
[0101] Additionally, in some implementations, the base station 104 may determine whether to activate or deactivate a model. However, the inference time of the model at the user device (e.g., UE) side may be determined by the user device 102. In such situations, some information exchange may be performed between the user device 102 and the base station 104 to align their understanding.
[0102] Additionally, in some implementations, the base station 104 may send a command to the user device 102, where the command indicates to the user device 102 to activate one or more models. In response to receiving the command, the user device 102 may assess the inference time of the to be activated model (s) and report the inference time to the base station 104. In other words, the user device 102 may report the inference time of the one or more models that are to be activated. In some of these implementations, the inference time reported by the user device 102 may be valid at least for a time duration Tvalid. In particular of these implementations, the time duration Tvalid may be indicated by the user device 102 or may be configured by the base station 104.
[0103] Additionally, in some implementations, after receiving the report from the user device 102, the base station 104 may confirm whether to activate the one or multiple models based on the inference time. If the base station 104 confirms to activate the one or more models (or a subset of these models) , the user device 102 may start activating the corresponding models. However, if the base station 104 confirms not to activate the one or more models (or a subset of these models) , the user device 102 may stop activating the corresponding models.
[0104] In event that the user device 102 has already activated some models previously, then in response to receiving the command to activate the one or multiple models, the user device 102 may report the inference time of the previously activated model and the inference time of the one or multiple models that are to be activated to the base station 104. This is because if the user device 102 activates the one or multiple models indicated by the command, the inference time of the previously activated models may be impacted.
[0105] Additionally, in some implementations, the base station 104 may send a command to the user device 102, where the command indicates to the user device 102 to activate one or more models, and the command indicates an expected inference time of the one or more models. In some of these implementations, in response to receiving the command, the user device 102 may assess the inference time of the to-be activated model (s) and report to the base station 104 whether the expected inference time can be guaranteed or not. In particular of these implementations where the user device 102 sends an assessment report to the base station 104, the assessment report includes at least one of the following: the expected inference time can be satisfied; or the expected inference time can’t be satisfied. In even that the expected inference time can be satisfied, the user device 102 may report to the base station 104 that the model is being activated. Additionally, in event that the expected inference time cannot be satisfied, the user device 102 may report to the base station 104 that the model is not activated, or may report to the base station 104 that the model is deactivated.
[0106] Additionally, in some implementations, the base station 104 may configure or indicate a reporting period to the user device 102. In response, the user device 102 may report the inference time to the base station 104 following or according to the period. Because the inference time may be impacted by various factors, e.g., overheating, number of activated models, memory resource shortage and etc, it may be beneficial for the user device 102 to report the inference time of the activated models periodically. In addition or alternatively, in some implementations, the inference time reported by the user device 102 may be valid at least for a valid time duration Tvalid. In any of various of these implementations, the valid time duration Tvalid may be indicated by the user device 102 or may be the same as the period configured by the base station 104. In addition or alternatively, in some implementations, the user device 102 may report the inference time together with the valid time duration Tvalid to the base station following the period.
[0107] In some implementations, each of the models may be in different states. That is, at a given point in time, each model may be one of a plurality of states, and may be in different states at different points in time. The models being in different states at different times may be implemented to minimize computing power consumption. In particular of these implementations, a given model may be in the following three states.
[0108] A first state is an activated state. In the activated state, a user device 102 may perform model inference and model monitoring for the models. In some implementations, in the activated state, the user device may perform model inference for an activated state inference time period and may perform model monitoring for an activated state model monitoring period
[0109] A second state is a deactivated state. In the deactivated state, the user device 102 may not perform model inference. Additionally, in some implementations, in the deactivated state, the user device 102 may not perform model monitoring. In other implementations, in the deactivated state, the user device 102 may perform model monitoring for a deactivated state model monitoring period For at least some of these implementations, the deactivated state model monitoring period is longer or larger than the activated state model monitoring period In other of these implementations, the deactivated state monitoring period is shorter or smaller than the activated state model monitoring period
[0110] A third state is an inactive state. In some implementations, in the inactive state, the user device 102 may not perform model inference. In other implementations, in the inactive state, the user device 102 may perform model inference. In these other implementations, in the inactive state, the user device 102 may perform model inference for an inactive state model inference period In some of these other implementations, the inactive state model inference period is longer or larger than the activated state model inference period i.e., In other of these other implementations, the inactive state model inference period is shorter or smaller than the activated state model inference period i.e., Additionally, in the inactive state, the user device may perform model monitoring, such as for an inactive state model monitoring period In some implementations, the inactive state model monitoring period is longer or larger than the activated state model monitoring period i.e., In other implementations, the inactive state model monitoring period is smaller or shorter than the activated model monitoring period i.e.,
[0111] Additionally, in some implementations, the base station 104 may configure or indicate at least one model pattern to the user device 102. In some of these implementations, a model pattern includes a pattern period and an active time duration, where the active time duration is smaller than the pattern period. The active time duration may start from the beginning of each pattern period or start from an offset relative to the beginning of each pattern period. In some of these implementations, the offset is also configured by the base station 104. During the active time duration, the model is in the activated state. During times outside of the active time duration, the model is in inactive state. A time duration outside of the active time duration may be referred to as an inactive time duration.
[0112] Fig. 5 is a timing diagram, illustrating an example of a model pattern. In the example in Fig. 5, the base station 104 may configure the pattern period, active duration time, and offset to be 100 ms, 50 ms and 10 ms, respectively. In this case, the active duration time in the first period may start from the 10 ms offset, and correspondingly may end at the 60 ms. The other time durations inactive time durations, including one extending from 0 ms to 10 ms, and another extending from 60 ms to 100 ms. The same pattern will be repeated in the subsequent pattern periods.
[0113] In addition or alternatively, in some implementations, a timer may be configured by the base station 104 to control the model. In some of these implementations, the timer may be implemented with a timer restart condition. When the timer restart condition is satisfied, the user device 102 may restart the timer. In particular of these implementations, the timer restart condition may include at least one of the following: the user device 102 receives a command to activate the model; the user device 102 receives a command to perform model monitoring; the user device 102 reports the model monitoring result; the user device 102 reports the inference time; the user device 102 reports the model activation report; or the user device 102 reports the inference result.
[0114] In some implementations, if the time restart condition is not satisfied during one time unit, the user device 102 may decrement the timer at the end of the time unit. The time unit may be one frame, one subframe, one slot, one sub-slot, one second, one millisecond, or one period, as non-limiting examples in any of various implementations. In addition or alternatively, the time unit may be configured by the base station 104 in any of various implementations.
[0115] Additionally, in some implementations, when the timer expires, the user device 102 may switch the model to inactive state. In other implementations, when the timer expires, the user device 102 may switch the model to the deactivated state.
[0116] Additionally, in some implementations, the base station 104 may send a command to the user device 102 that instructs the user device 102 to switch one or more models between two of the states and / or for one or more models to be in a particular one of the states. For example, the command may indicate to the user device 102 to switch a model to the deactivated state, the inactivate state, or the activated state. Fig. 6 shows a schematic diagram of an overall state switching of the model states.
[0117] In addition or alternatively, in some implementations, a user device 102 may indicate two sets of models to the base station 104, including a first set and a second set. Each set may include one or more models. The first set may be available at the user device 102. The second set may be available at a server (e.g., an over-the-top (OTT) server that the user device 102 can access and / or with which the user device 102 can communicate. To use the model (s) in the second set, the user device 102 may access the server to download the model (s) . In contrast, for a model in the first set, the user device 102 already has the model (e.g., has already downloaded the model) , and in turn, the user device 102 does not have to access a server to obtain the model. Correspondingly, to use a model in the second set, the user device 102 may require additional time to download the models the before the user device 102 can activate it.
[0118] In addition or alternatively, in some other implementations, a user device 12 may indicates two sets of models to the base station 104, including a first set and a second set. Each set may include one or more models. In some of these implementations, for a model in the first set, the user device 102 may incur a first time delay to activate the model. For a model in the second set, the user device 102 may incur the first time delay plus a second time delay to activate the model. In particular of these implementations, the first time delay is for the user device 102 to prepare for the model activation. The second time delay is for the user device 102 to download the model (s) .
[0119] In addition or alternatively, in some implementations, the base station 104 may indicate two sets of models to the user device 102, including a first set and a second set. Each set may include one or more models. The first set may be configured by the base station 104. The second set may be indicated by downlink control information (DCI) or a medium access control (MAC) -control element (CE) . In addition or alternatively, the models of the second set are selected from the first set of models. The base station 104 may activate one or more models from the second set for the user device 102, and / or indicate to the user device 102 which of the model (s) of the second set are to be activated.
[0120] In addition or alternatively, in some implementations, the first set of models are in the deactivated state, and the second set of models are in the inactive state. In some of these implementations, the base station 104 may activate one or more models from the second set, such as by switching those one or more models of the second set from the inactive state to the activated state.
[0121] In some implementations or situations, a user device 102 may not have enough computing power to perform one or more operations associated with a model, such as model training for example. In some of these implementations, the user device 102 may request computing power from the base station 104 to perform the one or more operations. However, to perform the one or more operations, the user device 102 may first acquire a data set for the one or more operations. Using model training as an example, the user device 102 may first acquire the dataset for model training.
[0122] Additionally, in some implementations, the user device 102 may send a dataset request to the base station 104. The dataset request may include one or more of the following information: a request identification (ID); a model input and output requirement; a dataset resolution requirement; or dataset size requirements
[0123] The request ID may be an index. When the base station 104 sends back the dataset information, the base station 104 may also send the same request ID to the user device 102. In turn, the user device 102 may understand that the dataset information is for the previous request. In some implementations, the request ID may also include the vendor ID of the user device 102.
[0124] The model input and output requirement may include the dataset dimension information of the input data and output data. In addition or alternatively, the model input and output requirement may also include the data quantization requirement. For example, the user device 102 may require the input dimension as 12x14, which represents the modulation symbol in one resource block (RB) (e.g., 12 resource elements (REs) ) in one slot (e.g., 14 OFDM symbols) . The quantization requirement may include linear or vector quantization.
[0125] The dataset resolution requirement may include an expected resolution, such as 1 sample per 1 ms or that the figure is accurate to two decimal places.
[0126] The data size requirement may include the total size of the dataset, or the total number of samples of the dataset.
[0127] In some implementations, in response to receipt of a dataset request, the base station 104 may send back a dataset request response to the user device 102. Additionally, in response to the dataset request, the base station 104 may identify a dataset that satisfies the requirements in the request. If the base station 104 identifies a dataset that satisfies the requirements, the dataset request response may include at least one of the following information of the dataset: a request identification (ID); dataset information; a dataset address; or dataset ID. The dataset information may include at least one of: a dataset size, a dimension, a resolution, input information, or output information, in some implementations. Additionally, the dataset address may indicate where the dataset is saved or stored in memory. In any of various implementations, the user device 102 and / or the base station 104 may access the dataset via this dataset address. Additionally, the dataset ID is the identification or index of the dataset.
[0128] In some implementations, in event that the base station 104 does not identify any dataset that satisfies the requirements, the dataset request response includes the request ID and error feedback. The error feedback may include an error cause, the error cause being that the base station could not identify any dataset satisfying the requirements of the request.
[0129] Additionally, in response to the dataset request response from the base station 104, the user device 102 may send a report to the base station 104 to indicate that the dataset request response has been received and / or to indicate that the dataset information has been received.
[0130] Additionally, in some implementations, the user device 102 and the base station 104 may perform information exchange for model training. In some of these implementations, the user device 102 may send a mode training request to the base station 104, where the request includes at least one of the following information: a request ID; a dataset ID; one or more model requirements; or a computing power requirement.
[0131] In some implementations, the request ID is an index. When the base station 104 sends back the model training information, the base station 104 may also send the same request ID to the user device 102. In turn, the user device 102 may understand that the model training information is for the previous request. In some implementations, the request ID may also include a vendor ID of the user device 102. The dataset ID may identify the dataset used for the model training. Model requirements may include at least one of: a model structure, a number of layers, or applicable conditions, as non-limiting examples. The applicable conditions may include applicable area information, applicable frequency / band / carrier information, and / or applicable time information, as non-limiting examples. The computing power requirement may include an expected computing power of the model and / or a maximum computing power of the model. In some implementations, the computing power may be represented by Floating Point Operations Per Second (FLOPs) .
[0132] Additionally, in some implementations, the base station 104 may send back a model training response to the user device 102. In event that the base station 104 successfully trains a model that satisfies the requirements, the model training response includes at least one of the following information: a request ID, a model ID, model information, or computing power information. The model ID may identify the model. The model information may include the model itself or include an address where the model is stored in memory, which in turn may allow the user device 102 to download the model from the memory. The computing power information may include the required computing power for the model to perform model inference. On the other hand, if the base station 104 does not train the model that satisfies the requirements, the model training response may include the request ID and error feedback. The error feedback may include the error cause, i.e., that the base station 104 could not train a model that satisfies the requirements.
[0133] In response to the model training response from base station 104, the user device 102 may send a report to the base station 104 to indicate that the model training response has been received and / or to indicate that the model information has been received.
[0134] The description and accompanying drawings above provide specific example embodiments and implementations. The described subject matter may, however, be embodied in a variety of different forms and, therefore, covered or claimed subject matter is intended to be construed as not being limited to any example embodiments set forth herein. A reasonably broad scope for claimed or covered subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices, components, systems, or non-transitory computer-readable media for storing computer codes. Accordingly, embodiments may, for example, take the form of hardware, software, firmware, storage media or any combination thereof. For example, the method embodiments described above may be implemented by components, devices, or systems including memory and processors by executing computer codes stored in the memory.
[0135] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment / implementation” as used herein does not necessarily refer to the same embodiment and the phrase “in another embodiment / implementation” as used herein does not necessarily refer to a different embodiment. It is intended, for example, that claimed subject matter includes combinations of example embodiments in whole or in part.
[0136] In general, terminology may be understood at least in part from usage in context. For example, terms, such as “and”, “or”, or “and / or,” as used herein may include a variety of meanings that may depend at least in part on the context in which such terms are used. Typically, “or” if used to associate a list, such as A, B or C, is intended to mean A, B, and C, here used in the inclusive sense, as well as A, B or C, here used in the exclusive sense. In addition, the term “one or more” as used herein, depending at least in part upon context, may be used to describe any feature, structure, or characteristic in a singular sense or may be used to describe combinations of features, structures or characteristics in a plural sense. Similarly, terms, such as “a,” “an,” or “the,” may be understood to convey a singular usage or to convey a plural usage, depending at least in part upon context. In addition, the term “based on” may be understood as not necessarily intended to convey an exclusive set of factors and may, instead, allow for existence of additional factors not necessarily expressly described, again, depending at least in part on context.
[0137] Reference throughout this specification to features, advantages, or similar language does not imply that all of the features and advantages that may be realized with the present solution should be or are included in any single implementation thereof. Rather, language referring to the features and advantages is understood to mean that a specific feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of the present solution. Thus, discussions of the features and advantages, and similar language, throughout the specification may, but do not necessarily, refer to the same embodiment.
[0138] Furthermore, the described features, advantages and characteristics of the present solution may be combined in any suitable manner in one or more embodiments. One of ordinary skill in the relevant art will recognize, in light of the description herein, that the present solution can be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments of the present solution.
[0139] The subject matter of the disclosure may also relate to or include, among others, the following aspects:
[0140] A first aspect includes a method for wireless communication that includes: transmitting, by a user device, a nominal inference time of an M number of models to a base station, wherein M is an integer of two or more; receiving, by the user device, an indication of a subset models from among the M number of models activated by the base station, wherein the subset comprises an N number of models activated by the base station, wherein N is an integer of one or more, and N is less than M; determining, by the user device an actual inference time of the subset based on the subset of models that are activated and the nominal inference time.
[0141] A second aspect includes a method for wireless communication that includes: receiving, by a base station, a nominal inference time of an M number of models from a user device, wherein M is an integer of two or more; activating, by the base station, a subset of models from among the M number of models, wherein the subset comprises an N number of models activated by the base station, wherein N is an integer of one or more, and N is less than M; and transmitting, by the base station, an indication of the subset of models that is activated to a user device.
[0142] A third aspect includes any of first or second aspects, and further includes wherein the actual inference time is determined based on the nominal inference time and the N number of models of the subset activated by the base station.
[0143] A fourth aspect includes any of the first or second aspects, and further includes wherein the actual inference time is determined based on the nominal inference time, the N number of models of the subset activated by the base station, and an additional time unit.
[0144] A fifth aspect includes the fourth aspect, and further includes wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to:
[0145] where Ti is a nominal inference time of the i-th model, N is the number of models of the subset, TΔ is the additional time unit, and i is an integer where 1≤i≤N.
[0146] A sixth aspect includes any of the first or second aspects, and further includes wherein the actual inference is determined based on the nominal inference time, the number of models of the subset, a total artificial intelligence processing unit (APU) for the user device, and an APU requirement of each model of the subset.
[0147] A seventh aspect includes the sixth aspect, and further includes wherein the actual interference time comprises an i-th actual inference time of an i-th model, and the i-th actual inference time is determined according to:
[0148] or
[0149] where Ti is a nominal inference time of the i-th model, Atotal is the total APU for the user device, Ai is an APU requirement of the i-th model, is a total APU requirement of the subset of models, and i is an integer where 1≤i≤N.
[0150] An eighth aspect includes any of the sixth or seventh aspect, and further includes wherein the total APU for the user device and the APU requirement for the i-th model are reported by the user device, indicated by the base station, or predefined.
[0151] A ninth aspect includes any of the first or second aspects, and further includes wherein the actual inference time is determined based on the nominal inference time, the N number of models activated by the base station, and a nominal number of activated models.
[0152] A tenth aspect includes the ninth aspect, and further includes wherein the actual inference time is determined further based on an additional time unit.
[0153] An eleventh aspect includes the tenth aspect, and further includes wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to:
[0154] where Ti is a nominal inference time of the i-th model, N is the number of models of the subset, TΔ is the additional time unit, i is an integer where 1≤i≤N, and P is the nominal number of activated models.
[0155] A twelfth aspect includes any of the first or second aspects, and further includes wherein the actual inference time is determined based on the nominal inference time and a number of layers.
[0156] A thirteenth aspect includes the twelfth aspect, and further includes wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to:
[0157] where Ti is a nominal inference time of the i-th model, L is the number of layers, i is an integer where 1≤i≤N.
[0158] A fourteenth aspect includes the twelfth aspect, and further includes wherein the actual inference time is determined further based a maximum number of layers.
[0159] A fifteenth aspect includes the fourteenth aspect, and further includes wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to:
[0160] where Ti is a nominal inference time of the i-th model, L is the number of layers, Lmax is the maximum number of layers, and i is an integer where 1≤i≤N.
[0161] A sixteenth aspect includes any of the first or second aspects, and further includes wherein the actual inference time is determined based on the nominal inference time and a number of time instances of a model input.
[0162] A seventeenth aspect includes the sixteenth aspect, and further includes wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to:
[0163] where Ti is a nominal inference time of the i-th model, I is the number of time instances, and i is an integer where 1≤i≤N.
[0164] An eighteenth aspect includes the sixteenth aspect, and further includes wherein the actual inference time is determined further based on a maximum number of time instances of a model input.
[0165] A nineteenth aspect includes the eighteenth aspect, and further includes wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to:
[0166] where Ti is a nominal inference time of the i-th model, I is the number of time instances, Imax is the maximum number of time instances, and i is an integer where 1≤i≤N.
[0167] A twentieth aspect includes any of the first through nineteenth aspects, and further includes: receiving, by the user device, a command from the base station, wherein the command indicates to activate the subset of models; and sending, by the user device, a report of the actual inference time of the subset.
[0168] A twenty-first aspect includes any of the first through twentieth aspects, and further includes: transmitting, by the base station, a command to the user device, wherein the command indicates to activate the subset of models; and receiving, by the base station, a report of the actual inference time of the subset.
[0169] A twenty-second aspect includes any of the twentieth or twenty-first aspects, and further includes wherein the actual inference time is valid for a valid time duration Tvalid, wherein the valid time duration Tvalid is indicated by the user device or configured by the base station.
[0170] A twenty-third aspect includes any of the twentieth through twenty-second aspects, and further includes wherein the command further indicates an expected inference time of the subset of models.
[0171] A twenty-fourth aspect includes the twenty-third aspect, and further includes: transmitting, by the user device, an assessment report to the base station, the assessment report indicating whether or not the expected inference time can be satisfied.
[0172] A twenty-fifth aspect includes any of the twenty-third or twenty-fourth aspects, and further includes: receiving, by the base station an assessment report from the user device, the assessment report indicating whether or not the expected inference time can be satisfied.
[0173] A twenty-sixth aspect includes any of the twentieth through twenty-fifth aspects, and further includes wherein the user device sends the report of the actual inference time of the subset according to a reporting period indicated by the base station.
[0174] A twenty-seventh aspect includes any of the twentieth through twenty-sixth aspects, and further includes wherein the base station indicates, to the user device, a reporting period according to which the user device is to report the actual inference time.
[0175] A twenty-eighth aspect includes any of the twenty-sixth or twenty-seventh aspects, and further includes wherein the reporting period is equal to a valid time duration Tvalid that the actual inference time is valid.
[0176] A twenty-ninth aspect includes the twenty-eighth aspect, and further includes: transmitting, by the user device, the valid time duration with the actual inference time according to the reporting period.
[0177] A thirtieth aspect includes any of the twenty-eighth or twenty-ninth aspects, and further includes: receiving, by the base station, the valid time duration with the actual inference time according to the reporting period.
[0178] A thirty-first aspect includes a method for wireless communication that includes: determining, by a user device, a state of a plurality of states for a model; and determining, by the user device, whether or not to perform at least one operation associated with the model according to the state, wherein the at least one operation comprises at least one of: model inference or model monitoring.
[0179] A thirty-second aspect includes the thirty-first aspect, and further includes wherein the plurality of states comprises: an activated state in which the user device performs model inference for an activated state inference period and model monitoring for an activated state monitoring period a deactivated state in which the user device does not perform model inference and does not perform model monitoring, or in which the user device performs model monitoring for a deactivated state monitoring period that is longer than the activated state monitoring period and an inactive state in which the user device: does not perform model inference and performs model monitoring, where the model monitoring is performed for an inactive state monitoring period that is longer or shorter than the activated state monitoring period or performs model inference and performs model monitoring, where the model inference is performed for an inactive state inference period that is longer than the activated state inference period
[0180] A thirty-third aspect includes any of the thirty-first or thirty-second aspects, and further includes: receiving, by the user device, a model pattern, the model pattern comprising a pattern period and an active time duration, wherein the activate time duration is smaller than the pattern period.
[0181] A thirty-fourth aspect includes the thirty-third aspect, and further includes wherein the user device receives the model pattern from the base station.
[0182] A thirty-fifth aspect includes any of the thirty-third or thirty-fourth aspects, and further includes wherein the active time duration starts from a beginning of each pattern period or from an offset relative to the beginning of each pattern period.
[0183] A thirty-sixth aspect includes any of the thirty-third or thirty-fourth aspects, and further includes wherein the model is in an activated state during the active time duration, and the model is in an inactive state during an inactive time duration.
[0184] A thirty-seventh aspect includes any of the thirty-first through thirty-sixth aspects, and further includes wherein a timer configured by the base station controls the model, wherein the user device restarts the timer in response to a timer restart condition being satisfied.
[0185] A thirty-eighth aspect includes the thirty-seventh aspect, and further includes wherein the timer restart condition comprises at least one of: the user device receives a command to activate the model; the user device receives a command to perform model monitoring; the user device reports a model monitoring result; the user device reports an inference time; the user device reports a model activation report; or the user device reports an inference result.
[0186] A thirty-ninth aspect includes any of the thirty-seventh or thirty-eighth aspects, and further includes: in response to the timer restart condition not being satisfied during a single time unit, the user device decrements the timer at an end of the single time unit, wherein the single time unit comprises: one frame, one subframe, one slot, one sub-slot, one second, one millisecond, or one period.
[0187] A fortieth aspect includes any of the thirty-seventh through thirty-ninth aspects, and further includes wherein the user device switches the model to an inactivate state when the timer expires.
[0188] A forty-first aspect includes any of the thirty-first through fortieth aspects, and further includes wherein the user device indicates a first set of one or more models and a second set of one or more models to the base station.
[0189] A forty-second aspect includes the forty-first aspect, and further includes wherein the set is available at the user device, and the second set is available at a server of the user device.
[0190] A forty-third aspect includes any of the forty-first or forty-second aspects, and further includes wherein the first set requires a first time delay for the user device to activate the one or more models of the first set, and the second set requires the first time delay plus a second time delay for the user device to activate the one or more models of the second set, wherein the first time delay is for the user device to prepare for model activation, and the second time delay is for the user device to download the one or more models of the second set.
[0191] A forty-fourth aspect includes any of the thirty-first through fortieth aspects, and further includes wherein the user device receives an indication of a first set of one or more models and a second set of one or more models from the base station, wherein at least one of: the first set is configured by the base station; the second set is indicated by downlink control information (DCI) or a medium access control (MAC) -control element (CE); the second set is selected from the first set of models; or the base station indicates to the user device that at least one of the one or more models of the second set is activated.
[0192] A forty-fifth aspect includes a wireless communications apparatus that includes at least one processor and a memory, wherein the at least one processor is configured to cause the apparatus to perform any of the first through forty-fourth aspects.
[0193] A forty-sixth aspect includes a computer program product that includes a computer-readable program medium comprising code stored thereupon, the code, when executed by a processor, causing the processor to implement any of the first through forty-fourth aspects.
[0194] In addition to the features mentioned in each of the independent aspects enumerated above, some examples may show, alone or in combination, the optional features mentioned in the dependent aspects and / or as disclosed in the description above and shown in the figures.
Claims
1.A method for wireless communication, the method comprising:transmitting, by a user device, a nominal inference time of an M number of models to a base station, wherein M is an integer of two or more;receiving, by the user device, an indication of a subset models from among the M number of models activated by the base station, wherein the subset comprises an N number of models activated by the base station, wherein N is an integer of one or more, and N is less than M;determining, by the user device an actual inference time of the subset based on the subset of models that are activated and the nominal inference time.2.A method for wireless communication, the method comprising:receiving, by a base station, a nominal inference time of an M number of models from a user device, wherein M is an integer of two or more;activating, by the base station, a subset of models from among the M number of models, wherein the subset comprises an N number of models activated by the base station, wherein N is an integer of one or more, and N is less than M; andtransmitting, by the base station, an indication of the subset of models that is activated to a user device.3.The method of claim 1, wherein the actual inference time is determined based on the nominal inference time and the N number of models of the subset activated by the base station.4.The method of claim 1, wherein the actual inference time is determined based on the nominal inference time, the N number of models of the subset activated by the base station, and an additional time unit.5.The method of claim 4, wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to: where Ti is a nominal inference time of the i-th model, N is the number of models of the subset, TΔ is the additional time unit, and i is an integer where 1≤i≤N.6.The method of claim 1, wherein the actual inference is determined based on the nominal inference time, the number of models of the subset, a total artificial intelligence processing unit (APU) for the user device, and an APU requirement of each model of the subset.7.The method of claim 6, wherein the actual interference time comprises an i-th actual inference time of an i-th model, and the i-th actual inference time is determined according to: or where Ti is a nominal inference time of the i-th model, Atotal is the total APU for the user device, Ai is an APU requirement of the i-th model, is a total APU requirement of the subset of models, and i is an integer where 1≤i≤N.8.The method of claim 6, wherein the total APU for the user device and the APU requirement for the i-th model are reported by the user device, indicated by the base station, or predefined.9.The method of claim 1, wherein the actual inference time is determined based on the nominal inference time, the N number of models activated by the base station, and a nominal number of activated models.10.The method of claim 9, wherein the actual inference time is determined further based on an additional time unit.11.The method of claim 10, wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to: where Ti is a nominal inference time of the i-th model, N is the number of models of the subset, TΔ is the additional time unit, i is an integer where 1≤i≤N, and P is the nominal number of activated models.12.The method of claim 1, wherein the actual inference time is determined based on the nominal inference time and a number of layers.13.The method of claim 12, wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to: where Ti is a nominal inference time of the i-th model, L is the number of layers, i is an integer where 1≤i≤N.14.The method of claim 12, wherein the actual inference time is determined further based a maximum number of layers.15.The method of claim 14, wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to: where Ti is a nominal inference time of the i-th model, L is the number of layers, Lmax is the maximum number of layers, and i is an integer where 1≤i≤N.16.The method of claim 1, wherein the actual inference time is determined based on the nominal inference time and a number of time instances of a model input.17.The method of claim 16, wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to: where Ti is a nominal inference time of the i-th model, I is the number of time instances, and i is an integer where 1≤i≤N.18.The method of claim 16, wherein the actual inference time is determined further based on a maximum number of time instances of a model input.19.The method of claim 18, wherein the actual inference time comprises an i-th actual inference time of an i-th model, wherein the i-th actual inference time is determined according to: where Ti is a nominal inference time of the i-th model, I is the number of time instances, Imax is the maximum number of time instances, and i is an integer where 1≤i≤N.20.The method of claim 1, further comprising:receiving, by the user device, a command from the base station, wherein the command indicates to activate the subset of models; andsending, by the user device, a report of the actual inference time of the subset.21.The method of claim 2, further comprising:transmitting, by the base station, a command to the user device, wherein the command indicates to activate the subset of models; andreceiving, by the base station, a report of the actual inference time of the subset.22.The method of any of claims 20 or 21, wherein the actual inference time is valid for a valid time duration Tvalid, wherein the valid time duration Tvalid is indicated by the user device or configured by the base station.23.The method of any of claims 20 or 21, wherein the command further indicates an expected inference time of the subset of models.24.The method of claim 23, further comprising:transmitting, by the user device, an assessment report to the base station, the assessment report indicating whether or not the expected inference time can be satisfied.25.The method of claim 23, further comprising:receiving, by the base station an assessment report from the user device, the assessment report indicating whether or not the expected inference time can be satisfied.26.The method of claim 20, wherein the user device sends the report of the actual inference time of the subset according to a reporting period indicated by the base station.27.The method of claim 21, wherein the base station indicates, to the user device, a reporting period according to which the user device is to report the actual inference time.28.The method of any of claims 26 or 27, wherein the reporting period is equal to a valid time duration Tvalid that the actual inference time is valid.29.The method of claim 28, further comprising:transmitting, by the user device, the valid time duration with the actual inference time according to the reporting period.30.The method of claim 28, further comprising:receiving, by the base station, the valid time duration with the actual inference time according to the reporting period.31.A method for wireless communication, the method comprising:determining, by a user device, a state of a plurality of states for a model; anddetermining, by the user device, whether or not to perform at least one operation associated with the model according to the state, wherein the at least one operation comprises at least one of: model inference or model monitoring.32.The method of claim 31, wherein the plurality of states comprises:an activated state in which the user device performs model inference for an activated state inference periodand model monitoring for an activated state monitoring perioda deactivated state in which the user device does not perform model inference and does not perform model monitoring, or in which the user device performs model monitoring for a deactivated state monitoring periodthat is longer than the activated state monitoring periodandan inactive state in which the user device:does not perform model inference and performs model monitoring, where the model monitoring is performed for an inactive state monitoring periodthat is longer or shorter than the activated state monitoring periodorperforms model inference and performs model monitoring, where the model inference is performed for an inactive state inference periodthat is longer than the activated state inference period33.The method of claim 31, further comprising:receiving, by the user device, a model pattern, the model pattern comprising a pattern period and an active time duration, wherein the activate time duration is smaller than the pattern period.34.The method of claim 33, wherein the user device receives the model pattern from the base station.35.The method of any of claims 33 or 34, wherein the active time duration starts from a beginning of each pattern period or from an offset relative to the beginning of each pattern period.36.The method of any of claims 33 or 34, wherein the model is in an activated state during the active time duration, and the model is in an inactive state during an inactive time duration.37.The method of claim 31, wherein a timer configured by the base station controls the model, wherein the user device restarts the timer in response to a timer restart condition being satisfied.38.The method of claim 37, wherein the timer restart condition comprises at least one of:the user device receives a command to activate the model;the user device receives a command to perform model monitoring;the user device reports a model monitoring result;the user device reports an inference time;the user device reports a model activation report; orthe user device reports an inference result.39.The method of claim 37, in response to the timer restart condition not being satisfied during a single time unit, the user device decrements the timer at an end of the single time unit, wherein the single time unit comprises: one frame, one subframe, one slot, one sub-slot, one second, one millisecond, or one period.40.The method of claim 37, wherein the user device switches the model to an inactivate state when the timer expires.41.The method of claim 31, wherein the user device indicates a first set of one or more models and a second set of one or more models to the base station.42.The method of claim 41, wherein the set is available at the user device, and the second set is available at a server of the user device.43.The method of claim 41, wherein the first set requires a first time delay for the user device to activate the one or more models of the first set, and the second set requires the first time delay plus a second time delay for the user device to activate the one or more models of the second set, wherein the first time delay is for the user device to prepare for model activation, and the second time delay is for the user device to download the one or more models of the second set.44.The method of claim 31, wherein the user device receives an indication of a first set of one or more models and a second set of one or more models from the base station, wherein at least one of:the first set is configured by the base station;the second set is indicated by downlink control information (DCI) or a medium access control (MAC) -control element (CE) ;the second set is selected from the first set of models; orthe base station indicates to the user device that at least one of the one or more models of the second set is activated.45.A wireless communications apparatus comprising at least one processor and a memory, wherein the at least one processor is configured to cause the apparatus to perform a method of any of claims 1 to 44.46.A computer program product comprising a computer-readable program medium comprising code stored thereupon, the code, when executed by a processor, causing the processor to implement a method of any of claims 1 to 44.
Citation Information
Patent Citations
Multi-version inference model deployment method, device and system in edge computing environment
CN111459505A
Slice by slice ai / ML model inference over communication networks
US20230275812A1
User equipment machine learning service continuity
US20230422117A1
Methods, devices, and medium for communication
WO2023092349A1
Method and apparatus for presenting ai and ML media services in wireless communication system
WO2023214809A1