Apparatuses, devices and methods for error-tolerant inference in wireless systems

Error-tolerant inference methods with dynamic I-QoS frameworks address latency and overhead issues in wireless networks, enhancing power efficiency and reducing operational costs.

WO2026157019A1PCT designated stage Publication Date: 2026-07-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-03-28
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing wireless networks, such as 5G NR, introduce latency and energy consumption due to multiple retransmissions and robust error corrections, particularly in AI inference tasks, leading to excessive network overheads.

Method used

Implementing error-tolerant inference methods with controlled error tolerance, using a novel I-QoS framework for dynamic adaptation of transmission parameters, allowing flexible error-handling strategies, and optimizing inference pipelines to reduce latency and overhead while maintaining quality.

Benefits of technology

Reduces latency, energy consumption, and network overheads in wireless systems by enabling error-tolerant inference with improved power usage effectiveness and simplified power infrastructure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025085664_30072026_PF_FP_ABST
    Figure CN2025085664_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods, apparatuses, devices, and systems for error-tolerant inference in a wireless system to improve power usage effectiveness, simplify power infrastructure, and / or reduce overall operational cost. A network device coordinating an inference task may transmit a request for the inference task. Each computing device may receive an input for at least a portion of an inference task, in accordance with first information. The first information may configure one or more parameters associated with inference quality of the interference task. Each computing device may perform the at least a portion of the inference task in accordance with the first information, and transmit an output of the at least a portion of the inference task in accordance with the first information. The network device may receive a final output of the inference task in accordance with the first information.
Need to check novelty before this filing date? Find Prior Art

Description

APPARATUSES, DEVICES AND METHODS FOR ERROR-TOLERANT INFERENCE IN WIRELESS SYSTEMSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present disclosure claims priority to and the benefit of U.S. Provisional Application No. 63 / 749,808 filed in the U.S. Patent and Trademark Office on January 27, 2025, which is incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] The present disclosure relates generally to wireless communications, and in particular to apparatuses, devices and methods for error-tolerant inference in a wireless system.BACKGROUND

[0003] Some wireless networks, such as 5th generation (5G) new radio (NR) network, provide communication categories like enhanced mobile broadband (eMBB) and ultra-reliable low-latency communication (URLLC) , for example, to achieve highly reliable data delivery through multiple retransmissions and robust error corrections. However, such approaches may introduce additional latency and energy consumption, and trigger excessive network overheads. Such drawbacks are particularly presented when performing artificial intelligence (AI) inference tasks in distributed wireless network systems.

[0004] Therefore, there is a need for new apparatuses, devices and methods for inference in a wireless network system.SUMMARY

[0005] Aspects of the present disclosure provide methods, apparatuses, devices and systems to reduce latency, energy consumption, and network overheads, as well as specific methods, apparatuses, devices, and systems that enable error-tolerant operations in a wireless system, for example performing error-tolerant inference tasks while maintaining acceptable inference quality in a distributed wireless system.

[0006] According to an aspect of the disclosure, there is provided a method for error-tolerant inference in a wireless system. The method may include transmitting a request for an inference task; and receiving an output of the inference task in accordance with first information. The first information may configure one or more parameters associated with inference quality of the interference task.

[0007] The proposed method may enable an error-tolerant inference in a wireless system. For example, by virtue of some aspects of the present disclosure, wireless transmissions (e.g., transmission of intermediate inference data between computing devices) may be implemented with controlled level of error tolerance, instead of enforcing near-perfect reliability. The error-tolerant system ensures that the overhead of wireless communication demands (which may be caused by deployments in wireless communication systems for inference) stays below the overhead of centralized cooling and special power supply demands in an alternative centralized datacenter, thereby, for example, yielding improved power usage effectiveness, simplifying power infrastructure, and / or reducing overall operational cost. In another example, by virtue of some aspects of the present disclosure, a novel inference-related quality of service (I-QoS) framework may be introduced, which may incorporate inference-quality metrics and dynamic policies, to continuously optimize transmission parameters (e.g., hybrid automatic repeat request (HARQ) , modulation and coding scheme (MCS) ) and / or resource allocation. Accordingly, the network may dynamically and flexibly adapt error tolerance and link configuration in real time (or in near real time) and ensure that inference pipelines operate at minimal overhead while maintaining acceptable quality. In another example, by virtue of some aspects of the present disclosure, the proposed method may enable dynamic and / or flexible error-handling strategy adjustment (e.g., selectively allowing certain errors) , thereby reducing delay for time-critical inference tasks and / or providing a standardized way to incorporate inference metrics into HARQ / MCS control at network layers.

[0008] In some implementations, the method may further include based on at least one of the metric used for the evaluation of the inference quality of the inference task, a topology of computing devices performing at least part of the inference task, or a number of the computing devices performing at least part of the inference task, determining at least one of: the transmission error rate, or the throughput or latency range.

[0009] In some implementations, the method may further include transmitting, to each computing device, the first information.

[0010] In some implementations, the first information is transmitted via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) .

[0011] In some implementations, transmitting the request for the inference task includes transmitting, to each computing device, a respective request for the inference task.

[0012] In some implementations, each computing device is included in a respective inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task.

[0013] In some implementations, more than one of the plurality of computing devices are included in a same inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task.

[0014] In some implementations, for each computing device, at least one of: the computing device receives an input for the respective portion of the inference task in accordance with the first information; and the computing device transmits an output of respective portion of the inference task in accordance with the first information.

[0015] In some implementations, the plurality of computing devices include one or more user equipments (UEs) .

[0016] In some implementations, the method may further include receiving, from a requesting device, an inference task request including at least part of the first information; and transmitting, to the requesting device, the output of the inference task in accordance with the first information.

[0017] In some implementations, the request for the inference task is generated based on the received inference task request.

[0018] According to an aspect of the present disclosure, there is provided an apparatus including means to perform the method illustrated in this disclosure. The apparatus may be a network device or a module / chipset in the network device. For example, the apparatus includes a processor configured to cause the processor to perform a method consistent with the implementations described above and / or elsewhere in the present disclosure. In another example, the apparatus includes a processor coupled with a computer-readable medium. The computer-readable medium stores thereon computer executable instructions that when executed cause the processor or the apparatus to perform a method consistent with the implementations described above and / or elsewhere in the present disclosure. Non-limiting examples of the apparatus are a base station (BS) , a transmission and receive point (TRP) , a network node, a network function, and / or any other suitable network devices, nodes, or apparatuses. In some implementations, the apparatus includes a chip, e.g., an IC chip, a modem chip (also referred to as a baseband chip) , an SoC chip, and / or an SIP chip. In some implementations, the apparatus may include circuitry such as an FPGA, a GPU, or an ASIC, that performs the methods. More generally, the apparatus may include one or more units to perform a method as described above or elsewhere in the present disclosure. The term “units” is used in a broad sense and may be referred to by any of various names, including for example, modules, components, elements, means, etc. The units may be implemented using hardware, software, firmware or any combination thereof.

[0019] According to an aspect of the disclosure, there is provided a method for error-tolerant inference in a wireless system. The method may include receiving an input for at least a portion of an inference task in accordance with first information. The first information may configure one or more parameters associated with inference quality of the interference task. The method may further include performing the at least a portion of the inference task in accordance with the first information, and transmitting an output of the at least a portion of the inference task in accordance with the first information.

[0020] The benefit of the proposed method may refer to the benefit of previous aspect and implementations.

[0021] In some implementations, the method may further include receiving, from a network device coordinating the inference task, the first information.

[0022] In some implementations, the network device coordinating the inference task is a base station.

[0023] In some implementations, the first information is received via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) .

[0024] In some implementations, the method may further include receiving, from the network device coordinating the inference task, a request for the at least a portion of the inference task.

[0025] In some implementations, each computing device is included in a respective inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task.

[0026] In some implementations, more than one of the plurality of computing devices are included in a same inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task.

[0027] In some implementations, the plurality of computing devices include one or more user equipments (UEs) .

[0028] In some implementations, based on any previous aspects and implementations, the one or more parameters associated with inference quality of the interference task indicate at least one of: a transmission error rate to be satisfied during the inference task; a throughput or latency range to be satisfied during the inference task; a type of the inference task; a metric used for evaluation of the inference quality of the inference task; a metric threshold value used for the evaluation of the inference quality of the inference task; an inference quality evaluation interval for the inference task; priority of the inference task; one or more communication parameters used for transmission of data associated with the interference task; or one or more adjustment policies, for parameter adjustment or resource allocation, based on the evaluation of the inference quality of the inference task.

[0029] In some implementations, based on any previous aspects and implementations, at least one of the transmission error rate or the throughput or latency range is determined based on at least one of: the metric used for the evaluation of the inference quality of the inference task; a topology of computing devices performing at least part of the inference task; or a number of the computing devices performing at least part of the inference task.

[0030] In some implementations, based on any previous aspects and implementations, the one or more communication parameters include at least one of: at least one parameter associated with a hybrid automatic repeat request (HARQ) protocol, or at least one parameter associated with a modulation and coding scheme (MCS) .

[0031] In some implementations, based on any previous aspects and implementations, the at least one parameter associated with the HARQ protocol includes at least one of: a HARQ ignorance rate indicative of a portion that HARQ retransmission is deactivated; or an override ignorance flag indicating that the HARQ retransmission is fully activated regardless of the HARQ ignorance rate.

[0032] In some implementations, based on any previous aspects and implementations, the metric used for the evaluation of the inference quality of the inference task includes at least one of: a perplexity metric measuring uncertainty associated with the inference task; an accuracy metric based on accuracy associated with the inference task; a metric based on Δ-probability associated with the inference task; or a metric based on Kullback–Leibler (KL) divergence associated with the inference task.

[0033] In some implementations, based on any previous aspects and implementations, the inference quality of the inference task is evaluated using at least one of: a predefined input of the inference task; or a predefined error-free output of the inference task.

[0034] In some implementations, based on any previous aspects and implementations, the inference quality of the inference task is evaluated, and wherein the one or more parameters are adjusted and / or resources to the interference task are reallocated based on at least one of: the evaluated inference quality of the inference task, or the one or more adjustment policies.

[0035] In some implementations, based on any previous aspects and implementations, the one or more parameters that are adjusted include at least one of the transmission error rate, the throughput or latency range, or the one or more communication parameters.

[0036] In some implementations, based on any previous aspects and implementations, the at least one parameter associated with the MCS is adjusted based on at least one of: the transmission error rate, or the throughput or latency range.

[0037] In some implementations, based on any previous aspects and implementations, when the one or more parameters are adjusted, the method may further include receiving the adjusted one or more parameters via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) .

[0038] In some implementations, based on any previous aspects and implementations, the inference quality of the inference task is evaluated periodically based on the inference quality evaluation interval.

[0039] In some implementations, based on any previous aspects and implementations, the inference task is collectively performed by a plurality of computing devices, each computing device performing a respective portion of the inference task.

[0040] In some implementations, based on any previous aspects and implementations, performing the at least a portion of the inference task includes performing the respective portion of the inference task. According to an aspect of the present disclosure, there is provided a computing apparatus including means to perform the method illustrated in this disclosure. For example, the computing apparatus includes a processor configured to cause the processor to perform a method consistent with the implementations described above and / or elsewhere in the present disclosure. In another example, the computing apparatus includes a processor coupled with a computer-readable medium. The computer-readable medium stores thereon computer executable instructions that when executed cause the processor or the computing apparatus to perform a method consistent with the implementations described above and / or elsewhere in the present disclosure. Non-limiting examples of the computing apparatus are a user equipment (UE) , any other suitable terminal devices or modules / chips in the apparatus. In some implementations, the computing apparatus includes a chip, e.g., an IC chip, a modem chip (also referred to as a baseband chip) , an SoC chip, and / or an SIP chip. In some implementations, the computing apparatus may include circuitry such as an FPGA, a GPU, or an ASIC, that performs the methods. More generally, the computing apparatus may include one or more units to perform a method as described above or elsewhere in the present disclosure. The term “units” is used in a broad sense and may be referred to as any of various names, including for example, modules, components, elements, means, etc. The units may be implemented using hardware, software, firmware or any combination thereof.

[0041] According to an aspect of the disclosure, there is provided a method for error-tolerant inference in a wireless system. The method may include transmitting an inference task request including at least part of first information. The first information may configure one or more parameters associated with inference quality of an interference task. The method may further include receiving an output of the inference task in accordance with the first information.

[0042] The benefit of the proposed method may refer to the benefit of previous aspect and implementations.

[0043] In some implementations, the one or more parameters associated with inference quality of the interference task indicate at least one of: a throughput or latency range to be satisfied during the inference task; a type of the inference task; a metric used for evaluation of the inference quality of the inference task; a metric threshold value used for the evaluation of the inference quality of the inference task; an inference quality evaluation interval for the inference task; priority of the inference task; or one or more communication parameters used for transmission of data associated with the interference task.

[0044] In some implementations, the inference task request is transmitted to a network device coordinating the inference task, and the output of the inference task is received from the network device coordinating the inference task.

[0045] According to an aspect of the present disclosure, there is provided a requesting apparatus including means to perform the method illustrated in this disclosure. For example, the requesting apparatus includes a processor configured to cause the processor to perform a method consistent with the implementations described above and / or elsewhere in the present disclosure. In another example, the requesting apparatus includes a processor coupled with a computer-readable medium. The computer-readable medium stores thereon computer executable instructions that when executed cause the processor or the requesting apparatus to perform a method consistent with the implementations described above and / or elsewhere in the present disclosure. Non-limiting examples of the requesting apparatus are a user equipment (UE) , any other suitable terminal devices or modules / chips in the apparatus. In some implementations, the requesting apparatus includes a chip, e.g., an IC chip, a modem chip (also referred to as a baseband chip) , an SoC chip, and / or an SIP chip. In some implementations, the requesting apparatus may include circuitry such as an FPGA, a GPU, or an ASIC, that performs the methods. More generally, the requesting apparatus may include one or more units to perform a method as described above or elsewhere in the present disclosure. The term “units” is used in a broad sense and may be referred to as any of various names, including for example, modules, components, elements, means, etc. The units may be implemented using hardware, software, firmware or any combination thereof.

[0046] According to an aspect of the present disclosure, there is provided a computer program product. The computer program product includes a computer program (also referred to as code or an instruction) . When the computer program is run or executed, a computer is enabled or caused to perform a method as described above or elsewhere in the present disclosure.

[0047] According to an aspect of the present disclosure, there is provided a computer-readable storage medium. The computer-readable storage medium stores computer executable instructions that, when executed, cause a computer to perform a method as described above or elsewhere in the present disclosure. The computer-readable storage medium may be non-transitory.

[0048] According to an aspect of the present disclosure, there is provided a computer-program. The computer program includes computer executable instructions that, when executed, cause a computer to perform a method as described above or elsewhere in the present disclosure.

[0049] In some aspects of the present disclosure, there is provided an apparatus for implementing (or configured to perform) any of the method aspects as disclosed above or elsewhere in the present disclosure. In one example, there is provided an apparatus comprising a communication unit configured to transmit a request for an inference task, and receive an output of the inference task in accordance with first information. The first information may configure one or more parameters associated with inference quality of the interference task. In another example, there is provided an apparatus comprising a communication unit and a processing unit. The communication unit may be configured to receive an input for at least a portion of an inference task in accordance with first information, and transmit an output of the at least a portion of the inference task in accordance with the first information. The first information may configure one or more parameters associated with inference quality of the interference task. The processing unit may be configured to perform the at least a portion of the inference task in accordance with the first information. In another example, there is provided an apparatus comprising a communication unit configured to transmit an inference task request including at least part of first information, and receive an output of the inference task in accordance with the first information. The first information configuring one or more parameters associated with inference quality of an interference task. In another example, an apparatus comprising one or more processors and an interface circuit. The interface circuit may comprise one or more transceivers, and / or may be configured to transmit a request for an inference task, and receive an output of the inference task in accordance with first information. The first information may configure one or more parameters associated with inference quality of the interference task. In another example, an apparatus comprising one or more processors and an interface circuit. The interface circuit may comprise one or more transceivers, and / or may be configured to receive an input for at least a portion of an inference task in accordance with first information, and transmit an output of the at least a portion of the inference task in accordance with the first information. The one or more processors may be configured to perform the at least a portion of the inference task in accordance with the first information. The first information may configure one or more parameters associated with inference quality of the interference task. In another example, an apparatus comprising one or more processors and an interface circuit. The interface circuit may comprise one or more transceivers, and / or may be configured to transmit an inference task request including at least part of first information, and receive an output of the inference task in accordance with the first information. The first information configuring one or more parameters associated with inference quality of an interference task.

[0050] In some aspects of the present disclosure, there is provided a device for implementing (or configured to perform) any of the method aspects as disclosed above or elsewhere in the present disclosure.

[0051] In some aspects of the present disclosure, there is provided an element / chipset system including means (e.g., at least one processor) to implement the method implemented by (or at) a UE or any suitable terminal device of the present disclosure. The apparatus / chipset system may be the terminal device or a module / component in the terminal device. In details, the at least one processor may execute instructions stored in a computer-readable medium to implement the method.

[0052] In some aspects of the present disclosure, there is provided an element / chipset system including means (e.g., at least one processor) to implement the method implemented by (or at) a network device of the present disclosure. The apparatus / chipset system may be a network device (e.g., a BS, a TRP, or any other suitable network device) , a module / component in the network device, or a network function in the network device. In details, the at least one processor may execute instructions stored in a computer-readable medium to implement the method.

[0053] In some aspects of the present disclosure, there is provided a system including a computing device of the present disclosure (or an element in (or at) a computing device of the present disclosure) , a requesting device of the present disclosure (or an element in (or at) a requesting device of the present disclosure) , and a network device of the present disclosure (or an element in (or at) a network device of the present disclosure) . The computing device, requesting device, and / or network device may include at least one of a UE, a BS, a TRP, any suitable terminal device, any suitable network device, any suitable network node, any suitable network function, and / or any elements thereof. The computing device, the requesting device, and the network device may be considered a computing apparatus, a requesting device, and a network apparatus. The computing device (or computing apparatus) , the requesting device (or requesting apparatus) , and the network device (or network apparatus) may be configured to perform methods as described above or elsewhere in the present disclosure.

[0054] In some aspects of the present disclosure, there is provided a method performed by a system including a computing device of the present disclosure (or an element in (or at) a computing device of the present disclosure) , a requesting device of the present disclosure (or an element in (or at) a requesting device of the present disclosure) , and a network device of the present disclosure (or an element in (or at) a network device of the present disclosure) . The computing device, requesting device, and / or network device may include at least one of a UE, a BS, a TRP, any suitable terminal device, any suitable network device, any suitable network node, any suitable network function, and / or any elements thereof. The computing device, the requesting device, and the network device may be considered a computing apparatus, a requesting device, and a network apparatus. The computing device (or computing apparatus) , the requesting device (or requesting apparatus) , and the network device (or network apparatus) may be configured to perform methods as described above or elsewhere in the present disclosure.BRIEF DESCRIPTION OF THE DRAWINGS

[0055] For a more complete understanding of the present implementations, and the advantages thereof, reference is now made, by way of example, to the following descriptions taken in conjunction with the accompanying drawings, in which:

[0056] FIG. 1 is a schematic diagram of a communication system in which implementations of the present disclosure may occur.

[0057] FIG. 2 is another schematic diagram of a communication system in which implementations of the present disclosure may occur.

[0058] FIG. 3 is a block diagram illustrating an example of an apparatus wirelessly communicating with another apparatus in a communication system in which implementations of the present disclosure may occur.

[0059] FIG. 4 is a block diagram illustrating an example of an apparatus in which implementations of the present disclosure may occur.

[0060] FIG. 5 is a block diagram illustrating another example of an apparatus in which implementations of the present disclosure may occur.

[0061] FIG. 6 is a schematic diagram illustrating an example wireless network in which an inference pipeline is deployed, in accordance with implementations of the present disclosure.

[0062] FIG. 7 is a schematic diagram illustrating an example wireless network in which multiple inference pipelines are deployed, in accordance with implementations of the present disclosure.

[0063] FIG. 8 is a schematic diagram illustrating an example wireless network in which an example pipeline that includes computing devices and a requesting device, in accordance with implementations of the present disclosure.

[0064] FIG. 9 is a schematic diagram illustrating an example wireless network in which an example pipeline that includes computing device groups and a requesting device, in accordance with implementations of the present disclosure.

[0065] FIG. 10 is a signal flow diagram illustrating an example process for inference pipeline initialization and inference quality management, in accordance with implementations of the present disclosure.

[0066] FIG. 11 is a schematic diagram illustrating an example wireless network with multiple inference pipelines in which different inference tasks are respectively assigned, in accordance with implementations of the present disclosure.

[0067] FIGs. 12A and 12B are signal flow diagrams illustrating an example method of inference quality management for multiple inference tasks, in accordance with implementations of the present disclosure.

[0068] FIG. 13 is a schematic diagram illustrating an example wireless network in which multiple inference tasks are concurrently carried out, in accordance with implementations of the present disclosure.

[0069] FIG. 14 is a flow diagram illustrating an example method for modulation and coding scheme (MCS) adaptation, in accordance with implementations of the present disclosure.

[0070] FIG. 15 is a schematic diagram illustrating an example periodic metric check procedure, in accordance with implementations of the present disclosure.

[0071] FIG. 16 is a schematic diagram illustrating an example method of inference quality management using an Inference-related Quality of Service (I-QoS) profile, in accordance with implementations of the present disclosure.

[0072] FIG. 17 is a flow diagram illustrating an example method for pre-emptive adjustment of parameters associated with inference quality, in accordance with implementations of the present disclosure.

[0073] FIG. 18 is a flow diagram illustrating an example method for inference quality management using a history mapping table, in accordance with implementations of the present disclosure.

[0074] FIG. 19 is a flow diagram illustrating an example method of proactive response to channel deterioration, in accordance with implementations of the present disclosure.

[0075] FIG. 20 is an example graded table correlating channel conditions and hybrid automatic repeat request (HARQ) ignorance rate, and an example graph corelating signal-to-noise ratio (SNR) levels to the corresponding pre-emptive HARQ ignorance adjustments, in accordance with implementations of the present disclosure.

[0076] FIG. 21 is an example decision history table that show mapping between measured channel SNR, decisioned HARQ ignorance rates, and the measured inference quality metrics, in accordance with implementations of the present disclosure.

[0077] FIG. 22 is a schematic diagram illustrating an example wireless network in which multiple inference pipelines with distinct I-QoS profiles run concurrently, in accordance with implementations of the present disclosure.

[0078] FIG. 23 is a schematic diagram illustrating an example inference quality management for multiple tasks with different priority levels, in accordance with implementations of the present disclosure.

[0079] FIG. 24 is a flow diagram illustrating an example process for MCS adjustments for priority scheduling and HARQ ignorance, in accordance with implementations of the present disclosure.

[0080] FIG. 25 is a schematic diagram illustrating an example interval management for multiple pipelines associated with difference metrics, in accordance with implementations of the present disclosure.

[0081] FIG. 26 is a block diagram illustrating example interval adaptations based on stability of inference quality tests, in accordance implementations of the present disclosure.

[0082] FIG. 27 is a diagram illustrating example MCS / HARQ tier classes associated with a metric and a metric threshold value configured in the I-QoS profile, in accordance with implementations of the present disclosure.

[0083] FIG. 28 is a diagram illustrating three example tiers for inference quality management, specifying combinations of MCS and HARQ ignorance rates, in accordance with implementations of the present disclosure.

[0084] FIGs. 29A and 29B are signal flow diagrams illustrating an example method for HARQ ignorance rate control, in accordance with implementations of the present disclosure.

[0085] FIGs. 30A and 30B illustrate example HARQ and MCS interactive adjustments, in accordance with implementations of the present disclosure.

[0086] FIG. 31 is a schematic diagram illustrating an example data-aware HARQ ignorance control using partial cyclic redundancy check (CRC) , in accordance with implementations of the present disclosure.

[0087] FIG. 32 is a diagram illustrating an example signaling design of DCI including an inference tolerated error level (ITEL) field, in accordance with implementations of the present disclosure.

[0088] FIG. 33 is a diagram illustrating an example I-QoS class definition in RRC signaling, in accordance with implementations of the present disclosure.

[0089] FIG. 34 is a diagram illustrating example HARQ ignorance signaling fields, in accordance with implementations of the present disclosure.

[0090] FIG. 35 is a signal flow diagram illustrating an example method for user equipment (UE) capability negotiation, in accordance with implementations of the present disclosure.

[0091] FIG. 36 is a signal flow diagram illustrating an example method for error-tolerant inference in a wireless network, in accordance with implementations of the present disclosure.DETAILED DESCRIPTION

[0092] For illustrative purposes, specific example implementations will now be explained in greater detail below in conjunction with the figures.

[0093] The implementations set forth herein represent information sufficient to practice the claimed subject matter and illustrate ways of practicing such subject matter. Upon reading the following description in light of the accompanying figures, those of skill in the art will understand the concepts of the claimed subject matter and will recognize applications of these concepts not particularly addressed herein. It should be understood that these concepts and applications fall within the scope of the disclosure and the accompanying claims.

[0094] Moreover, it will be appreciated that any module, component, or device disclosed herein that executes instructions may include or otherwise have access to a non-transitory computer / processor readable storage medium or media for storage of information, such as computer / processor readable instructions, data structures, program modules, and / or other data. A non-exhaustive list of examples of non-transitory computer / processor readable storage media includes magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, optical disks such as compact disc read-only memory (CD-ROM) , digital video discs or digital versatile discs (i.e., DVDs) , Blu-ray DiscTM, or other optical storage, volatile and non-volatile, removable and non-removable media implemented in any method or technology, random-access memory (RAM) , read-only memory (ROM) , electrically erasable programmable read-only memory (EEPROM) , flash memory or other memory technology. Any such non-transitory computer / processor storage media may be part of a device / apparatus or accessible or connectable thereto. Computer / processor readable / executable instructions to implement a method, an application or a module described herein may be stored or otherwise held by such non-transitory computer / processor readable storage media.

[0095] Aspects of the present disclosure generally relate to wireless communications, and more specifically to methods, apparatuses, devices, and systems that enable error-tolerant operations in a wireless system, for example performing error-tolerant inference tasks while maintaining acceptable inference quality in a distributed wireless system.

[0096] FIGs. 1 to 5 and following below provide context for a network and device (s) that may be in the network and that may implement aspects of the present disclosure.

[0097] FIG. 1 is a schematic illustration of an example communication system according to an implementation of the present disclosure, there is shown a communication system 100 that includes a radio access network (RAN) 120, one or more communication electronic devices (EDs) 110a, 110b, 110c, 110d, 110e, 110f, 110g, 110h, 110i, 110j (collectively referred to as 110) , a core network 130, a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160. The RAN 120 may include, but is not limited to, a future generation RAN, or a legacy RAN such as, but not limited to, 5th generation (5G) , 4th generation (4G) , 3rd generation (3G) or 2nd generation (2G) radio access network. The RAN 120 may be, for example, an evolved universal mobile telecommunications system (UMTS) Terrestrial Radio Access Network (E-UTRAN) , a NextGen RAN (NG RAN) , or some other type of RAN. Examples of RAN 120 based on the evolution of telecommunications standards include, but are not limited to, GSM (Global System for Mobile Communications) and code division multiple access (CDMA) for 2G, universal mobile telecommunications system (UMTS) based on wideband code division multiple access (WCDMA) and CDMA2000 for 3G, long-term evolution (LTE) and WiMAX (Worldwide Interoperability for Microwave Access) for 4G, and new radio (NR) for 5G. In some implementations, The RAN 120 may use any radio access technology (RAT) in the wireless interface between the one or more EDs 110 and the RAN 120. In some implementations, the term “radio access” may refer to the future generation air interface standards which may include both terrestrial networks (TNs) and non-terrestrial networks (NTNs) . These networks will be described in greater detail below in conjunction with various implementations. The one or more communication EDs 110 (also referred to as “user equipment” ) are configured to connect (e.g., communicatively couple) with each other or to one or more network nodes 170a, 170b (collectively referred to as 170) in the RAN 120. The core network (CN) 130 is a part of the communication system 100 and consists of network nodes (e.g., 170a, 170b) which provide support for the network features and telecommunication services. In some implementations, the CN 130 may be dependent on the RAT used in the communication system 100. In other implementations, the CN 130 may be access-agnostic, i.e., the CN 130 may be independent of the RAT used in the communication system 100. There are different types of CN 130, for different 3GPP system generations. For example, the CN 130 is the evolved packet core (EPC) in 4G, also known as the evolved packet system (EPS) . In another example, the CN 130 is the 5G Core (5GC) which was developed as part of the 5G System (5GS) . The CN 130 also enables integration of different 3GPP and non-3GPP access types. In some implementations and referring to FIG. 1, the CN 130 also provides the interface towards external networks that may include the PSTN 140, the Internet 150, and other networks 160 in the communication system 100.

[0098] In general, the communication system 100 facilitates interaction between multiple wireless or wired elements. The communication system 100 may transmit different types of content, such as voice, data, video, and / or text, through different transmission methods such as, but not limited to, broadcast, multicast, groupcast, and unicast. Additionally, the communication system 100 operates by allocating and / or sharing resources, such as carrier spectrum bandwidth, among its constituent elements.

[0099] The communication system 100 may provide a wide range of communication services and applications including, but not limited to, Enhanced Mobile Broadband (eMBB) services, ultra-reliable low-latency communication (URLLC) services, massive machine type communication (mMTC) services, integrated sensing and communication (ISAC) , immersive communication, ultra-massive machine-type communication (uMTC) , hyper reliable and low-latency communication, ubiquitous connectivity, integrated AI and communication, and other services that may be provided by a future generation communication system. The communication system 100 may provide other services and applications such as, but not limited to, earth monitoring, remote sensing, passive sensing and positioning, navigation and tracking, autonomous delivery and mobility and the like.

[0100] The communication system 100 may include a terrestrial communication system (or network) and / or a non-terrestrial communication system (or network) . The communication system 100 may provide a high degree of availability and robustness through a joint operation of the terrestrial communication system and the non-terrestrial communication system. For example, integrating a non-terrestrial communication system (or components thereof) into a terrestrial communication system may result in a heterogeneous network including multiple layers. The heterogeneous network may achieve better overall performance through efficient multi-link joint operation, more flexible functionality sharing, and faster physical layer link switching between terrestrial networks and non-terrestrial networks. The terrestrial communication system and the non-terrestrial communication system may be considered as sub-systems of the communication system 100.

[0101] FIG. 2 illustrates another example communication system 100 according to an implementation of the present disclosure, there is shown the communication system 100 includes EDs 110a, 110b, 110c, 110d (collectively referred to as ED 110) , RANs 120a, 120b, one or more CNs 130, a PSTN 140, the Internet 150, and other networks 160. Additionally, the communication system 100 may also include a non-terrestrial network (NTN) 120c. The RANs 120a and 120b may include network nodes 170a and 170b respectively. Examples of network nodes 170a, 170b include base stations, which may be generally referred to as terrestrial network (TN) devices or terrestrial transmit and receive points (T-TRPs) 170a and 170b (collectively referred to as 170) . In this context, the terms "TRP" and "base station" are used interchangeably unless otherwise specified. For simplicity, this disclosure primarily refers to network nodes as base stations; however, unless explicitly stated otherwise, references to TRP are considered non-limiting and interchangeable. The T-TRPs 170a, 170b may be base stations mounted on a building or tower. In one implementation, the NTN 120c includes a RAN node such as a base station 172, which may be generally referred to as an NTN device, a non-terrestrial node, a non-terrestrial network device, a non-terrestrial base station, or a non-terrestrial transmit and receive point (NT-TRP) 172.

[0102] In some implementations, the NT-TRP 172 is not attached to the ground, for example, as in the case of an airborne base station. An airborne base station may be implemented using communication equipment supported or carried by a flying device. For example, a flying device may include, but is not limited to, an airborne platform (such as a blimp or an airship) , balloon, drone (such as a quadcopter) , and other types of aerial vehicles. In some implementations, an airborne base station may be supported or carried by an unmanned aerial system (UAS) or an unmanned aerial vehicle (UAV) , such as a drone. An airborne base station may be a moveable or mobile base station that may be flexibly deployed in different locations to meet network demand. A satellite base station is another example of a non-terrestrial base station. A satellite base station may be implemented using communication equipment supported or carried by a satellite. A satellite base station may also be referred to as an orbiting base station. High altitude platforms are yet another example of non-terrestrial base stations, including international mobile telecommunication base stations.

[0103] As referred to herein, and unless specified otherwise, a “TRP” may also refer to a T-TRP or an NT-TRP, a “T-TRP” may also refer to a “TN TRP” , and an “NT-TRP” may also refer to an “NTN TRP” . The NTN 120c may be considered a RAN, sharing operational aspects with RANs 120a, 120b. The NTN 120c may include at least one NTN device and at least one corresponding terrestrial network device. The at least one NTN device may function as a transport layer device and the at least one corresponding terrestrial network device may function as a RAN node, communicating with the ED 110 via the NTN device. Additionally, there may be an NTN gateway on the ground (referred to as a terrestrial network device) that also functions as a transport layer device facilitating communication with both the NTN device and the RAN node. The RAN node may communicate with the ED 110 via the NTN device and the NTN gateway. In some implementations, the NTN gateway and the RAN node may be located within the same device.

[0104] A base station 170 (also referred to as a TRP as stated above) is a network element within a radio access network responsible for radio transmission and reception in one or more cells to or from the ED (such as a user equipment (UE) ) . In different implementations, the base station 170 may also be known as a base transceiver station (BTS) , a radio base station, a network node, a network device, a device on the network side, a transmit / receive node, a Node B, an evolved NodeB (eNodeB or eNB) , a Home eNodeB, a next Generation NodeB (gNB) , a transmission point (TP) , a site controller, an access point (AP) , a wireless router, a relay station, a terrestrial node, a terrestrial network device, a terrestrial base station, a non-terrestrial node, a non-terrestrial network device, a non-terrestrial base station, and a positioning node, among other possibilities. The base station 170 may be a macro base station (BS) , a pico BS, a relay node, a donor node, or combinations thereof. When the base station 170 performs (or is configured to perform) a method described herein, it may be interpreted as the base station itself, one or more modules (or units) in the base station, a circuit or chip, or a combination thereof, performing the method. For example, the circuit or chip may include a modem chip, also referred to as a baseband chip, a system on chip (SoC) including a modem core, system in package (SIP) ) , and the like, and may be responsible for one or more communication functions within the base station.

[0105] The EDs 110a-110d and TRPs 170a-170b, 172 are examples of communication equipment configured to implement some or all of the operations and / or implementations described herein. The T-TRP 170a forms part of the RAN 120a, which may include other TRPs, and / or other devices. Also, the TRP 170b forms part of the RAN 120b, which may include other TRPs, and / or devices. Each TRP 170a, 170b may transmit and / or receive wireless signals within a particular geographic region or area, sometimes referred to as a “cell” or a “coverage area” . The TRPs 170a-170b may be responsible for allocating and / or configuring resources and transmission and / or reception in a set of cell (s) . A cell is a radio network object that may be uniquely identified by a cell identification that is broadcasted over a geographical region or area from base stations associated with the cell. A cell may work in either FDD or TDD mode. A cell may be further divided into cell sectors, and a base station 170a-170b may, for example, employ one or more transceivers to provide services to one or more sectors. Some implementations may include pico or femto cells if supported by the radio access technology. In some implementations, one or more transceivers may be used for each cell, such as with multiple-input multiple-output (MIMO) technology. The number of RANs 120a-120b shown is merely an example. Any number of RANs may be contemplated when designing the communication system 100.

[0106] A base station may be a single element, as shown in the figures, or multiple elements distributed throughout the corresponding RAN, or otherwise configured. In some implementations, a plurality of RAN nodes coordinate to assist the ED 110 in implementing radio access, and different RAN nodes separately implement and handle different functions of the base station. For example, the RAN node may be a central unit (CU) , a distributed unit (DU) , a CU-control plane (CP) , a CU-user plane (UP) , or a radio unit (RU) etc. The CU and the DU may be separately deployed, or included within the same element (i.e., a baseband unit (BBU) ) . The RU may be included in a radio frequency device or a radio frequency unit (i.e., a remote radio unit (RRU) , an active antenna unit (AAU) , or a remote radio head (RRH) ) . In different systems, the CU (or the CU-CP and the CU-UP) , the DU, or the RU may be known by different names, but their functions are understood by a person skilled in the art. For example, in an open radio access network (ORAN) system, a CU may be referred to as an open CU (O-CU) , a DU may be referred to as an open DU (O-DU) , and a CU-CP may be referred to as an open CU-CP (O-CU-CP) . The CU-UP may also be referred to as an open CU-UP (O-CU-UP) , and the RU may also be referred to as an open RU (O-RU) . Any one of the CU (or the CU-CP, the CU-UP) , the DU, and the RU may be implemented using a software module, a hardware module, or a combination of a software module and a hardware module.

[0107] Furthermore, communication between different devices / apparatuses in various implementations of this disclosure may refer to direct communication (that is, without the need of forwarding by another device / apparatus) or may refer to communication (s) between different devices / apparatuses via another device / apparatus (that is, requiring forwarding by another device / apparatus) . Alternatively, such communication (s) may involve one functional unit inside a device / apparatus using another functional unit within the device / apparatus to communicate with another device / apparatus. In other words, phrases such as "sending (or transmitting) information to... (an ED or a base station) " in this disclosure may be understood as a destination endpoint of the information being an ED or a base station, including, sending / transmitting information directly or indirectly to an ED or a base station. Similarly, phrases like "receiving information from... (an ED or a base station) " may be understood as a source endpoint of the information being an ED or a base station, including directly or indirectly receiving information from an ED or a base station. Between the source endpoint that sends the information and the destination endpoint, necessary processing such as, but not limited to, format conversion, digital-to-analog conversion, amplification, and filtering may be performed on the information. However, the destination endpoint may understand valid information from the source endpoint. A similar understanding applies to other descriptions in this disclosure without reiterating details already described. In the present disclosure, the terms "send" and "transmit" may be used interchangeably in different implementations of this disclosure.

[0108] The ED 110 is used to connect people, objects, machines, and other entities. The ED 110 may be widely used in various scenarios including, but not limited to, cellular communications, device-to-device (D2D) , vehicle to everything (V2X) , peer-to-peer (P2P) , machine-to-machine (M2M) , MTC, internet of things (IoT) , virtual reality (VR) , augmented reality (AR) , mixed reality (MR) , metaverse, digital twin, industrial control, self-driving, remote medical, smart grid, smart furniture, smart office, smart wearable, smart transportation, smart city, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, and autonomous delivery and mobility.

[0109] Each ED 110 represents any suitable end user device for wireless operation and may include such devices (or may be referred to as, but not limited to) a user equipment (UE) or a user device or a terminal device, a wireless transmit / receive unit (WTRU) , a mobile station, a fixed or mobile subscriber unit, a cellular telephone, a station (STA) , an MTC device, a personal digital assistant (PDA) , a smartphone, a laptop, a computer, a tablet, a wireless sensor, a consumer electronics device, a smart book, a vehicle, a car, a truck, a bus, a train, or an IoT device, wearable devices (such as a watch, a pair of glasses, head mounted equipment, etc. ) , an industrial device, or an apparatus (such as a module, modem, or chip) in the forgoing devices, among other possibilities. Future generation EDs 110 may be referred to by other terms. When an ED 110 performs (or is configured to perform) a method described herein, it may be interpreted as the ED itself, one or more modules (or units) in the ED, a circuit or chip, or a combination thereof, performing the method. For example, the circuit or chip may include a modem chip, also referred to as a baseband chip, a system on chip (SoC) including a modem core, or system in package (SIP) ) , and the like, and may be responsible for one or more communication functions in the ED.

[0110] Each ED 110 connected to TRPs 170a-170b, and / or TRPs 172 may be dynamically or semi-statically turned-on (i.e., established, activated, or enabled) , turned-off (i.e., released, deactivated, or disabled) and / or configured in response to one or more of: connection availability and connection necessity.

[0111] Any ED 110 may be alternatively or additionally configured to interface, access, or communicate with any of the TRPs 170a, 170b and 172, the Internet 150, the CN 130, the PSTN 140, the other networks 160, or any combination thereof. In some examples, the ED 110a may communicate an uplink (UL) and / or downlink (DL) transmission over a terrestrial air interface 190a with station-TRP 170a. In some examples, the EDs 110a, 110b, 110c, and 110d may also communicate directly with one another via one or more sidelink (SL) air interfaces 190b. In some examples, the EDs 110a, 110d may communicate using a UL and / or a DL transmission over a non-terrestrial air interface 190c with NT-TRP 172.

[0112] An air interface (such as, for example, 190a, 190b, 190c) generally includes a number of components and associated parameters that collectively specify how a transmission is to be sent and / or received over a wireless communications link between two or more communicating devices such as EDs and base station (s) . For example, an air interface may include one or more components defining the waveform (s) , frame structure (s) , multiple access scheme (s) , protocol (s) , coding scheme (s) and / or modulation scheme (s) for conveying information (such as, data) over a wireless communications link. The air interfaces 190a and 190b may use similar communication technology, that may include any suitable radio access technology.

[0113] The non-terrestrial air interface 190c may enable communication between the EDs 110a, 110d and one or more NT-TRPs 172 via a wireless link or simply a link. For some examples, the link is a dedicated connection for unicast transmission, a connection for broadcast transmission, or a connection between a group of EDs 110 and one or more NT-TRPs 172 for multicast transmission.

[0114] The TRPs 170a-170b, 172 may communicate with one another over one or more air interfaces 190e, 190f using wireless communication links (such as radio frequency (RF) , microwave, infrared (IR) , etc. ) or wired communication links. The air interfaces 190e, 190f may utilize any suitable radio access technology, and may be substantially similar to the air interfaces 190a, 190c over which the EDs 110a-110d communicate with one or more of the TRP 170a-170b, 172 or they may be substantially different. For example, the communication system 100 may implement one or more channel access methods, such as time division multiple access (TDMA) , frequency division multiple access (FDMA) , code division multiple access (CDMA) , Single Carrier Frequency Division Multiple Access (SC-FDMA) , Low Density Signature Multicarrier Code Division Multiple Access (LDS-MC-CDMA) , non-orthogonal multiple access (NOMA) , pattern division multiple access (PDMA) , lattice partition multiple access (LPMA) , resource spread multiple access (RSMA) , and sparse code multiple access (SCMA) .

[0115] The RANs 120a and 120b are in communication with the CN 130 to provide the EDs 110a, 110b, and 110c with various services such as voice, data, multimedia, and other services. The RANs 120a and 120b and / or the CN 130 may be in direct or indirect communication with one or more other RANs (not shown) , which may or may not be directly served by the CN 130, and may employ different radio access technologies from RAN 120a and / or RAN 120b. The CN 130 may also serve as a gateway access between (i) the RANs 120a and 120b and / or the EDs 110a, 110b, and 110c, and (ii) other networks (such as the PSTN 140, the Internet 150, and the other networks 160) . In addition, some or all of the EDs 110a, 110b, and 110c may include functionality for communicating with different wireless networks over different wireless links using different wireless technologies and / or protocols. For example, the EDs 110a, 110b, and 110c communicate using different cellular communications protocols, such as, but not limited to, a Global System for Mobile Communications (GSM) protocol, a code-division multiple access (CDMA) network protocol, a Push-to-Talk (PTT) protocol, a PTT over Cellular (POC) protocol, a universal mobile telecommunications system (UMTS) protocol, a 3GPP long term evolution (LTE) protocol, a fifth generation (5G) protocol, a new radio (NR) protocol, and the like. Instead of wireless communication (or in addition thereto) , the EDs 110a, 110b, and 110c may communicate using wired communication channels to a service provider or switch (not shown) , and / or to the Internet 150. The PSTN 140 may include circuit switched telephone networks for providing plain old telephone service (POTS) . The Internet 150 may include a network of computers and subnets (intranets) or both, and incorporate protocols, such as internet protocol (IP) , transmission control protocol (TCP) , user datagram protocol (UDP) . The EDs 110a, 110b, and 110c may be multimode devices capable of operation according to multiple radio access technologies, and may incorporate one or multiple transceivers necessary to support such.

[0116] In addition, the communication system 100 may comprise a sensing agent (not shown) to manage the sensed data from ED 110 and / or any one of TRPs 170a, 170b, 172. In one implementation, the sensing agent may be part of any one of TRPs 170a, 170b, 172. In another implementation, the sensing agent is a separate node that may communicate with the CN 130 and / or the RAN 120 (such as any one of TRPs 170a, 170b, 172) .

[0117] Additional details regarding the EDs 110, T-TRP 170, and NT-TRP 172 are known to those of skill in the art. As such, these details are omitted here.

[0118] FIG. 3 is a schematic illustration showing an example of an apparatus 310 wirelessly communicating with another apparatus 320 within a communication system (e.g., the communication system 100) according to an implementation of the present disclosure. The apparatus 310 may be an electronic device (such as ED 110) . The apparatus 320 may be a network node (e.g., the network node 170) such as a T-TRP 170 or an NT-TRP 172. Although only one apparatus 310, and one other apparatus 320 are shown in the figure, the number of apparatus 310 and / or the number of apparatus 320 may vary, potentially including one or more of each. For example, a single ED 110 may be served by a single T-TRP 170 (or a single NT-TRP 172) , or by multiple T-TRPs 170 (or multiple NT-TRPs 172) . Similarly, a single ED 110 may be served by one or more T-TRPs 170 and one or more NT-TRPs 172. Similarly, a single T-TRP 170 (or a single NT-TRP 172) may serve one or more EDs 110.

[0119] The apparatus 310 may include one or more processors 210. For clarity and to avoid overcrowding the illustration, only a single processor 210 is illustrated. The apparatus 310 may further include a transmitter 201 and a receiver 203 coupled to one or more antennas 204. For clarity, only a single antenna 204 is illustrated. One, some, or all of the antennas 204 may alternatively be panels. In some implementations, the transmitter 201 and the receiver 203 are separate from each other. In other implementations, the transmitter 201 and the receiver 203 may be integrated into a single unit, for example, as a transceiver. The transceiver is configured to modulate data or other content for transmission by the one or more antennas 204 or a network interface controller (NIC) . The transceiver may also be configured to demodulate data or other content received by the one or more antennas 204. A transceiver may include any suitable structure for generating signals for wireless or wired transmission and / or for processing signals received through wireless or wired communication. Each antenna 204 includes any suitable structure for transmitting and / or receiving wireless or wired signals. The apparatus 310 may include a memory 208. In some implementations, the apparatus 310 may include multiple memories 208. Only a single transmitter 201, receiver 203, processor 210, memory 208, and antenna 204 is illustrated for simplicity, but the apparatus 310 may include one or more other components. In some implementations of the present disclosure, the transceiver (or transmitter 201 and / or receiver 203) may be viewed as an interface circuit.

[0120] The memory 208 is configured to store instructions used to perform operations described herein. The memory 208 may also be configured to store data that is used, generated, or collected by the apparatus 310. For example, the memory 208 may store software instructions or modules configured to implement some or all of the functionalities and / or operations described herein and that which are executed by the one or more processors 210.

[0121] The apparatus 310 may further include one or more input / output devices (not shown) or interfaces. The input / output devices or interfaces facilitate interaction with a user or other devices in the network. Each input / output device or interface includes suitable components for facilitating transmission of information to a user and reception of information from a user, and for various network interface communications. Such components may include, but are not limited to, a speaker, microphone, keypad, keyboard, display, touch screen, and the like.

[0122] The processor 210 may be configured to perform (or control the apparatus 310 to perform) operations (or methods) described herein as being performed by the apparatus 310. For example, the processor 210 performs or controls the apparatus 310 to perform the operations of: a) receiving one or more transport blocks (TBs) , b) using a resource for decoding at least one of the received TBs, c) releasing the resource for decoding another of the received TBs, and / or d) receiving configuration information configuring a resource. Specifically, the operations may include tasks related to: preparing a transmission for UL transmission to the apparatus 320, processing DL transmissions received from the apparatus 320, and handling SL transmission to and from another apparatus 310. Processing operations related to preparing a transmission for UL transmission may include operations such as, but not limited to, encoding, modulating, transmit beamforming, and generating symbols for transmission. Processing operations related to processing DL transmissions may include operations such as, but not limited to, receive beamforming, demodulating and decoding received symbols. Processing operations related to processing SL transmissions may include operations such as, but not limited to, transmit / receive beamforming, modulating / demodulating and encoding / decoding symbols. Depending upon the implementation, a DL transmission may be received by the receiver 203, possibly using receive beamforming, and the processor 210 may extract signaling from the DL transmission (such as by detecting and / or decoding the signaling) . An example of signaling may be a reference signal transmitted by the apparatus 320. In some implementations, the processor 210 implements the transmit beamforming and / or the receive beamforming based on the indication of beam direction, such as beam angle information (BAI) , received from the apparatus 320. In some implementations, the processor 210 may be configured to perform operations relating to network access (such as initial access) and / or downlink synchronization, which includes operations for detecting a synchronization sequence, decoding and obtaining the system information, and the like. In some implementations, the processor 210 may perform channel estimation, such as using a reference signal received from the apparatus 320.

[0123] Although not illustrated, in some implementations, the processor 210 may either be a part of the transmitter 201 or a part of the receiver 203 or a part of both the transmitter 201 and the receiver 203. Although not illustrated, in some implementations, the memory 208 may be a part of the processor 210.

[0124] The processor 210, along with the processing components of the transmitter 201 and the receiver 203 may each be implemented by one or more processors that may be the same or different. These processors are configured to execute instructions stored in a memory (such as in the memory 208) .

[0125] The apparatus 320 includes one or more processors 260 (only one processor 260 is illustrated) . The apparatus 320 may further include one or more transmitters 252 and one or more receivers 254 coupled to one or more antennas 256. Only a single antenna 256 is illustrated to avoid clutter in the illustration. One, some, or all of the antennas 256 may alternatively be panels. In some implementations, the transmitter 252 and the receiver 254 are separate from each other. In other implementations, the transmitter 252 and the receiver 254 may be integrated into a single unit such as, for example, as a transceiver. The apparatus 320 may further include a memory 258. In some implementations, the apparatus 320 may include multiple memories 258. The apparatus 320 may further include a scheduler 253. Only a single transmitter 252, receiver 254, processor 260, memory 258, antenna 256 and scheduler 253 are illustrated for simplicity, however the apparatus 320 may include one or more other components. In the present disclosure, in some implementations, the transceiver (or transmitter 252 and / or receiver 254) may be viewed as an interface circuit.

[0126] In some implementations, various components of the apparatus 320 may be distributed. For example, some of the modules of the apparatus 320 may be located remotely from the equipment housing the antennas 256 for the apparatus 320 (and therefore also may be viewed as one or more nodes) . These modules, which may be considered as one or more nodes, may be coupled to the equipment that houses the antennas 256 over a communication link (not shown) , sometimes referred to as front haul, such as the common public radio interface (CPRI) . Therefore, in some implementations, the term apparatus 320 may also refer to network-side nodes that perform processing operations such as, but not limited to, determining the location of the apparatus 310, resource allocation (scheduling) , message generation, and encoding / decoding, and that which are not necessarily part of the equipment that houses the antennas 256 of the apparatus 320. The nodes may also be coupled to other apparatuses 320. In some implementations, the apparatus 320 may actually be a plurality of nodes that are operating together to serve the apparatus 310, such as through the use of coordinated multipoint transmissions, or through the use of an ORAN system as described above in the disclosure.

[0127] The processor 260 is configured to perform operations including those related to: preparing a transmission for DL transmission to the apparatus 310, processing an UL transmission received from the apparatus 310, preparing a transmission for backhaul transmission to another apparatus 320, and processing a transmission received over backhaul from another apparatus 320. Processing operations related to preparing a transmission for DL or backhaul transmission may include operations such as, but not limited to, encoding, modulating, precoding (such as MIMO precoding) , transmit beamforming, and generating symbols for transmission. Processing operations related to processing received transmissions in the UL or over backhaul may include operations such as, but not limited to, receive beamforming, demodulating received symbols, and decoding received symbols. The processor 260 may also be configured to perform operations relating to network access (such as initial access) and / or DL synchronization, such as generating the content of synchronization signal blocks (SSBs) , generating the system information, and the like. In some implementations, the processor 260 is further configured to generate an indication of beam direction, such as BAI, which may be scheduled for transmission by the scheduler 253 which will be described below. In some implementations, the processor 260 implements the transmit beamforming and / or receive beamforming based on beam direction information (such as BAI) received from another apparatus 320. The processor 260 is configured to perform other network side processing operations described herein, such as, but not limited to, determining the location of the apparatus 310, determining where to deploy another apparatus 320, and the like. In some implementations, the processor 260 may generate signaling data, to configure one or more parameters of the apparatus 310 and / or one or more parameters of another apparatus 320. Any signaling data generated by the processor 260 is sent by the transmitter 252. In some implementations, the apparatus 320 implements physical layer processing. In some implementations, the apparatus 320 may perform higher layer functions such as those at the medium access control (MAC) or radio link control (RLC) layers in addition to physical layer processing. In the apparatus 320, the scheduler 253 may be coupled to the processor 260 or integrated within the processor 260. In some implementations, the scheduler 253 may be integrated within the apparatus 320 or may be operated separately from the apparatus 320. The scheduler 253 may schedule UL, DL, SL, and / or backhaul transmissions, including issuing scheduling grants and / or configuring scheduling-free (such as “configured grant” ) resources.

[0128] The apparatus 320 may further include a memory 258 that is configured to store instructions for performing the operations described herein. The memory 258 may also store data that is used, generated, or collected by the apparatus 320. For example, the memory 258 may store software instructions or modules configured to implement some or all of the functionalities and / or implementations described herein and that which are executed by the processor 260.

[0129] Although not illustrated, the processor 260 may be implemented as part of the transmitter 252 and / or a part of the receiver 254. Although not illustrated, in some implementations, the processor 260 may implement the scheduler 253 and the memory 258 may be implemented as part of the processor 260.

[0130] The processor 260, the scheduler 253, the processing components of the transmitter 252, and the processing components of the receiver 254 may each be implemented by the same or different processors that are configured to execute instructions stored in a memory, such as in the memory 258.

[0131] The apparatus 320 and / or the apparatus 310 may include other components, not shown or described herein for the sake of clarity.

[0132] Note that the term “signaling” , as used herein, may alternatively be referred to as control signaling, control message, control information, or message for simplicity. Signaling between a base station (such as the TRP 170a, 170b, 172) and a UE or sensing device (such as ED 110) , or signaling between a different UE or sensing device (such as between ED 110a and ED 110b) may be carried in physical layer signaling (also referred to as dynamic signaling) , which is transmitted in a physical layer control channel. For DL, the physical layer signaling may be known as downlink control information (DCI) which is transmitted in a physical downlink control channel (PDCCH) . For UL, the physical layer signaling may be known as uplink control information (UCI) which is transmitted in a physical uplink control channel (PUCCH) . For SL, signaling between different UEs or sensing devices (such as between ED 110a and ED 110b) may be known as SL control information (SCI) which is transmitted in a physical sidelink control channel (PSCCH) . Signaling may be carried in a higher layer (such as higher than physical layer) signaling, which is transmitted in a physical layer data channel, such as in a physical downlink shared channel (PDSCH) for downlink signaling, in a physical uplink shared channel (PUSCH) for uplink signaling, and in a physical sidelink shared channel (PSSCH) for SL signaling. Higher layer signaling may also be referred to as static signaling, or semi-static signaling. The higher layer signaling may include radio resource control (RRC) protocol signaling or media access control -control element (MAC-CE) signaling. Signaling may be included in a combination of physical layer signaling and higher layer signaling.

[0133] It should be noted that in the present disclosure, “information” , when different from “message” , may be carried within a single message, or may be carried in multiple separate messages.

[0134] FIG. 4 illustrates an example apparatus 410 according to an implementation of the present disclosure. The apparatus 410 may be a communication device or an apparatus implemented in a communication device such as the ED 110 or the TRPs 170a, 170b, 172. For example, the apparatus 410 implemented in an ED may be an integrated circuit, which in some instances may be referred to as a chip, a modem, a modem chip, a baseband chip, or a baseband processor. In some implementations, one or more integrated circuits may be packaged into a system-on-chip, a system-in-package, or a multi-chip module. The apparatus 410 may include one or more integrated circuits and other discrete components. In some implementations, the apparatus 410 may be a module within the ED 110, or within the apparatus 310. In some implementations, the apparatus 410 may be a module within one of the TRPs 170a, 170b, 172, or the apparatus 320.

[0135] In an example, the apparatus 410 may include one or more processors 411, and an interface circuit 412. The apparatus 410 may further include a memory 413. The one or more processors 411 are configured to process signals and execute one or more communication protocols. The memory 413 is configured to store at least a part of corresponding computer program instructions and / or data. In an example, the one or more processors 411 execute the computer program instructions stored in the memory 413 to implement related operations (for example, inputting, outputting, receiving, and transmitting) in the method implementations disclosed herein. In some implementations, the memory 413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 413 is configured to store all of the corresponding computer program instructions and / or data for execution by the one or more processors 411. In some implementations, the memory 413 being configured to store the corresponding computer program instructions and / or data may mean that the memory 413 is configured to store a part of the corresponding computer program instructions and / or data. For example, the part of the corresponding computer program instructions and / or data may include computer program instructions and / or data that need to be currently executed by the one or more processors 411. Thus, the memory 413 may store different parts of computer program instructions and / or data for a plurality of times for the one or more processors 411 to perform related operations in the method implementations disclosed herein. As a communication interface, the interface circuit 412 is configured to implement communication with another component. For example, the interface circuit 412 may communicate a signal with another apparatus or system, such as a radio frequency processing apparatus or another processor. The signal may include or carry information intended as a payload, such as user data, control information, etc. The signal may also include or carry information useful to a receiver, but not necessarily as a payload, such as a pilot signal or reference signal. Communicating the signal may include transmitting the signal to another component or device. Communicating the signal may additionally or alternatively include receiving the signal from another component or device. Transmitting the signal may include outputting the signal to a component or device that is directly or indirectly coupled to the interface circuit 412. Receiving the signal may include inputting or obtaining the signal from a component or device that is directly or indirectly coupled to the interface circuit 412. Optionally, to reduce a load of the one or more processors, a baseband signal processing circuit 414 may be also disposed to implement processing of at least a part of the baseband signals, including signal demodulation, modulation, encoding, decoding, or the like.

[0136] The apparatus 410 may be the processor 210 (or 260) within the apparatus 310 (or 320) , in some scenarios, or may be included within the processor 210 (or 260) within the apparatus 310 (or 320) in some scenarios. The apparatus 410 may be a baseband chip or may include a baseband chip. In some implementations, the apparatus 410 may be independently packaged into a chip. In some implementations, the apparatus 310 (or 320) includes different types of chips. The apparatus 410 may be packaged into a processor chip (for example, an SoC chip or an SIP chip) with the different types of chips. In some implementations, the apparatus 410 may be packaged into a chip with some or all of circuits of a radio frequency processing system that may further be included in the apparatus 310 (or 320) .

[0137] FIG. 5 illustrates an example apparatus 510 according to an implementation of the present disclosure. The apparatus 510 may include corresponding modules or units configured to implement methods and / or implementations described herein. In some implementations, the apparatus 510 includes a processing unit 512 and a communication unit 513. Optionally, the apparatus 510 may further include a storage unit 511 configured to store apparatus program code (or instructions) and / or data.

[0138] The apparatus 510 may be an ED side apparatus, for example, an ED or a module in an ED, or a circuit or a chip responsible for a communication function in an ED. In some implementations, the apparatus 510 may be the apparatus 310. The processing unit 512 may be the processor 210. The communication unit 513 may comprise a receiving unit and / or a transmitting unit. The receiving unit and / or the transmitting unit may be the transmitter 201 and / or the receiver 203 respectively. The storage unit 511 may be the memory 208.

[0139] The apparatus 510 may be a base station side apparatus, for example, a base station or a module in a base station, or a circuit or a chip responsible for a communication function in a base station. In some implementations, the apparatus 510 may be the apparatus 320. The processing unit 512 may be the processor 260 (the scheduler 253 may also be included) . The communication unit 513 may comprise a receiving unit and / or a transmitting unit. The receiving unit and / or the transmitting unit may be the transmitter 252 and / or the receiver 254 respectively. The storage unit 511 may be the memory 258.

[0140] In some implementations, when the apparatus 510 is an ED 110 or a module in an ED 110, a function of the apparatus 510 may be implemented by one or more processors. Specifically, the processor may include a modem chip, or a system on chip (SoC) chip or an SIP chip that includes a modem core. A function of the communication unit 513 may be implemented by a transceiver circuit.

[0141] In some implementations, when the apparatus 510 is a circuit or a chip that is responsible for a communication function in an ED 110, such as a modem chip, a system on chip (SoC) chip or an SIP chip that includes a modem core –a function of the processing unit 512 may be implemented by a circuit system within the chip which includes one or more processors. A function of the communication unit 513 may be implemented by an interface circuit or a data transceiver circuit on the chip.

[0142] It may be understood that the units in the apparatus 510 may be logical or functional. Each function may correspond to one functional unit, or two or more functions may be integrated into a single functional unit. In actual implementation, all or some of the units may be integrated into a single physical entity, or may be distributed across different physical entities. In addition, the functional units may be implemented in the form of hardware, software, or a combination of hardware and software. Whether a function is implemented in the form of hardware or software depends on particular applications and design constraint conditions of the technical solutions. A person skilled in the art may use different methods to implement the described functions for specific applications, but it should not be considered that the implementation goes beyond the scope of this disclosure.

[0143] In an example, a functional unit in any one of the apparatuses may be configured as one or more integrated circuits for implementing the methods disclosed herein, for example, as one or more application-specific integrated circuits (application-specific integrated circuits, ASICs) , one or more central processing units (CPUs) , one or more microprocessors or microprocessor units (MPUs) , one or more microcontrollers or microcontroller units (MCUs) , one or more digital signal processors (DSPs) , one or more field programmable gate arrays (FPGAs) , or a combination of these.

[0144] In an example, the storage unit 511 may include a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, and / or a register.

[0145] A processor may be referred to as a processor system, an application processor, a baseband processor, a processor circuit, or a processor core. The processor may include one or a combination of one or more central processing units (CPUs) , one or more digital signal processors (DSPs) , one or more microprocessors (microprocessor units, MPUs) , one or more microcontrollers (microcontroller units, MCUs) , one or more graphics processing units (GPUs) , one or more field programmable gate arrays (FPGAs) , one or more artificial intelligence processors (AI processors) , or one or more neural network processing units (NPUs) .

[0146] A memory or a storage unit may include one or more of the following storage media: a random access memory (RAM) , a static random access memory (static RAM, SRAM) , a dynamic random access memory (dynamic RAM, DRAM) , a phase-change memory (PCM) , a resistive random access memory (resistive RAM, ReRAM) , a magnetoresistive random access memory (magnetoresistive RAM, MRAM) , a ferroelectric random access memory (ferroelectric RAM, FRAM) , a cache, a register, a read-only memory (ROM) , a flash memory (flash memory) , an erasable programmable read-only memory (erasable programmable ROM, EPROM) , a hard disk, and the like. In an example, computer program instructions used to execute implementations may be stored in a non-volatile memory, for example, at least a part of a memory or storage unit (for example, one or more of a ROM, a flash memory, an EPROM, or a hard disk) . When a terminal runs, a part or all of corresponding computer program instructions may be loaded to a memory that has a higher transmission speed with the processor, for example, at least a part of a memory or a storage unit (for example, one or more of a RAM, an SRAM, a DRAM, a PCM, a ReRAM, an MRAM, a FRAM, a cache, or a register) , so that the processor executes the computer program instructions to perform the steps in the method implementations disclosed herein.

[0147] High-performance artificial intelligence (AI) model inference, for example large language models (LLMs) and complex multimodal networks, may rely on extensive resources of hardware accelerated processing unit (HAPU) such as GPUs, CPUs, tensor processing units (TPUs) , NPUs or other forms of processing units.

[0148] Some implementations harness a large number of HAPUs at a single site (e.g., a core network location) , resulting in various challenges, for example those illustrated below: 1. Intensive thermal management: Concentrated HAPU clusters generate considerable heat. Maintaining a suitable  operating temperature requires robust cooling systems, which in turn consume large amounts of energy. This leads to high power usage effectiveness (PUE) , indicating that a significant portion of total energy expenditure is devoted to cooling rather than computation. 2. Specialized power infrastructure: Centralized HAPU installations often demand specialized, costly, high-capacity  power supplies. These arrangements increase both capital expenditures (CapEx) and operational expenditures (OpEx) . 3. Specialized Super-Computer Servers: Centralized Servers are often comprised of massive bleeding-edge super- computer servers to handle huge amount of computation requests that they receive. These super-computers are not always available and reduce the scalability and availability of the computation. As the IoT technologies advance, the demand for computation increases substantially but these super-computing servers are not available everywhere.

[0149] These factors significantly drive up costs and complexity. To mitigate these issues, a distributed deployment model for HAPUs has emerged as a promising alternative. Instead of one massive HAPU cluster at a single location, multiple smaller HAPU network devices (or even HAPU-equipped UEs, or computing UEs) may be deployed closer to end-user devices or at various distributed edge locations. Each smaller computing UE is simpler to cool-often using standard environments-and may rely on normal electrical loads. This distribution directly reduces cooling requirements, lowers PUE, and avoids the complexity of specialized power infrastructure.

[0150] However, distributed Computing UEs (or distributed HAPUs) deployment introduces new challenges. Since computations are spread across multiple network devices, intermediate inference layers and partial model computations may be exchanged over wireless links. To ensure that the thermal and power savings from distributed deployment are not overshadowed, the wireless transmission overhead costs (e.g., in terms of both energy and bandwidth usage) may remain lower than the costs avoided by not having a centralized HAPU cluster. To ensure this, the quality requirements of the model inference may be carefully and appropriately taken into consideration. For example, while high throughput and low latency are necessary, perfect reliability may not always be needed. Therefore, in case of distributed inference, the inference pipelines -where split partial AI model computations are distributed over multiple Computing UEs or edge nodes-may exchange a large volume of intermediate inference data over wireless links, while attaining ultra-low latency, efficient resource utilization, and a nuanced balance between reliability and flexibility.

[0151] Some wireless networks, such as 5th generation (5G) NR, introduce communication classes like Enhanced mobile broadband (eMBB) and ultra-reliable low-latency communication (URLLC) for different mission-critical services. Especially for short-latency mission, URLLC emphasizes low latency and near-perfect data delivery through excessive retransmissions and robust error correction. While this approach ensures extremely high reliability, it typically incurs additional latency and energy consumption, which is counterproductive in scenarios where minor inference errors are tolerable. Such strict reliability requirements may prevent the network from flexibly reducing overhead when some errors are acceptable.

[0152] In the present disclosure, the 5G, NR, and / or 5G NR may be used as non-limiting examples of existing wireless network systems. Some or all aspects illustrated in the present disclosure in connection with 5G, NR, and / or 5G NR may be similarly applicable to other existing wireless networks, such as LTE, 4G, 3G, and / or 2G radio access network.

[0153] In other words, the URLLC's perfectionist approach may not always be the best fit for large-scale distributed inference. For example, under the URLLC, reliability may not be lowered to cut down overhead unless permissible. Without the ability to tolerate minor errors, wireless transmission costs (latency and energy) might diminish the benefits gained from distributed HAPU deployments.

[0154] Existing link adaptation and scheduling methods, for example those described in 3GPP specifications and related literature, primarily respond to channel indicators (e.g., CQI, SNR) and static QoS parameters, lacking mechanisms to integrate application-level inference quality metrics (e.g., Perplexity for linguistic tasks, Accuracy for vision tasks) . Without such integration, systems may miss opportunities to reduce latency by ignoring certain HARQ retransmissions in scenarios where inference quality would remain sufficiently good with a tolerable BER.

[0155] The term “QoS” may refer to “quality of service, and the term “BER” may refer to “bit error rate” . Further, the term “CQI” may refer to “channel quality indicator” , and the term “SNR” may refer to “signal-to-noise ratio” .

[0156] As future generation networks emerge, there is an industry expectation to support more flexible QoS frameworks and protocol designs that allow inference-level metrics to guide physical (PHY) and medium access control (MAC) adaptations, including HARQ ignorance rate and modulation coding scheme (MCS) adjustments and Inference-QoS (I-QoS) -based signaling.

[0157] Aspects of the present disclosure are illustrated in context of a distributed inference framework that: ● Provides Thermal and Power Efficiency via Distributed HAPUs: Spreading HAPUs across multiple computing  UEs (or edge nodes) reduces local heat density and allows reliance on standard power conditions rather than expensive, specialized infrastructures. ● Builds up a Wireless Interconnection for Distributed Inference: Pipelines of inference computations may form  between UEs, and be coordinated by a base station (for example, a T-TRP) , thereby enabling intermediate inference data exchange among Computing UEs. For example, one Requesting UE requests an inference task, the base station (e.g., T-TRP) assigns parts of the inference workload to one or more Computing UEs with HAPUs, and these Computing UEs transmit intermediate inference data to each other. The T-TRP may primarily serve as a coordination and routing point rather than a heavy compute node. ● Employs Adaptive Reliability: While connecting Computing UEs via wireless links eases cooling and power  demands, the intermediate data of the model inference may need to be exchanged reliably. However, perfect reliability (e.g., URLLC approaches) leads to retransmissions and energy / latency overhead. There is an opportunity to allow controlled error tolerance in these intermediate transmissions. Some inference tasks may tolerate minor data errors or packet losses. Others may not. Thus, unlike prior art that aims for near-perfect transmission, implementations of the present disclosure consider inference-level quality metrics to guide the allowable error tolerance. ● Utilizes Inference-Level Quality Metrics (e.g., Perplexity) : To ensure that communication overhead does not  outweigh the benefits of distributed deployment, the system dynamically adjusts error tolerance using inference-level quality metrics like Perplexity and the like. Because quality metric might need to be measured periodically to account for time varying wireless conditions, the network may incorporate mechanisms to periodically revalidate inference quality. Such continuous or intermittent measurements ensure that the chosen error tolerance remains aligned with actual inference performance.

[0158] Accordingly, a method addressing error tolerance is provided in the present disclosure. The method provides an "error tolerance" framework. Instead of enforcing near-perfect reliability, it allows controlled levels of noise and minor errors in wireless transmissions, which may be guided by inference-level quality metrics (e.g., Perplexity) . By carefully selecting when to accept small errors, the system ensures that the overhead of wireless inter-node communication stays below the overhead of centralized cooling and special power supply demands. As a result, a viable shift from large, centralized HAPU clusters to distributed HAPU deployments interconnected by wireless systems is enabled, thereby yielding improved PUE, simplifying power infrastructure, and reducing overall operational cost.

[0159] The present disclosure further introduces a novel Inference-related QoS (I-QoS) framework. I-QoS profiles incorporate inference-quality metrics and dynamic policies to continuously optimize transmission parameters, including HARQ, MCS, and resource allocation, based on real-time inference performance and channel conditions. With the I-QoS framework / profile, the network may adapt error tolerance and link configuration in real time (or in near real time) and ensure that inference pipelines operate at minimal overhead while maintaining acceptable quality.

[0160] The present disclosure further provides control mechanisms to dynamically adjust HARQ and MCS error-handling strategies based on inference-quality metrics. By selectively allowing certain errors, the system reduces delay for time-critical inference tasks. Building on the above-noted concepts like I-QoS framework / profile and dynamic quality-driven optimization, this approach provides a standardized way to incorporate inference metrics into HARQ / MCS control at the PHY / MAC layers.

[0161] Previous technologies, as documented in standard 3GPP releases and related literature, focus on stringent reliability and low latency (e.g., URLLC) without considering inference-level needs. HARQ retransmissions are always triggered upon errors, aiming for near-perfect delivery. While link adaptation and scheduling optimizations exist, they do not incorporate inference-quality metrics that may justify occasionally ignoring HARQ retransmissions to gain latency advantages at the cost of tolerable error.

[0162] The absence of inference-aware QoS management leads to inefficiencies –either over-protecting the link when minor errors are acceptable or failing to improve conditions when inference metrics degrade. No known published solution integrates inference metrics into QoS decision-making in real-time. Consequently, current protocols may not differentiate between inference tasks that may tolerate minor errors and those requiring strict accuracy, resulting in potentially suboptimal latency and resource usage for AI-driven inference workloads.

[0163] Consequently, existing solutions have the following characteristics: ● Strict adherence to near-perfect reliability even when minor errors may be acceptable. ● Lack of protocol elements (RRC IEs, MAC CEs, DCI fields) designed for inference-oriented HARQ and QoS  negotiation. ● Inability to dynamically adjust transmission parameters (including HARQ ignorance and MCS parameters) based  on real-time inference-quality metric evaluations. ● No support for per-task customizable HARQ / MCS strategies aligned with evolving AI inference workloads.

[0164] The above characteristics of the existing solutions may result in at least the following disadvantages: ● Excessive retransmissions and latency due to striving for near-perfect reliability. ● High energy consumption that negates the potential benefits of distributed HAPU deployments. ● Lack of adaptive, inference-driven error tolerance to optimize performance and cost.

[0165] The present disclosure addresses the following technical problems: ● Integrating distributed HAPU resources, including HAPU-equipped UEs (computing UEs) , to form inference  pipelines, where requesting UEs request inference tasks and other UEs serve as Computing UEs providing collaborative inference computation, and the BS manages the inference pipelines. ● Ensuring that wireless transmission overhead (latency, energy, bandwidth usage) remains lower than the avoided  central cooling and special power infrastructure costs. ● Inference-Centric QoS Integration: Incorporating inference-quality metrics into QoS frameworks to dynamically  set and adjust error tolerance. Utilizing inference-level quality metrics like Perplexity, which may be measured periodically, to maintain proper configuration in changing wireless conditions. ● Per-Task Customization: Each inference task may require distinct metrics and thresholds. The invention  described in the present disclosure allows defining separate I-QoS profiles per task type. ● Dynamically adjusting communication reliability based on task-specific error tolerance requirements.  Continuously refining HARQ / MCS settings, and resource allocation policies based on periodic inference metric checks and channel conditions. ● Enabling protocols and standard signaling that integrate inference-quality metrics into QoS negotiation, HARQ  ignorance adjustments, and dynamic PHY / MAC layer adaptations. ● Efficiency and Scalability: By allowing acceptable minor errors, reducing unnecessary retransmissions, and / or  adjusting frequency of metric checks, the system lowers overhead and meets evolving AI inference demands. ● Latency Reduction: Achieve lower latency for time-critical inference pipelines by minimizing unnecessary  retransmissions. ● Ensuring compatibility with future networks (e.g., 6G network) , allowing flexible task-specific configurations,  changing resource allocation, and keeping error tolerance aligned with evolving AI inference requirements.

[0166] The present disclosure proposes a set of protocol enhancements and procedures for future wireless network systems (e.g., 6G network) , based on existing wireless network frameworks (e.g., 5G NR frameworks) , to support Error-Tolerant Inference scenarios. Aspects of the present disclosure include: ● UE Role Differentiation: Some UEs act as Computing UEs (with HAPUs) providing inference services to  requesting UEs. The computing UEs and the requesting UEs may be orchestrated by the BS. This facilitates distributed inference computation and results in standard cooling / power conditions. ● I-QoS Profiles: A new QoS class integrated with inference metrics, specifying fields like Metric_Type,  Metric_Interval, Baseline_Error_Tolerance, Metric_Threshold, and Adjustment_Policies. ● Adaptive Error Tolerance: Allowing controlled targeted BER or packet loss that are tailored to each inference  task's sensitivity. ● Hybrid automatic repeat request (HARQ) Adaptations: Introducing HARQ ignorance rates or HARQ-free modes  (which may be activated when permissible) . ● Introducing MCS adjustment approaches that cooperate with the HARQ ignorance mechanisms to achieve error  acceptance. ● Dynamic Quality-Driven Optimization: Using periodic metric measurements to guide HARQ, MCS, and  scheduling adjustments. This ensures that latency and reliability remain aligned with actual inference needs. ● Task-Specific and Channel-Aware Policies: The T-TRP correlates channel conditions with inference quality- metric results to anticipate performance changes, thereby ensuring stable and efficient distributed inference pipelines. Different I-QoS profiles yield distinct HARQ / MCS behaviors, which may be tailored to each inference scenario. ● Channel and Priority Awareness: Error-control mechanisms also account for channel conditions and task priority,  thereby ensuring robust and responsive adjustments. There are procedures defined for dynamic QoS negotiation that consider MCS adjustments, to improve the effect of HARQ ignorance control on resource prioritization. ● Standardized Signaling: New fields in RRC, MAC, and DCI messages enable a unified, extensible approach for  vendors and operators. ● Providing standardized DCI formats that allow MCS overrides or adjustments in real time to thereby support the  real-time HARQ ignorance override mechanism.

[0167] In some implementations, the I-QoS profile may be included in first information that may be transmitted between a network device (e.g., BS, T-TRP) , computing devices (e.g., computing UEs) , and a requesting device (e.g., requesting UE) . The fields (e.g., one or more parameters associated with inference quality) included in the I-QoS profile (or first information) may indicate at least one of: a metric used for evaluation of the inference quality of the inference task, an inference quality evaluation interval for the inference task, a transmission error rate to be satisfied during the inference task; a metric threshold value used for the evaluation of the inference quality of the inference task; or one or more adjustment policies. Put another way, the first information may configure one or more parameters associated with inference quality of the interference task.

[0168] The present disclosure pertains to a 6G RAN environment where multiple EDs (UEs) exist. Some UEs request inference services, and these UEs may be referred to as “Requesting UE” . Some UEs are hosting HAPU resources, and operate under cellular network, form inference pipelines that compute partial computations of the AI model inference. These UEs may be referred to as “Computing UE” . The BS (T-TRP) acts as a coordinator, assigns inference tasks, and routes intermediate data between Computing UEs.

[0169] With distributed HAPUs, each Computing UE is thermally manageable and uses standard power lines for its power supply. Intermediate layers of the model are transmitted over wireless links with controlled error tolerance.

[0170] For example, if a UE requests an LLM-based inference, the T-TRP distributes computation (s) across multiple Computing UEs. The T-TRP obtains the error-tolerance for computing UEs based on their pipeline’s task and channel conditions, and communicates the obtained error-tolerance via RRC / MAC signaling. Periodically, the T-TRP triggers a known test input, measures the inference quality metric (e.g., Perplexity) , and adjusts HARQ and modulation / coding scheme (MCS) strategies if the quality does not meet task requirements. This ensures that the overhead remains balanced, and the PUE advantages are retained.

[0171] In the present disclosure, the “known test input” may include a predefined input of the inference task, a predefined input of the inference quality test (or evaluation) , or the likes. In some implementations, an output of the inference task that is performed with the predefined input may be predefined. In some implementation, the inference quality of the inference task may be evaluated using the predefined input and / or the predefined error-free output of the inference task. The evaluation of the inference quality may be referred to as inference quality check, inference quality test, inference quality metric measurement, or other similar expressions.

[0172] Each task has a specific Inference-QoS (I-QoS) profile containing the information about the task requirements and corresponding parameters. The I-QoS profile may include one or more of the following items: ● Metric_Type: Specifies which inference-quality metric to use (e.g., Perplexity, Accuracy, Δ-Probability,  Kullback–Leibler (KL) ) . ● Metric_Interval: How often the T-TRP triggers a test input for metric measurement. ● Metric_Thresholds: Defines acceptable metric ranges or limits. ● Baseline_Error_Tolerance: Initial acceptable code block error ratio or similar tolerance. ● Adjustment_Policies: Instructions on how to alter HARQ, MCS / coding, and resource allocations when metrics  deviate from thresholds. ● Task_Class_Indicator: Associates the I-QoS profile with a particular inference task type (e.g., LLM_QA,  Object_Detection) . ● Resource_Priority_Level: Indicates priority handling among concurrent tasks.

[0173] In the present disclosure, in some implementations: ● “Metric_Type” or similar expressions may be considered or referred to as a metric used for evaluation of the  inference quality of the inference task; ● “Metric_Interval” or similar expressions may be considered or referred to as an inference quality evaluation  interval for the inference task; ● “Metric_Thresholds” or similar expressions may be considered or referred to as a metric threshold value used for  the evaluation of the inference quality of the inference task; ● “Baseline_Error_Tolerance” or similar expressions may be considered or referred to as a transmission error rate  to be satisfied during the inference task; ● “Adjustment_Policies” or similar expressions may be considered or referred to as one or more adjustment  policies, for parameter adjustment or resource allocation, based on the evaluation of the inference quality of the inference task; ● “Task_Class_Indicator” or similar expressions may be considered or referred to as a type of the inference task;  and ● “Resource_Priority_Level” or similar expressions may be considered or referred to as priority of the inference  task.

[0174] In the present disclosure, in some implementations: ● “Perplexity” may be considered or referred to as a perplexity metric measuring uncertainty associated with the  inference task; ● “Accuracy” may be considered or referred to as an accuracy metric based on accuracy associated with the  inference task; ● “Δ-Probability” may be considered or referred to as a metric based on Δ-probability associated with the inference  task; and ● “Kullback–Leibler (KL) ” may be considered or referred to as a metric based on KL divergence associated with  the inference task.

[0175] The T-TRP uses the I-QoS profile to obtain and manage the quality of inference and provide the corresponding guidelines for acceptable error tolerance levels, metric intervals, HARQ ignorance and MCS strategies.

[0176] This scenario is applicable to distributed AI inference in RAN environments, including non-terrestrial and heterogeneous networks, enabling low-latency inferencing for robotics, autonomous driving, and real-time analytics tasks.

[0177] FIG. 6 illustrates, in a schematic diagram, an example wireless network 600 in which an inference pipeline 610 is deployed, in accordance with implementations of the present disclosure. The wireless network 600 includes a T-TRP 601, a computing UE_B 612, and a computing UE_C 614.

[0178] The T-TRP 601 may act as a coordinator and assign an inference task to the inference pipeline 610 which includes the computing UE_B 612 and the computing UE_C 614. In other words, the T-TRP 601 may coordinate the inference task performed at the inference pipeline 610, and manage the computing UE_B 612 and the computing UE_C 614. The computing UE_B 612 and the computing UE_C 614 may collectively perform the inference task that is assigned to them. Put another way, each of the computing UE_B 612 and the computing UE_C 614 may perform a respective portion of the inference task assigned thereto.

[0179] The T-TRP 601 may use the I-QoS profile to obtain and manage the quality of inference and provide the corresponding guidelines for acceptable error tolerance levels, metric intervals, HARQ ignorance and MCS strategies. A non-limiting example fields and / or parameters included in the I-QoS profile are shown in the RRC I-QoS configuration frame 650. For the purpose of illustration, in the present disclosure, “I-QoS profile” , “I-QoS framework” , “I-QoS configuration frame” and / or other similar expressions may be interchangeably used in the present disclosure.

[0180] The I-QoS profile 650 may configure one or more parameters associated with inference quality of the interference task. As shown in FIG. 6, the RRC I-QoS configuration frame 650 may include fields associated with (or parameters associated with) message type, computing device identifier (e.g., UE-ID) , task class indicator, metric type, metric interval, baseline error tolerance, metric threshold, adjustment policies, resource priority level, and / or padding bits.

[0181] In FIG. 6, the task class indicator (or other similar parameter indicating a task type) included in the I-QoS profile 650 indicates that “LLM Q&A” is the type of the inference task to be performed by the computing UEs 612 and 614 in the pipeline 610. This indication is forwarded to the computing devices in the pipeline 610 as illustrated in FIG. 6. Specifically, the T-TRP 601 may forward this task class indicator to the computing UE_B 612, which in turn forward the task class indicator to the computing UE_C 614. The computing UE_C 614 may forward this indication back to the T-TRP 601, for example to indicate that the “LLM Q&A” type of inference task is performed.

[0182] In accordance with the information (e.g., fields or parameters) included in the RRC I-QoS configuration frame 650, evaluation of the inference quality of the inference task (e.g., a quality metric test) may be performed. For example, intermediate inference data may be exchanged between the devices in accordance with the information included in the I-QoS configuration frame 650. After the evaluation of the inference quality of the inference task is completed, the T-TRP 601 may adjust, for example, at least one parameter associated with a HARQ protocol, and / or at least one parameter associated with a MCS (e.g., adjust necessary HARQ / MCS settings) . The evaluation of the inference task and HARQ / MCS setting adjustment may be performed periodically as illustrated in FIG. 6.

[0183] FIG. 7 illustrates, in a schematic diagram, an example wireless network 700 in which an inference pipeline 710 and an inference pipeline 720 are deployed, in accordance with implementations of the present disclosure. The wireless network 700 includes a T-TRP 701, a computing UE_B 712, a computing UE_C 714, and a computing UE_D 722, and a computing UE_E 724.

[0184] The inference pipeline 710 includes the computing UE_B 712 and the computing UE_C 714. The computing UE_B 712 and the computing UE_C 714 may collectively perform the inference task that is assigned to the inference pipeline 710. Similarly, the inference pipeline 720 includes the computing UE_D 722 and the computing UE_E 724. The computing UE_D 722 and the computing UE_E 724 may collectively perform the inference task that is assigned to the inference pipeline 720.

[0185] The T-TRP 701 may transmit, to each of the computing UEs 712, 714, 722, 724, the I-QoS profile to obtain and manage the quality of inference. The T-TRP 701 may also provide, to each of the computing UEs 712, 714, 722, 724, the corresponding guidelines for acceptable error tolerance levels, metric intervals, HARQ ignorance and MCS strategies. The I-QoS profile, the corresponding guidelines (e.g., information relating to adjusted one or more parameters associated with HARQ / MCS) , and / or other similar information may be transmitted to each of the computing UEs 712, 714, 722, 724 via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) , as illustrated in FIG. 7 or elsewhere in the present disclosure.

[0186] The T-TRP 701 may be similar to the T-TRP 601 illustrated above or elsewhere in the present disclosure, and the computing UE_B 712, the computing UE_C 714, the computing UE_D 722, and the computing UE_E 724 may be similar to the computing UE_B 612 and the computing UE_C 614. Accordingly, additional details regarding the T-TRP 701, and the computing device UE_B 712, the computing UE_C 714, and the computing UE_D 722, and the computing UE_E 724 are omitted here.

[0187] Example devices which apply the method illustrated in the present disclosure may include at least one of the followings. ● Enhanced BS / gNB Products: The enhanced BS / gNB products may feature software-defined HARQ adaptations  and inference-quality-based parameter setting; integrate I-QoS logic and dynamic adaptation capabilities; and / or implement new RRC IEs, MAC CEs, and DCI field handling for I-QoS-driven HARQ Ignorance and MCS adjustments. ● HAPU-Equipped UEs (Computing UEs or edge nodes) : Premium devices or industrial nodes that offer inference  layers, form pipelines (and / or parallel computing collaboration) with other computing UEs, and / or support dynamic HARQ / MCS changes and report metric test results. ● Software Solutions for RAN / Edge Controllers: The software solutions for RAN / Edge controllers may manage  error tolerance, periodic quality metric (perplexity or other metrics) checks, and dynamic resource allocation.

[0188] The following disclosure describes details of the method in a wireless-based distributed inference system that leverages multiple Computing UEs with HAPUs to perform partial inference tasks. The system employs a novel Inference-QoS (I-QoS) framework and methods for dynamically optimizing transmission parameters (HARQ, MCS, resource allocation) based on inference-quality metrics. The I-QoS framework extends beyond 5G QoS classes by introducing fields and logic that reflect the unique requirements of distributed AI inference tasks. While Perplexity is a prime example metric for linguistic AI tasks, I-QoS may flexibly incorporate various metrics depending on task types and domains. The I-QoS framework may show how inference-quality metrics may drive both HARQ-related decisions and Modulation and Coding Scheme (MCS) adjustments cooperatively, offering controllable tolerance for transmission errors and potentially reducing latency and resource overhead. By enabling selective HARQ retransmission control-referred to as HARQ ignorance-and complementary MCS adaptations, these implementations pave the way for future network system enhancements that consider inference-level quality indicators.

[0189] In one implementation, a scenario of distributed Inference Pipeline is provided.

[0190] The T-TRP sends an RRC signaling (e.g., RRC Reconfiguration message) with an “I-QoS Profile” that includes an “Error Tolerance Parameter” (e.g., 1%code block error) and / or a “Throughput / Latency Tolerance Parameter” (e.g., 0.5 millisecond latency) .

[0191] FIG. 8 illustrates, in a schematic diagram, an example wireless network 800 in which an example pipeline that includes computing devices and a requesting device, in accordance with implementations of the present disclosure. The wireless network 800 includes a T-TRP 801, a requesting UE_A 803, a computing UE_B 812, and a computing UE_C 814.

[0192] The T-TRP 801 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UE_B 812 and the computing UE_C 814 may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 801, and the computing device UE_B 812 and the computing UE_C 814 are omitted here.

[0193] The requesting UE_A 803 may transmit, to the T-TRP 801, an inference task request along with other information (e.g., part or all of the I-QoS profile illustrated above or elsewhere in the present disclosure, “Model-Inference-Request” profile illustrated below or elsewhere in the present disclosure) . The requesting UE_A 803 may receive an output of the inference task, for example, from the last computing device (e.g., UE_C 814) in the pipeline associated with the inference task.

[0194] FIG. 8 shows an example pipeline forming: UE_A 803 → UE_B 812 → UE_C 814 → UE_A 803, exchanging intermediate inference data wirelessly.

[0195] The computing UE_B 812 and the computing UE_C 814 may collectively perform the inference task that is assigned to them. The requesting UE_A 803 may transmit an inference task request to the T-TRP 801. The T-TRP 801 may trigger the computing UE_B 812 and the computing UE_C 814 so that these computing UEs may collectively perform the inference task assigned to them. In some implementations, the computing UE_B 812 and the computing UE_C 814 may be considered or included in a respective inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task. The T-TRP 801 may receive the output of the inference task from the last computing UE (UE_C 814) in the pipeline, and may transmit the output to the requesting UE_A 803.

[0196] The T-TRP 801 configures HARQ ignorance rates via signaling (e.g., MAC CEs) so that minor errors do not trigger retransmissions. The T-TRP 801 injects a known test input, measures an inference-quality metric (e.g., Perplexity) . If the metric degrades, the T-TRP 801 tightens error tolerance or select more robust MCS settings to restore quality. This ensures that wireless overhead stays below the cost of centralized cooling / power infrastructures. In some implementations, the T-TRP 801 periodically measures the inference-quality metric.

[0197] Put another way, in some implementations, the T-TRP 801 may evaluate the inference quality of the inference task (e.g., inference quality metric check) . For example, the T-TRP 801 may set up the communication error tolerances (e.g., a transmission error rate to be satisfied during the inference task) , and regularly check the quality metrics. Based on the evaluation of inference quality metrics, the T-TRP 801 may adjust one or more parameters related to communication error tolerances. For example, the T-TRP 801 may adjust one or more parameters associated with the MCS to restore the inference quality.

[0198] FIG. 9 illustrates, in a schematic diagram, an example wireless network 900 in which an example pipeline that includes computing device groups and a requesting device, in accordance with implementations of the present disclosure. The wireless network 900 includes a T-TRP 901, a requesting UE_A 903, a first computing device group (UE group) 910 that includes a UE_B 912, and a computing UE_C 914, and a second computing device group (UE group) 920 that includes a UE_B 922, and a computing UE_C 924.

[0199] The T-TRP 901 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UE_B 912, the computing UE_C 914, and the computing UE_B 922, and the computing UE_C 924 may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Further, the requesting UE_A 903 may be similar to the requesting UEs (e.g., UE_A 803) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 901, the computing UEs 912, 914, 922, 924, and the requesting UE_A 903 are omitted here.

[0200] FIG. 9 shows another example pipeline forming: UE_A 903 → first UE group 910 (UE_B 912 and UE_C 914) → second UE group 920 (UE_B 922 and UE_C 924) → UE_A 903. As shown in FIG. 9, the computing UEs may form different inference cluster topologies. For example, the model may form a pipeline between groups of computing UEs, and inside each group, parallel computations occur between collaborating computing UEs.

[0201] More specifically, as noted above, the pipeline includes the first computing UE group 910 and the second computing UE group 920. The first UE group 910 may include the computing UE_B 912 and the computing UE_C 914. The computing UE_B 912 and the computing UE_C 914 perform inference computations in parallel, for the part of the inference task (or the inference task) assigned to the first UE group 910. In some implementations, the computing UE_B 912 and the computing UE_C 914 may be considered or included in the first inference layer 911 of the AI / ML model associated with the inference task. Similarly, the second UE group 920 may include the computing UE_B 922 and the computing UE_C 924. The computing UE_B 922 and the computing UE_C 924 perform inference computations in parallel, for the part of the inference task (or the inference task) assigned to the second UE group 920. In some implementations, the computing UE_B 922 and the computing UE_C 924 may be considered or included in the second inference layer 921 of the AI / ML model associated with the inference task.

[0202] In any case, the protocols for Error Tolerant inference apply to computing UEs regardless of the topology of inference. Furthermore, the connection between the computing nodes or computing groups 910 and 920 may be either through the base-station 901 or point-to-point (e.g., via sidelink) .

[0203] Put another way, in some implementations, the T-TRP 901 may transmit, to each of the computing UEs 912, 914, 922, 924 in the first and second computing UE groups 910 and 920, the I-QoS profile and / or intermediate data associated with the inference task or inference computation. The I-QoS profile may include, for example, communication error tolerance parameters (e.g., a transmission error rate to be satisfied during the inference task) . In some implementations, different I-QoS profiles may be transmitted to the first and second computing UE groups 910 and 920, respectively. In other words, the I-QoS profile transmitted to the computing UEs 912, 914 in the first computing UE group 910 may be different from that transmitted to the computing UEs 922, 924 in the second computing UE group 920.

[0204] In some other implementations, while not explicitly described in FIG. 9, transmission of the I-QoS profile may involve sidelink communication. For example, the T-TRP 901 may transmit a first I-QoS profile to the computing UE 912 (or computing UE 914) , and the computing UE 912 (or computing UE 914) may forward, via sidelink, the first I-QoS profile to the computing UE 914 (or computing UE 912) . Similarly, the T-TRP 901 may transmit a second I-QoS profile to the computing UE 922 (or computing UE 924) , and the computing UE 922 (or computing UE 924) may forward, via sidelink, the second I-QoS profile to the computing UE 924 (or computing UE 922) .

[0205] In another implementation, a protocol-level integration for task and quality indicators is provided. As shown in FIG. 10, it may include at least one of the following steps.

[0206] The steps for the protocol-level integration for task and quality indicators will be illustrated with reference to FIG. 10, which illustrates, in a signaling diagram, an example process 1000 for inference pipeline initialization and inference quality management, in accordance with implementations of the present disclosure.

[0207] The wireless network system in which the process 1000 is performed may include a T-TRP 1001, a computing UE_B 1002b and a computing UE_C 1002c (collectively referred to as 1002) , and a requesting UE_A 1003. The T-TRP 1001 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UEs 1002 may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Further, the requesting UE_A 1003 may be similar to the requesting UEs (e.g., UE_A 803, UE_A 903) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 1001, the computing UEs 1002, and the requesting UE_A 1003 are omitted here.

[0208] The requesting UE (e.g., UE_A 1003 in FIG. 10) , at 1010, sends a “Model-Inference-Request” (e.g., non-access stratum (NAS) information element (IE) : I-QoS) to the T-TRP 1001. This profile may include: ● A “Task Class Indicator” (e.g., LLM_QA or Object_Detection) defining the inference scenario. ● An “Inference Quality Indicator Metric” (e.g., Δ-perplexity) specifying the chosen quality evaluation metric to  be used for evaluation of the inference accuracy. ● An “Inference Quality Threshold Value” (e.g., 3%relative Δ-perplexity) specifying acceptable inference quality  value. ● A “Throughput / Latency Indicator” (e.g., 50 generated token per second) specifying acceptable overall  throughput. ● An “Inference Quality Measurement Interval” defining how frequently (e.g., every N frames) a known test input  is used to verify inference quality.

[0209] At least in some implementations, “Task Class Indicator” , “Inference Quality Indicator Metric” , “Inference Quality Threshold Value” , and / or “Inference Quality Measurement Interval” included in the above-noted “Model-Inference-Request” profile may be similar to “Task_Class_Indicator” , “Metric_Type” , “Metric_Thresholds” , and / or “Metric_Interval” illustrated above or elsewhere in the present disclosure.

[0210] In the present disclosure, “Throughput / Latency Indicator” included in the above-noted “Model-Inference-Request” profile or similar expressions may be considered or referred to as a throughput or latency range to be satisfied during the inference task.

[0211] Upon receiving the inference request, the T-TRP 1001, at 1015, sends an RRC signalling including, for example, “Computing-UE-Inference-Profile” to the involved computing UEs (for example, UE_B 1002b and UE_C 1002c in FIG. 10) . This profile may include: ● A “Channel Quality Indicator” specifying acceptable error ranges for computing UE’s communication channel  (e.g., 0.5%code block error) . ● A “Channel Throughput / Latency Indicator” specifying acceptable latency ranges for computing UE’s  communication channel (e.g., 0.5msec latency) .

[0212] At least in some implementations, “Channel Quality Indicator” included in the above-noted “Computing-UE-Inference-Profile” may be similar to “Baseline_Error_Tolerance” , illustrated above or elsewhere in the present disclosure. In some implementations, “Channel Throughput / Latency Indicator” included in the above-noted “Computing-UE-Inference-Profile” or similar expressions may be considered or referred to as a throughput or latency range to be satisfied during the inference task.

[0213] The computing UEs 1002 may interpret these parameters to adjust HARQ and MCS settings.

[0214] In other words, in some implementations, the computing UEs 1002 may use the parameters included in the above-note “Computing-UE-Inference-Profile” to determine whether to adjust HARQ and MCS settings.

[0215] At 1020, there may be N frames of inference task communications between the T-TRP 1001, the computing UEs 1002, and / or the requesting UE_A 1003. Some aspects of the inference task communications are illustrated above or elsewhere in the present disclosure.

[0216] After N frames, the T-TRP 1001, at 1025, sends a “MAC CE Inference-Quality-Test-Trigger” instructing the pipeline to process a known test input.

[0217] In some implementations, the known test input may be included in the “MAC CE Inference-Quality-Test-Trigger” transmitted from the T-TRP 1001 to the computing UE_B 1002b. In some implementations, the “MAC CE Inference-Quality-Test-Trigger” may be considered a request for the inference task (or inference quality test) . The computing UE_B 1002b may perform a portion of the inference quality test using the known test input received from the T-TRP 1001.

[0218] In some implementations, the “MAC CE Inference-Quality-Test-Trigger” may be generated based on “Model-Inference-Request” that the T-TRP 1001 received from the requesting UE_A 1003. More generally, the request for the inference task may be generated based on the inference task request that the T-TRP 1001 received from the requesting UE_A 1003.

[0219] At 1030, the test inference data may be transmitted through the pipeline. For example, at 1030, the computing UE_B 1002b may transmit, to the computing UE_C 1002b, the test inference data (e.g., intermediate inference data, an output of a portion of the inference test performed by the UE_B 1002b) . The computing UE_C 1002c may perform another portion of the inference quality test using the test inference data received from the computing UE_B 1002b. The other portion of the inference quality test performed by the computing UE_C 1002c may output a final inference result.

[0220] The T-TRP 1001, at 1035, retrieves the final inference result, and at 1040, compares the final inference result against a baseline. If quality worsens, the T-TRP 1001 lowers error tolerance or selects more robust MCS / forward error correction (FEC) codes. The final result, at 1035, is returned to the T-TRP 1001. The T-TRP 1001, at 1040, compares the chosen metric (e.g., Perplexity) against a baseline.

[0221] Then, the T-TRP 1001, at 1045, issues another MAC CE to lower error tolerance and possibly shift to a more robust MCS / FEC code, thereby ensuring inference quality remains acceptable.

[0222] For example, the T-TRP 1001, at 1045, may transmit a revised MAC CE that adjust HARQ / MCS settings (e.g., a revised MAC CE that adjust one or more parameters associated with HARQ / MCS) . In this way, the inference quality of the inference task performed by the computing UEs 1002 may remain acceptable.

[0223] In some implementations of the present disclosure, one or more parameters associated with HARQ, one or more parameters associated with MCS, or other similar expressions may be considered non-limiting examples of communication parameters used for transmission of data associated with the interference task. In some implementations of the present disclosure, the one parameter associated with the HARQ (or HARQ protocol) may include a HARQ ignorance rate and / or an override ignorance flag. The HARQ ignorance rate may indicate a portion that HARQ retransmission is deactivated, and the override ignorance flag may indicate that the HARQ retransmission is fully activated regardless of the HARQ ignorance rate. In some implementations, the HARQ ignorance rate may be referred to as HARQ Ignorance or Ignorance or other similar expressions, and the override ignorance flag may be referred to as “HARQ Ignorance Override” flag or other similar expressions.

[0224] This protocol-level signaling ensures that each task’s characteristics and acceptable error margin are explicitly defined and dynamically enforced, and accordingly a favorable balance between energy / latency savings and inference quality may be maintained.

[0225] As noted above or elsewhere in the present disclosure, the FIG. 10 shows a sequence diagram of the signaling involving the request of an inference task by requesting UE (UE_A) 1003 and signaling of pipeline initialization and quality management involving Computing UEs 1002 (UE_B 1002b and UE_C 1002c) .

[0226] In some inference tasks, the quality requirements of the task may change in real-time during the same task. For example, in the same Q&A LLM task and during the same session, the required inference quality to answer a hard question is generally much higher than when outputting the tokens of grammar structure word (e.g., conjunction or proposition words) . Therefore, the Error-Tolerant Inference system may employ methods for real-time update of the I-QoS profile, inference quality threshold, and corresponding error-control adjustments. In some implementations, the real-time inference quality thresholds may be updated by the requesting UE’s application layer or by some other measures. For example, an autonomous vehicle may increase the quality requirement threshold for object detection inference after sunset or during extreme weather conditions with poor visuals.

[0227] In another implementation, a scheme for varying tolerance by task, time, and HAPU UE roles is provided as shown in FIG. 11.

[0228] FIG. 11 illustrates, in a schematic diagram, an example wireless network 1100 with inference pipelines 1110 and 1120 in which different inference tasks are respectively assigned, in accordance with implementations of the present disclosure. The wireless network 1100 includes a T-TRP 1101, a computing UE_B 1112, a computing UE_C 1114, and a computing UE_D 1122, and a computing UE_E 1124.

[0229] The T-TRP 1101 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UE_B 1112, the computing UE_C 1114, the computing UE_D 1122, and the computing UE_E 1124 may be similar to may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 1101, and the computing UEs 1112, 1114, 1122, and 1124 are omitted here.

[0230] In a distributed inference environment, multiple inference tasks with different accuracy and latency requirements may coexist. For example, as shown in FIG. 11, one pipeline (e.g., Pipeline 1110 in FIG. 11 where UE_B 1112 and UE_C 1114 are involved as computing UEs) might handle an LLM question and answer (Q&A) task (which may tolerate a slightly higher error tolerance in intermediate layers) , while another pipeline (e.g., Pipeline 1120 in FIG. 11 where UE_D 1122 and UE_E 1124 are involved as computing UEs) might run a safety-critical object detection model (requiring stricter error control) . The T-TRP 1101 orchestrates several computing UEs 1112, 1114, 1122, 1124 with HAPUs to serve these distinct tasks concurrently.

[0231] When the T-TRP 1101 receives multiple inference requests, it assigns each task to the most suitable computing UEs. For each task, a distinct “I-QoS” profile is provided via, for example, RRC signaling. The “I-QoS” profile for each task may be completely different and separate from the other task pipelines that are simultaneously governed by the T-TRP 1101. For instance, one pipeline 1110 may be running an LLM Q&A task 1115 while another pipeline 1120 running object detection 1125. The quality metrics and other error-control parameters of I-QoS profile are also separate for each task. For example, for LLM Q&A task 1115, the quality metric error (rate) may be less than 5%and the latency may be less than 1 millisecond (ms) , for object detection task 1125, the quality metric error (rate) may be less than 0.05%and the latency may be less than 0.1 ms. These indicators may define allowable error rates (e.g., Comm. Error in FIG. 11) . For instance, the LLM Q&A task 1115 might allow a 0.1%code block error ratio, while the object detection task 1125 might only accept 0.02%. Each of these tasks 1115 and 1125 may be measured (for example, the inference-quality measurement procedures described above) using a chosen inference-quality metric (e.g., Perplexity for language tasks, or a custom metric for vision tasks) to ensure the error tolerance levels remain appropriate over time.

[0232] After each measurement interval, if the Perplexity (or another suitable metric) for the LLM Q&A pipeline 1110 remains stable, the T-TRP 1101 keeps the current HARQ ignorance settings. For the object detection pipeline 1120, if even a small change in the chosen quality metric occurs that indicates potential inference degradation, the T-TRP 1101 immediately sends a MAC CE to tighten error tolerance, reducing acceptable error rates. This per-task adaptability ensures that each inference workload benefits optimally from distributed GPU deployments without sacrificing necessary accuracy.

[0233] As conditions change (e.g., at peak traffic hours, wireless conditions may worsen) , the T-TRP 1101 may find that the LLM Q&A task 1115 may also reduce its error tolerance to maintain acceptable Perplexity. Conversely, during off-peak hours with stable channels, more relaxed error tolerance might be restored. By combining periodic metric measurements with per-task configurations, the system dynamically maintains the right balance between energy / latency savings and inference performance across a diverse set of tasks (e.g., LLM Q&A task 1115 and object detection task 1125) and pipelines (e.g., LLM Q&A pipeline 1110 and object detection pipeline 1120) .

[0234] In some implementations, The T-TRP 1101 may maintain a lookup table correlating different tasks with baseline Perplexity (or other metrics) thresholds. If tasks are added or removed, the system updates the lookup table accordingly.

[0235] Some tasks may require more frequent checks due to sensitivity, while others might reduce check frequency under stable conditions.

[0236] For example, the inference quality check / evaluation may be more frequently carried out for the object detection task 1125, for example, due to sensitivity. On the other hand, the inference quality check / evaluation may be less frequently carried out under stable conditions for the LLM Q&A task 1115.

[0237] In some implementations, in cases where processing efficiency requires or there is limited number of computing UEs, different pipelines may share some computing UEs, e.g., some of the computing UEs may be part of multiple different pipelines with different task classes.

[0238] In one implementation, other types of inference metrics and HARQ groupings are described as shown in FIGs. 12A and 12B.

[0239] FIGs. 12A and 12B illustrate an example method 1200 of inference quality management for multiple inference tasks, in accordance with implementations of the present disclosure.

[0240] The wireless network system in which the method 1200 is performed may include a T-TRP 1201, a computing UE_B 1202b, a computing UE_C 1202c, a computing UE_D 1202d a computing UE_E 1202e (collectively referred to as 1202) , and a verification database 1203. The T-TRP 1201 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UEs 1202 may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 1201 and the computing UEs 1202 are omitted here. The verification database 1203 may be configured to provide datasets to be used for the inference quality tests, for example, upon request of the T-TRP 1201.

[0241] Different inference workloads might require different metrics. Some tasks may use Δ-probability difference (comparing predicted vs. baseline probabilities) , top-k accuracy, or domain-specific metrics (e.g., mean average precision for object detection) . The T-TRP 1201, by design, is flexible and may accommodate these alternative inference-quality metrics. Furthermore, tasks may be grouped into HARQ Groups, where each group is associated with a particular quality requirement profile.

[0242] In an example scenario, FIGs. 12A and 12B show an example for two pipelines with T-TRP 1201 that defines two HARQ Groups. The first HARQ Group 1 is for language-based tasks using Perplexity as the metric. This group might allow a certain range of ITEL values for error tolerance, associated with a certain coding / modulation scheme. The second HARQ Group 2 is for vision-based tasks using accuracy-based or Δ-probability metrics. This group might enforce stricter error controls initially due to sensitivity to noise in intermediate feature maps, associated with stricter coding scheme, lower-order modulation.

[0243] As FIGs. 12A and 12B show, the T-TRP 1201 first transmits, at 1210, to the verification database 1203, a request for the datasets to be used for the inference quality tests. The T-TRP 1201 obtains, at 1215, the perplexity corpus dataset and also obtains, at 1220, the object detection datasets required for performing the inference quality tests. The Perplexity corpus may be any predefined text corpus that the T-TRP 1201 injects into an example LLM pipeline and calculates the perplexity for that corpus. The T-TRP 1201 then uses its predefined guidelines to check if the measured perplexity is in the acceptable range. Similarly for the Object Detection task, a predefined dataset may be used by the T-TRP 1201 to check the accuracy of a set of test object detection tasks.

[0244] At periodic intervals, the T-TRP 1201 triggers metric checks for each HARQ Group. For the first HARQ Group 1, a Perplexity test is run on a known linguistic test input. For the second HARQ Group 2, a known test image or sensor dataset is run through the pipeline to compute the chosen accuracy metric.

[0245] More specifically, the T-TRP 1201, at 1225, may prepare quality metric test (e.g., perplexity test) using the Perplexity corpus obtained from the database 1203. The T-TRP 1201, at 1230, may send an “inference quality metric test trigger” to the computing UE_B 1202b which may perform a portion of the inference quality test (e.g., Perplexity test) . After the portion of the inference quality test, at 1235, the computing UE_B 1202b may transmit the intermediate inference data (e.g., an output of a portion of the inference quality test performed by the UE_B 1202b) to the computing UE_C 1202b, which may provide a final output of the inference quality test. At 1240, the computing UE_C may transmit, to the T-TRP 1201, the final output of the inference quality test. Then, the T-TRP 1201, at 1245, may compare the final output of the inference quality test against a metric threshold value (e.g., baseline Perplexity, a perplexity threshold value) .

[0246] It is noted that steps 1230-1245 may be similar to steps 1025-1040 illustrated above or elsewhere in the present disclosure. Accordingly, additional details regarding these steps are omitted here.

[0247] Based on the comparison, the T-TRP 1201 may determine as to whether it will adjust one or more parameters associated with HARQ / MCS.

[0248] If the first HARQ Group 1’s Perplexity stays stable, no changes 1250 occur.

[0249] On the other hand, if the final output indicates that the inference quality worsens or the first HARQ Group 1’s Perplexity is unstable, then the T-TRP 1201, at 1255, may adjust one or more parameters associated with HARQ / MCS for improved reliability.

[0250] Steps for the second HARQ Group 2 may be carried out in a similar manner. Specifically, steps 1260-1285 may be similar to steps 1225-1250 illustrated above or elsewhere in the present disclosure. Accordingly, additional details regarding these steps are omitted here.

[0251] Step 1290 may be carried out in a manner similar to step 1255 (as illustrated above or elsewhere in the present disclosure) . However, to draw some distinction between the case of the first HARQ Group 1 and the case of the second HARQ Group 2, some additional details are provided below.

[0252] If the second HARQ Group 2’s accuracy metric drops below a threshold, the T-TRP 1201, at 1290, sends a MAC CE updating the HARQ Group’s maximum allowed error ratio, thereby ensuring stricter control. Over time, as conditions stabilize or tasks change, the T-TRP 1201 may relax or tighten these metrics group-wise, avoiding the complexity of individually configuring each logical channel.

[0253] In some implementations, this approach provides multiple dimensions of adaptation: ● Metric Flexibility: The system may adopt new metrics as AI models evolve. ● HARQ Grouping: Related tasks or channels may be managed collectively, reducing signaling overhead. ● Periodic Re-evaluation: Just like with Perplexity in implementations illustrated above or elsewhere in the present  disclosure, the T-TRP 1201 periodically reassesses these metrics, ensuring continuous alignment with performance goals.

[0254] In some implementations, a hierarchy of metrics may affect how the T-TRP 1201 adjusts parameters. For example, where if primary metric (Perplexity) is stable but a secondary metric (e.g., Δ-probability) worsens, the T-TRP 1201 may choose to adjust parameters accordingly (e.g., if the primary metric is relatively fixed, the secondary metric may be used to adjust the parameters, but any changes in the primary metric may take precedent over changes in the secondary metric) .

[0255] In some implementations, tasks are allowed to dynamically switch groups based on their evolving requirements or user-defined policies.

[0256] FIG. 13 illustrates, in a schematic diagram, an example wireless network 1300 in which multiple inference tasks are concurrently carried out, in accordance with implementations of the present disclosure. The wireless network 1300 includes a T-TRP 1301, a computing UE_B 1312, a computing UE_C 1314, and a computing UE_D 1322, and a computing UE_E 1324.

[0257] The inference pipeline 1310 includes the computing UE_B 1312 and the computing UE_C 1314, and may be associated with the LLM_QA task as illustrated in FIG. 13. Similarly, the inference pipeline 1320 includes the computing UE_D 1322 and the computing UE_E 1324, and may be associated with the Object_Detection task as illustrated in FIG. 13.

[0258] The T-TRP 1301 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UE_B 1312, the computing UE_C 1314, the computing UE_D 1322, and the computing UE_E 1324 may be similar to may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 1301, and the computing UEs 1312, 1314, 1322, and 1324 are omitted here.

[0259] In an example scenario shown in FIG. 13 where multiple tasks run concurrently, each task may be associated with its own I-QoS profile as follows: ● For LLM_QA Task: ○ Metric_Type: Perplexity ○ Baseline_Error_Tolerance: 1% ○ Metric_Interval: 100 frames ○ Metric_Threshold: Perplexity ≤ X ○ Adjustment_Policies: Tighten HARQ or shift to lower MCS if Perplexity worsens. ● For Object_Detection Task: ○ Metric_Type: Accuracy ○ Baseline_Error_Tolerance: 0.5% ○ Metric_Interval: 200 frames ○ Metric_Threshold: Accuracy ≥ Y% ○ Adjustment_Policies: If Accuracy drops, allocate more PRBs or choose stronger coding schemes.

[0260] Referring to FIG. 13, the value of X may be “0.05” and the value of Y may be “80 (%) ” .

[0261] Additional details regarding the parameters included in the I-QoS profile associated with each task are illustrated above or elsewhere in the present disclosure.

[0262] The T-TRP 1301 issues an RRC Reconfiguration for each task pipeline, detailing its I-QoS profile. After their respective intervals (100 frames for LLM_QA, 200 frames for Object_Detection) , the T-TRP 1301 triggers test inputs and obtains metric results. Each task’s metric outcome directly influences that task’s HARQ / MCS adjustments, thereby ensuring minimal overhead while meeting distinct quality goals.

[0263] Differences from Conventional QoS: Traditional QoS classes treat traffic categories uniformly. I-QoS customizes parameters and metric checks for each task’s unique demands, integrating multiple metrics simultaneously without fixed reliability extremes.

[0264] In one implementation, modulation and coding scheme (MCS) adaptation strategies 1400 are described as shown in FIG. 14.

[0265] The adaptation strategies 1400 include selecting and switching modulation and coding schemes, which may include at least one of the following steps.

[0266] Initial Assignment: Based on the task’s QoS and error tolerance, the T-TRP initially chooses 1410 a certain modulation (e.g., 16QAM) and coding scheme (e.g., moderate low density parity check (LDPC) code rate) for that task.

[0267] Periodic Metric Checks for Fine-Tuning: If Perplexity, Latency or another chosen metric does not meet the requirements, at the next interval, the T-TRP may instruct switching to a more robust MCS. Otherwise, if quality is stable, it might remain at a higher MCS code or modulation rate for better efficiency.

[0268] More specifically, with reference to FIG. 14, the T-TRP may perform, at 1420, a metric check test. In some implementations, the metric check test may be performed periodically.

[0269] The T-TRP may determine, at 1430, whether the metric (e.g., Perplexity, Latency, or another chosen metric) meets a threshold value associated with the metric. If the metric meets the metric threshold value, then the MCS setting may remain. If not, the T-TRP may adjust, at 1440, HARQ ignorance rate.

[0270] At 1450, the T-TRP may determine whether the metric (e.g., Perplexity, Latency, or another chosen metric) meets the metric threshold value. If the metric meets the metric threshold value, then the MCS setting may remain. If not, at 1460, the T-TRP may switch a modulation and coding scheme (e.g., switching to lower order modulation or more robust coding) .

[0271] In some implementations, switching the modulation and coding scheme may be carried out based on a tiered approach, for example, as illustrated below.

[0272] Tiered Approaches: The T-TRP might define tiers of MCS and move tasks between tiers depending on periodic quality measurements: ● Tier 1 (Looser) : 64QAM, higher moderate coding ● Tier 2 (Intermediate) : 16QAM, stronger coding ● Tier 3 (Strict) : QPSK with very strong coding

[0273] “QAM” refers to “quadrature amplitude modulation” , and “QPSK” refers to “quadrature phase shift keying” .

[0274] After a metric check, the T-TRP moves a task’s channel up or down these tiers as needed, maintaining a stable inference quality without overshooting on reliability.

[0275] In another implementation, I-QoS integration and inference-quality measurement procedures are described.

[0276] FIG. 15 illustrates the sequence diagram of a Periodic Metric Check procedure 1500 including injection of a default dataset inference into the pipeline and measuring the quality metrics.

[0277] The wireless network system in which the procedure 1500 is carried out may include a T-TRP 1501, a computing UE_B 1502, and a device 1503 including a verification database. The T-TRP 1501 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UE_B 1502 may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 1501 and the computing UE_B 1502 are omitted here. The device 1503 may be configured to provide datasets to be used for the inference quality tests, for example, upon request of the T-TRP 1501.

[0278] The BS (T-TRP) 1501 may send 1510 a request to a device 1503 with verification dataset (e.g., getDataset () ) and receives 1515 an inference-quality-test dataset from the device. The T-TRP 1501 then prepares 1520 a quality test inference process based on the received dataset.

[0279] Then, the T-TRP 1501 instructs, by sending 1525 an inference quality metric test trigger, one of the Computing UEs (e.g., UE_B 1502 in FIG. 15) to run a known verification input through the inference pipeline. This verification input is selected from a stored dataset or a small reference corpus that the T-TRP 1501 uses for quality checks.

[0280] In some implementations, the TRP 1501 may instruct the computing UEs (e.g., UE_B 1502) periodically, for example every 100 frames, the TRP 1501 may instruct the computing UEs (e.g., UE_B 1502) to run the verification input. Alternatively, the TRP 1501 may a period / interval to indicate the computing UEs (e.g., UE_B 1502) to run the verification input periodically.

[0281] After the verification input passes through the pipeline (potentially involving multiple Computing UEs and intermediate wireless transmissions) , the final output is returned 1530 to the T-TRP 1501. The T-TRP 1501 then computes an inference-quality metric by comparing 1535 this output against a known baseline result.

[0282] If the inference-quality metric remains within acceptable limits, the system continues 1540 operating under the current error tolerance and HARQ / MCS configurations.

[0283] However, if the inference-quality metric does not remain within acceptable limits or the metric threshold is violated, then T-TRP 1501 adjusts 1545 the HARQ / MCS parameters, for example, using DCI and MAC CE signaling.

[0284] For example, if the metric is worse than the metric threshold (e.g., decline in inference quality, the output deviates more significantly from the baseline than allowed) , at 1545, the T-TRP 1501 may send a MAC CE or issue an RRC Reconfiguration to tighten HARQ ignorance rates, utilize a more robust MCS scheme (e.g. switching from 16QAM to QPSK) or a lower code rate) , or otherwise change the error tolerance level. This dynamic feedback loop ensures that the error tolerance strategies remain aligned with the actual inference performance over time, even when wireless conditions and load factors change.

[0285] In some implementations, perplexity is as an example of inference-quality metric. Perplexity is a commonly used metric in language modeling tasks, often employed to gauge how well a model predicts a given sequence of tokens. In essence, perplexity measures the model’s uncertainty: e.g., the lower the perplexity, the more confidently the model predicts the next token in a sequence. The perplexity values obtained for completely different types of models are not comparable with each other, because different types of models have different flexibility and vocabulary ranges. However, when perplexity difference (ΔPPL) between the error-contaminated and the original model is measured for an LLM-based inference pipelines, it may serve as a clear indicator for inference quality implied by the noise. For example, if the pipeline’s intermediate errors cause the final LLM output to become less coherent or deviate from expected results, perplexity will rise, which may suggest decline in inference quality.

[0286] Specifically, perplexity (PPL) for a language model is defined as the exponential of the average negative log-likelihood of the test sample. Intuitively, a perfect predicting model (e.g., one that assigns high probability to the correct next token) would have a very low perplexity. A higher perplexity value means the model found the test input less predictable, and often correlates with degraded inference performance, when the source of the extra uncertainty is known to be the added noise.

[0287] It is noted that perplexity is just one possible metric, and different inference-quality metrics may be possibly used. For example, some tasks (e.g., vision-based inference or sensor fusion) might prefer accuracy-based or Δ-probability difference metrics or KL divergence. The method illustrated in the present disclosure is designed to permit flexible inference-quality metric to make the network runs well. If perplexity is chosen, it serves as a straightforward linguistic quality indicator for LLM tasks. For tasks unrelated to language modeling, an alternate metric may be employed, following the same principles of periodic checks and dynamic adjustments.

[0288] “KL” refers to Kullback–Leibler.

[0289] In some implementations, the frequency of the metric checks may be reduced to lower overhead under stable wireless conditions. For instance, if no perplexity-based adjustments were needed over several intervals, the T-TRP may extend the interval between measurements. Conversely, if conditions fluctuate frequently, more frequent checks may be beneficial.

[0290] Furthermore, if a single metric is insufficient, multiple metrics may be applied concurrently. For example, the system might use perplexity for language-based inference tasks and a Δ-probability difference metric for vision tasks, combining the results to form a comprehensive inference-quality profile. This multi-metric approach may help refine how error tolerance is set for each inference scenario, and ensure that all tasks remain well-served despite differences in their quality requirements.

[0291] In some implementations, to reduce the overhead of the inference tests, the T-TRP, instead of periodically checking the inference quality metrics (e.g., perplexity) , may periodically check the channel quality indices (e.g., block error rate) for each computing UE and only perform the inference quality test if a persisting degradation occurs in channel quality indices.

[0292] FIG. 16 illustrates, in a signal flow diagram, an example method 1600 of inference quality management using an I-QoS profile, in accordance with implementations of the present disclosure.

[0293] The wireless network system in which the process 1600 is performed may include a T-TRP 1601, a computing UE_B 1602b and a computing UE_C 1602c (collectively referred to as 1602) . Although not shown in FIG. 16, in some implementations, a requesting UE that request inference tasks may be involved in the process 1600. The T-TRP 1601 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UEs 1602 may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 1601 and the computing UEs 1602 are omitted here.

[0294] At 1610, the T-TRP 1601 may generate or obtain an I-QoS profile which may include one or more parameters associated with inference quality of the interference task.

[0295] For example, for an LLM Q&A task, the I-QoS profile may include one or more of the followings: ● Metric_Type: Perplexity ● Metric_Interval: Every 100 frames ● Baseline_Error_Tolerance: 1%code block error ratio ● Metric_Threshold: Perplexity ≤ X ● Adjustment_Policies: If metric worsens, reduce HARQ ignorance or switch to more robust coding.

[0296] The T-TRP 1601 sends 1615 an RRC Reconfiguration message to the UEs involved (requesting UE (not shown in FIG. 16) and Computing UEs 1602) , including these I-QoS fields. The UEs configure their HARQ settings and may choose an initial MCS tier that matches the baseline tolerance.

[0297] At 1620, there may be N frames of inference task communications between the T-TRP 1601, the computing UEs 1602, and / or the requesting UE (not shown in FIG. 16) . Some aspects of the inference task communications are illustrated above or elsewhere in the present disclosure.

[0298] After N frames, the T-TRP 1601 sends 1625 a MAC CE “Inference-Quality-Test-Trigger” to a Computing UE 1602b, which injects a known test input.

[0299] At 1630, the test inference data may be transmitted, through the pipeline, to the UE_C 1602c, which may provide a final output of the inference quality test.

[0300] The final output returns to the T-TRP 1601, which computes 1640 Perplexity.

[0301] Steps 1625 to 1640 illustrated above are similar to steps 1025 to 1040 illustrated above or elsewhere in the present disclosure. Accordingly, additional details regarding these steps are omitted.

[0302] If Perplexity is stable, no changes are needed. If Perplexity exceeds X, the T-TRP issues 1645 another MAC CE to lower error tolerance or adopt a lower-order modulation, thereby ensuring that inference quality remains acceptable while minimizing overhead.

[0303] Step 1645 illustrated above are similar to step 1045 illustrated above or elsewhere in the present disclosure. Accordingly, additional details regarding this step is omitted.

[0304] Differences from Conventional QoS: Unlike static 5G QoS profiles, the I-QoS profile includes inference-specific fields and policies that dynamically adjust conditions based on real-time metric evaluations, rather than relying solely on pre-set reliability targets.

[0305] In one implementation the channel-aware I-QoS parameter mappings and pre-emptive MCS / HARQ responses are described.

[0306] FIG. 17 illustrates, in a flow diagram, an example method 1700 for pre-emptive adjustment of parameters associated with inference quality, in accordance with implementations of the present disclosure.

[0307] The T-TRP may proactively response to sudden channel degradation (e.g., a sudden packet loss ratio (PLR) spike) . Beyond periodic inference-quality metric checks, the T-TRP leverages channel condition reports (e.g., CQI, SNR, PLR) to proactively adjust HARQ / MCS parameters before metric degradation occurs. The T-TRP may send a DCI with a "HARQ Ignorance Override" flag, immediately reducing the allowable HARQ ignorance level. Additionally, it may issue an "MCS Override" DCI to switch to a more robust MCS. By taking these measures early, the network maintains a better inference quality level despite deteriorating channel conditions.

[0308] In some implementations of the present disclosure, the “HARQ Ignorance Override” flag may refer to an override ignorance flag indicating that the HARQ retransmission is fully activated regardless of the HARQ ignorance rate.

[0309] For the purpose of illustration, it is presumed that the I-QoS profile for an LLM Q&A task (with Perplexity as the metric) states: ● Baseline_Error_Tolerance = 1% ● Metric_Interval = 100 frames ● Metric_Threshold: Perplexity ≤ X ● Adjustment_Policies: If Perplexity worsens, switch to a lower MCS or reduce HARQ ignorance.

[0310] Referring to FIG. 17, at 1710, the T-TRP may obtain (e.g., receive) channel quality metrics (e.g., CQI, SNR, PLR) .

[0311] At 1715, the TRP may determine whether channel quality (e.g., SNR) drops greater than an SNR change threshold value for the I-QoS profile (e.g., ΔSNRthrehold for I-QoS) . If SNR does not drop or drops less than the SNR change threshold value, then the process moves to 1725 which is illustrated below or elsewhere in the present disclosure.

[0312] However, if SNR suddenly drops by a certain margin (e.g., 5 dB below the normal operational point) , the T-TRP anticipates potential Perplexity degradation in the next check. At 1720, the T-TRP sends an updated DCI configuration selecting a slightly more robust MCS / coding scheme and a MAC CE to reduce HARQ ignorance. By doing so, even if the metric is not yet measured, the T-TRP prevents future quality drops.

[0313] At 1725, the T-TRP may run a normal inference task with (adjusted) parameters associated with HARQ / MCS (e.g., adjusted MCS setting, adjusted HARQ ignorance rate) . Then, the T-TRP may wait for the periodic test interval (e.g., Metric_Interval) which may be indicated in the I-QoS profile.

[0314] At 1730, the T-TRP may run Perplexity test with new adjustment (e.g., Perplexity under adjusted schemes) . At 1735, the T-TRP may determine if the Perplexity meets the metric threshold defined in the I-QoS profile (e.g., Metric_Threshold) .

[0315] If the Perplexity fails to meet the metric threshold defined in the I-QoS profile (e.g., if Perplexity worsens) , at 1740, the T-TRP may issue an “MCS Override” DCI to switch to a lower MCS or reduce HARQ ignorance rate.

[0316] If the Perplexity meets the metric threshold defined in the I-QoS profile, the T-TRP determines 1745 if the Perplexity under adjusted schemes is over-qualified (e.g., determine if ΔPerplexity is too small) . If not over-qualified, the procedure 1700 is finished.

[0317] If, at the next metric interval, the Perplexity under adjusted schemes shows more than enough quality, then the T-TRP may revert 1750 to the previous MCS or error tolerance and check the Perplexity again. This anticipatory mechanism ensures minimal overhead and stable inference quality, aligning I-QoS decisions not only with metric outcomes but also with instantaneous channel states.

[0318] This mechanism allows us to increase the time interval between periodic checks which may be time consuming and costly. Thus, the quality metric checks will respond to the large-scale gradual changes of the channel conditions, while the sudden fluctuations are handled by pre-emptive channel quality checks which are easier to obtain.

[0319] Additional Details for some implementations: ● The T-TRP may store a mapping table correlating SNR / PLR ranges to recommended MCS / HARQ tiers for each  I-QoS profile. ● If channel quality improves, the T-TRP may gradually relax restrictions, possibly sending a MAC CE indicating  a return to the baseline error tolerance or even allowing a slightly higher error margin if previous metric checks were consistently good.

[0320] Because there is a direct correlation between channel condition reports and the inference quality metric, in some implementations, the T-TRP may store and update a history mapping table between channel condition reports and the measured inference quality metrics. In future periodic checks, T-TRP will check if the current measured channel conditions match a previous known state with a difference margin. If yes, then there is no need for a new inference quality metric check and T-TRP deduces the inference metric from the table and adjusts the MCS / HARQ specs accordingly. As the system runs through different channel conditions over time, this method will reduce the overhead of measuring inference quality metric in the periodic Metric Checks.

[0321] In this regard, FIG. 18 illustrates, in a flow diagram, an example method 1800 for inference quality management using a history mapping table, in accordance with implementations of the present disclosure.

[0322] At 1810, the T-TRP may run a normal inference task with one or more parameters associated with inference quality of the inference task (e.g., parameters associated with HARQ / MCS) . The one or more parameters may be included in the I-QoS profile associate with the inference task. Then, the T-TRP may wait for the periodic test interval (e.g., an inference quality evaluation interval) which may be indicated in the I-QoS profile.

[0323] At 1815, the channel quality metrics (e.g., SNR, PLR) of the computing UEs included in the inference pipeline may be obtained by the T-TRP (measure the inference pipeline’s UEs channel quality metrics) .

[0324] At 1820, the T-TRP may determine whether the history mapping table has inference quality metrics that match the channel quality metrics.

[0325] If such metrics exist, the T-TRP may retrieve 1825 the matching Perplexity from the history mapping table. Then, the T-TRP may adjust 1840 one or more parameters associated with inference quality of the inference task (e.g., parameters associated with HARQ / MCS) according to the matching Perplexity.

[0326] On the other hand, if such metrics does not exist, the T-TRP may run 1830 the Perplexity test, and obtain the Perplexity. At 1835, the T-TRP may update the history mapping table by adding the mapping data between the channel quality metrics and the obtained Perplexity (and / or other inference quality metrics) . At 1840, the T-TRP may adjust one or more parameters associated with inference quality of the inference task (e.g., parameters associated with HARQ / MCS) according to the obtained Perplexity.

[0327] An example method 1900 of Proactive Response to Channel Deterioration is illustrated in FIG. 19.

[0328] For the purpose of illustration, it is presumed that PLR spikes suddenly at 1910.

[0329] Additionally or alternatively, it may presumed that sudden SNR drop is detected at 1910.

[0330] In such case as described for implementations related to the Dynamic HARQ ignorance, the T-TRP may suspect a future quality degradation even before the periodic check and immediately send, at 1915, a DCI “HARQ Ignorance Override” bit set to force a reduction in Ignorance (e.g., from 40%to 20%) .

[0331] It is also possible for the T-TRP quality controller module to additionally adjust, at 1920, the MCS using a “MCS override” in cooperation with the “HARQ ignorance override” , which will give a much better temporary control over the error correction mechanisms in case of a sudden channel quality drop. By generating another DCI message “MCS override” , at 1925, the T-TRP will signal the UE and will use a more robust MCS temporarily until the next metric check (e.g. switch from QAM-16 polar coding rate 0.5 to QPSK polar coding rate 0.4) .

[0332] This cooperative adjustment is beneficial mainly because the HARQ retransmissions cause a lot of overhead and delay. Therefore, if the response may be managed by a change in modulation or coding scheme, it will be more effective to apply the MCS overrides first. If the channel quality drop is not manageable by only adjusting MCS overrides, then the HARQ ignorance adjustments may also be involved.

[0333] Subsequent Metric Check occurs at 1930.

[0334] After the next interval, if the Δ-probability metric is still stable 1935, the T-TRP may partially restore 1940 Ignorance to 30%and restore 1940 the MCS to QAM-16 polar coding rate 0.4.

[0335] In other words, if the metric quality remains stable, the T-TRP may restore 1940 HARQ and MCS rules partially.

[0336] If the metric worsens 1945, the T-TRP at 1950 fully enforces low Ignorance and robust MCS until conditions improve.

[0337] Put another way, if the metric worsens 1945, the T-TRP may adjust, at 1950, the inference metrics to even stricter HARQ and MCS rules.

[0338] Integration with I-QoS: I-QoS may define channel thresholds. “If SNR < SNR_threshold” or “If PLR > PLR_threshold, ” Ignorance and MCS coding rate may be reduced by a specified amount, pre-emptively.

[0339] In some implementations, the T-TRP may use a graded approach where each PLR or SNR threshold corresponds to a defined ignorance adjustment step.

[0340] In some implementations, the T-TRP may maintain a decision history table correlating channel conditions and previous HARQ ignorance decisions with resulting inference metrics. If current conditions match a previously encountered state, reuse the known optimal ignorance setting without additional metric checks.

[0341] For example of allowing a graded approach: Each PLR level triggers a specific Ignorance adjustment step. For example, PLR > 1e-2 reduces Ignorance by 10%, PLR > 1e-1 reduces it by 20%, etc. FIG. 20 shows an example mapping using a graded table and a plot correlating SNR levels to the corresponding pre-emptive HARQ ignorance adjustments.

[0342] As there is a direct correlation between channel condition (e.g., SNR, PLR) and the inference quality metric (e.g., BLEU) , the T-TRP may store and update a history mapping table between channel condition reports, their corresponding decisioned HARQ ignorance rates, and the measured inference quality metrics. In subsequent periodic checks, T-TRP will check if the current measured channel conditions match a previous known state up to a difference margin. If yes, then there is no need for a new inference quality metric check and T-TRP reuses the previously-obtained HARQ ignorance rate from the table. As the system runs through different channel conditions over time, this method will reduce the overhead of measuring inference quality metric in the periodic Metric Checks.

[0343] “BLEU” refers to “bilingual language understudy” .

[0344] FIG. 21 illustrates an example decision history table 2100 that show mapping between measured channel SNR, decisioned HARQ ignorance rates, and the measured inference quality metrics, in accordance with implementations of the present disclosure.

[0345] In another implementation the priority-driven resource reallocation using I-QoS isprovided.

[0346] Specifically, FIG. 22 illustrates, in a schematic diagram, an example wireless network 2200 in which multiple inference pipelines with distinct I-QoS profiles run concurrently, in accordance with implementation of the present disclosure. The wireless network 2200 includes a T-TRP 2201, a computing UE_B 2212, a computing UE_C 2214, and a computing UE_D 2222, and a computing UE_E 2224.

[0347] The first inference pipeline 2210 includes the computing UE_B 2212 and the computing UE_C 2214, and the second inference pipeline 2220 includes the computing UE_D 2222 and the computing UE_E 2224.

[0348] The T-TRP 2201 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UE_B 2212, the computing UE_C 2214, the computing UE_D 2222, and the computing UE_E 2224 may be similar to may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 2201, and the computing UEs 2212, 2214, 2222, and 2224 are omitted here.

[0349] Having the example of FIG. 22, where multiple inference pipelines 2210 and 2220 with distinct I-QoS profiles run concurrently. The I-QoS profiles for the first and second pipelines 2210 and 2220 are presumed as follows: ● For the first Pipeline 2210 (LLM_QA) : I-QoS allows 1%error, Perplexity checks every 100 frames. ● For the second Pipeline 2220 (Autonomous Driving Vision Task) : I-QoS allows only 0.2%error, Accuracy  checks every 50 frames, Priority Level = High.

[0350] If, at the next Accuracy measurement, the second Pipeline 2220’s metric shows a drop below the threshold, the T-TRP 2201 detects urgent quality restoration needs. The T-TRP 2201 consults the I-QoS profile for the second Pipeline 2220 and sees that for Accuracy degradation, the T-TRP 2201 may allocate additional PRBs and possibly shift to a stricter code rate. Since the second Pipeline 2220 has high priority, the T-TRP 2201 dynamically reprioritizes resource scheduling, for example, as follows: 1. The T-TRP 2201 reduces some allocated PRBs from the first Pipeline 2210 (which remains stable and may afford  slightly higher error tolerance) . 2. If required, the T-TRP 2201 issues a MAC CE to affected computing UEs of the first and second Pipelines 2210  and 2220 to instruct them to use a lower MCS tier for higher reliability. 3. The T-TRP 2201 increases HARQ retransmission allowance or reduces HARQ ignorance rate specifically for  the second Pipeline 2220’s data channels. 4. In case that lower MCS or reduced HARQ ignorance would degrade the inference latency beyond its tolerance,  then the T-TRP 2201 may change the inference pipeline’s topology by using more capable computing UE that may fit model parts previously assigned to (other) multiple computing UEs, to thereby ensure faster and less error-tolerant inference.

[0351] This priority-based approach ensures that critical tasks will receive immediate quality improvements as dictated by I-QoS policies, while stable tasks operate under more flexible conditions, optimizing overall efficiency and performance.

[0352] Additional Details for some implementations: ● The T-TRP 2201 may maintain a priority ranking within the I-QoS profile itself. Tasks with higher priority get  immediate attention when metrics degrade. ● If multiple tasks degrade simultaneously, the T-TRP 2201 uses priority levels to decide which tasks get the  strictest improvements first.

[0353] An example scenario 2300 of Multiple Tasks with different Priority Levels is illustrated in FIG. 23, in accordance with implementations of the present disclosure.

[0354] A safety-critical inference task 2310 (e.g., industrial automation using an Accuracy metric) is high priority. Another LLM Q&A task 2320 (using BLEU) is low priority and stable.

[0355] Metric-Induced Priority Shift: If the high-priority Accuracy metric dips, T-TRP reduces HARQ Ignorance for the task 2310, ensuring more retransmissions. Simultaneously, T-TRP may increase HARQ Ignorance for the LLM Q&A task 2320 (since the task 2320 is stable and may tolerate more errors) to free resources to ensure the required inference throughput for high-priority task 2310 is preserved.

[0356] Resource Reallocation: By increasing Ignorance on stable tasks 2320, fewer retransmissions occur for them, thus freeing PRBs and scheduling opportunities. High-priority degraded tasks 2310 get more PRBs and lower Ignorance to quickly restore inference quality and throughput.

[0357] Alternatives for some implementations: ● Introduce a “Priority-based HARQ Ignorance Policy” in I-QoS: For critical tasks, Ignorance may be reduced to  near-zero upon metric drop and allocate freed resources from stable tasks. ● If multiple high-priority tasks degrade simultaneously, a weighted approach may be applied, thereby possibly  giving the most critical one minimal Ignorance first.

[0358] In some implementations, after adjusting the low-priority task’s HARQ ignorance, the T-TRP can also switch its MCS to a more robust coding scheme to prevent severe quality degradation. The interplay of HARQ ignorance and MCS adjustments ensures that while prioritizing one task, the network does not critically impair the other. Some networks might define how priority indicators and MCS / HARQ profiles are negotiated at the RRC level and how they may be changed dynamically via MAC CEs and DCI for just-in-time adaptation.

[0359] In this regard, FIG. 24 illustrates, in a flow diagram, an example process 2400 for MCS adjustments for priority scheduling and HARQ ignorance, in accordance with implementations of the present disclosure.

[0360] Metric-Induced Priority Shift: If the high-priority Accuracy metric dips 2410, T-TRP reduces 2415 HARQ Ignorance for that task, ensuring more retransmissions 2420 and adjust its MCS.

[0361] In other words, if the pipeline for a high priority task experiences 2410 quality degradation, then the T-TRP may reduce 2415 HARQ ignorance rate for the high priority task, to ensure more retransmissions 2420.

[0362] In this way, the inference quality may be restored 2425.

[0363] Resource Reallocation: As discussed above or elsewhere in the present disclosure, in this scenario (e.g., if the pipeline for a high priority task experiences 2410 quality degradation) , the T-TRP may attempt to increase 2430 the HARQ ignorance of the low priority Q&A task (since the low priority Q&A task is stable and may tolerate more errors) . By increasing 2430 Ignorance on the stable low-priority task, fewer retransmissions 2435 occur for it, thus freeing 2440 PRB resources and scheduling opportunities.

[0364] In other words, there may be more resources available, at 2440, for example due to fewer retransmissions 2435.

[0365] Thus, the T-TRP may reallocate 2445 the freed up PRBs to High-priority degraded task to quickly restore 2450 inference quality and throughput.

[0366] The freed up PRB may refer to the resources from the pipeline for the low priority task. These resources may be reallocated 2445 to the high-priority task where the throughput dropped at 2442. Accordingly, the inference quality and throughput for the high priority task may be restored at 2540.

[0367] MCS adjustment for low-priority task: After increasing 2430 the HARQ ignorance rate for the low priority Q&A task, there would be a possible quality drop 2455 for the low priority Q&A task. In that case, it would be preferable to adjust 2460 the modulation and coding schemes for the low priority Q&A task to a more robust scheme to preserve 2465 the quality of the low-priority task as well.

[0368] Put another way, if there is a possible quality drop 2455 for the low priority Q&A task, the T-TRP may reduce 2460 the coding rates at the pipeline for the low priority Q&A task.

[0369] This way, the MCS adjustment works in harmony with the HARQ adjustments to maintain 2465 the quality of the low-priority task in case of a resource reallocation.

[0370] In another implementation, the adaptive change of quality-test intervals is provided.

[0371] In particular, FIG. 25 illustrates, in a schematic diagram, example interval management for multiple pipelines associated with difference metrics, in accordance with implementations of the present disclosure.

[0372] The wireless network 2500 to which the multiple pipelines belong includes a T-TRP 2501, a computing UE_F 2512, a computing UE_G 2514, and a computing UE_B 2522, a computing UE_C 2524, a computing UE_D 2532, and a computing UE_E 2534.

[0373] The inference pipeline 2510 includes the computing UE_F 2512 and the computing UE_G 2514. The inference pipeline 2520 includes the computing UE_B 2522 and the computing UE_C 2524. The inference pipeline 2530 includes the computing UE_D 2532 and the computing UE_E 2534.

[0374] The T-TRP 2501 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UE_F 2512, the computing UE_G 2514, the computing UE_B 2522, the computing UE_C 2524, the computing UE_D 2532, and the computing UE_E 2534 may be similar to may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 2501, and the computing UEs 2512, 2514, 2522, 2524, 2532, and 2534 are omitted here.

[0375] I-QoS profiles may dynamically adjust metric checking intervals. If a task’s metrics remain stable over multiple checks, the T-TRP 2501 may increase the interval to reduce overhead. If metrics degrade, the T-TRP 2501 may shorten the interval for closer monitoring. This approach balances measurement overhead and responsiveness to quality changes. Different tasks may rely on different metrics and intervals. For instance: ● Pipeline 2510 (LLM_QA) : Perplexity metric, interval = 100 frames. Stable over several intervals. ● Pipeline 2510 (Object_Detection) : Accuracy metric, interval = 200 frames, but recently unstable accuracy  readings prompt more frequent checks. ● Pipeline 2520 (Another AI Task) : Δ-Probability metric every 50 frames because it is highly sensitive and often  fluctuates.

[0376] If Pipeline 2510 remains stable after multiple checks, the T-TRP 2501 may, according to the I-QoS profile, reduce the metric check frequency for the task assigned to the Pipeline 2510 to every 200 frames, lowering signaling overhead. On the other hand, since the Pipeline 2520’s accuracy recently dropped, the T-TRP 2501 temporarily increases its check frequency to every 100 frames until accuracy stabilizes.

[0377] The initial test interval of 50 frames for the Pipeline 2530 remains unchanged, for example, because the Δ-Probability metric is highly sensitive and often fluctuates.

[0378] This variable interval approach ensures I-QoS is not only metric-adaptive but also overhead-adaptive. By adjusting how often tests occur, the system may minimize unnecessary measurement costs while still ensuring timely detection of quality changes.

[0379] Additional Details for some implementations: ● I-QoS profiles may include rules like “If stable for 3 consecutive checks, double the metric interval. ”  ● Conversely, I-QoS profiles may include a rule of “If metric worsens 2 checks in a row, halve the interval for  closer monitoring. ”

[0380] FIG. 26 illustrates, in a block diagram, example interval adaptations based on stability of inference quality tests, in accordance implementations of the present disclosure.

[0381] Case 2600 for Interval Control: A pipeline uses an Accuracy metric checked every 100 frames. After 3 stable checks 2602, 2604, and 2606, I-QoS rules say 2607 “double the interval” to 200 frames, maintaining current HARQ Ignorance. This reduces signaling frequency and overhead, trusting stability.

[0382] Case 2650 for Unstable Scenarios: Another pipeline using Δ-probability metric is unstable. T-TRP halves 2653 its interval from 100 to 50 frames, allowing more frequent Ignorance updates. For instance, if Δ-probability keeps fluctuating, T-TRP may rapidly decrease Ignorance after each check until stability returns.

[0383] RRC Signaling: T-TRP may send RRC Reconfiguration to update the metric interval fields in I-QoS. MAC CE “Interval Update” may inform UEs of the new check frequency.

[0384] With reference to FIG. 26, in the case of 2600, after three consecutive stable tests 2602, 2604, and 2606, the T-TRP may transmit 2607, to UEs, an adjusted inference quality test interval (200 frames) via MAC-CE “Interval Update” . The subsequent inference quality check 2608 may be carried out 200 frames after the inference quality test 2606. In the case of 2650, after the first inference quality test 2652, which is unstable, the T-TRP may transmit 2653, to UEs, a reduced inference quality test interval (50 frames) via MAC-CE “Interval Update” . The subsequent inference quality checks 2654 and 2656 may be carried out every 50 frames, after the update of the inference quality test interval.

[0385] Alternatives for some implementations: ● Non-linear interval changes or dynamic scaling based on the magnitude of metric deviations (big metric shifts =  more frequent checks, small shifts = slightly more checks) .

[0386] In another implementation, a Tiering approach for MCS / HARQ mechanisms is described along with example tier classes.

[0387] In particular, example MCS / HARQ tier classes associated with a metric (e.g., Accuracy) and a metric threshold value (e.g., Y%) configured in the I-QoS profile is illustrated in FIG. 27.

[0388] I-QoS profiles may define multiple reliability tiers associated with distinct MCS / coding choices. For example, as partly shown in FIG. 27, the I-QoS profile for a vision task might list: ● Tier 1: 16QAM with medium FEC, suitable if measured Accuracy ≥ Y% ● Tier 2: QPSK with stronger FEC if measured Accuracy falls slightly below Y% ● Tier 3: QPSK or BPSK with very robust FEC if measured Accuracy severely degrades

[0389] “BPSK” refers to “binary phase-shift keying” .

[0390] In some implementation, Accuracy is considered severely degraded if the measured Accuracy << Y%.

[0391] When measured Accuracy drops just below Y%at the next check, the T-TRP refers to the I-QoS profile: it will move from Tier 1 to Tier 2. This involves sending a MAC CE adjusting HARQ settings (e.g., reducing HARQ ignorance rate) and a DCI signaling a lower MCS. If Accuracy worsens further in the subsequent interval, the T-TRP escalates to Tier 3. If Accuracy improves for several checks, the T-TRP might revert to Tier 2 or even Tier 1, re-gaining efficiency.

[0392] This tiered approach having multiple tiers, gives the system granularity. It does not jump straight from a loose configuration to an extremely strict one. Instead, it follows a gradual path, ensuring minimal disruption and optimal utilization of spectrum and computational resources.

[0393] Further, the tiering provides the system with a base setting that matches certain conditions, which may then be fine-tuned using DCI signaling to slightly change the HARQ ignorance or MCS code rates. For example, the system may, at the start of the inference task, measure a very noisy channel causing very low inference quality, and therefore may select the Tier 3 (strict tier) as its base configuration. Subsequently, with the periodic quality metric checks, the system might decide to slightly change the code rate or HARQ ignorance rate using DCI and MAC CE signaling within the same Tier 3 baseline. However, if a large-scale change occurs in the channel conditions, the T-TRP may order a Tier change to operate on a different tier as baseline.

[0394] Additional Details for some implementations: ● The T-TRP may store a tier mapping per I-QoS profile. Each tier corresponds to a certain range of metric values,  forming a state machine driven by metric outcomes. ● The UEs update their modulation, coding, and HARQ behavior instantly upon receiving updated DCI or MAC  CE instructions.

[0395] A three tier example 2800 for inference quality management is illustrated below and in FIG. 28.

[0396] In the following example, three tiers 2801-2803 are defined in the I-QoS profile, each specifying a combination of MCS and HARQ ignorance levels. For instance: ● “Tier 1” 2801: High MCS, moderate coding, HARQ Ignorance ~40% (for stable metric conditions) ● “Tier 2” 2802: QPSK, stronger coding, HARQ Ignorance ~20% (for minor metric degradation) ● “Tier 3” 2803: QPSK / BPSK, very robust coding, HARQ Ignorance = 0% (for severe metric drops)

[0397] As inference quality metrics fluctuate, the T-TRP may reassign tasks among these tiers using standardized MAC CEs (e.g., a “Tier Update” CE) or RRC Reconfiguration messages. For example, if the inference quality crosses a given threshold, the network may promote a task from “Tier 1” 2801 to “Tier 2” 2802 or jump directly to “Tier 3” 2803. This tiered approach allows for structured, scalable adaptation. Future networks might define how tier parameters and transitions are communicated, thereby enabling consistent multi-vendor operation.

[0398] Task-Agnostic: If a LLM Q&A pipeline uses Perplexity instead of Accuracy, the same tier transitions apply, just triggered by Perplexity thresholds. If stable, remain in “Tier 1” 2801 or “Tier 2” 2802; if worsens severely, “Tier 3” 2803 is applied.

[0399] These tiers 2801-2803 will work as the baseline operating point of the error handling modules. Then small changes of channel condition will be rectified by the MAC CE and DCI adjustments of HARQ ignorance rate and MCS coding rates.

[0400] Alternatives for some implementations: ● More than three tiers, or dynamic creation of tiers may be implemented based on changing network conditions. ● Tier definitions may also specify maximum allowed Ignorance increments / decrements per metric check to ensure  gradual changes.

[0401] In another implementation, the general HARQ ignorance mechanism and its integration with the I-QoS is provided.

[0402] In particular, FIGs. 29A and 29B illustrate, in a signal flow diagram, an example method 2900 for HARQ ignorance rate control, in accordance with implementations of the present disclosure.

[0403] The wireless network system in which the method 2900 is performed may include a T-TRP 2901, a computing UE_B 2902b and a computing UE_C 2902c (collectively referred to as 2902) . Although not shown in FIGs. 29A and 29B, in some implementations, a requesting UE that request inference tasks may be involved in the process 2900. The T-TRP 2901 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the computing UEs 2902 may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 2901 and the computing UEs 2902 are omitted here.

[0404] Initial Setup: The T-TRP 2901 receives, at 2910, a request for an inference task. Based on the I-QoS profile, the T-TRP 2901 identifies: ● Metric_Type: An inference-quality metric (e.g., Accuracy for a vision task) . ● Baseline_Error_Tolerance: e.g., 1%code block error. ● Metric_Interval: e.g., every 100 frames. ● HARQ Ignorance Policies: Start with a moderate Ignorance level (e.g., 30%) and adjust as metric checks occur.

[0405] Initial Signaling (RRC) : The T-TRP 2901 sends, at 2915, an RRC Reconfiguration message to involved UEs 2902, including I-QoS parameters and a “HARQ_Ignorance_Allowed” field indicating the maximum starting Ignorance. UEs 2902 store these parameters.

[0406] At 2920, there may be N frames of inference task communications between the T-TRP 2901, the computing UEs 2902, and / or the requesting UE (not shown in FIGs. 29A and 29B) . Some aspects of the inference task communications are illustrated above or elsewhere in the present disclosure.

[0407] Periodic Metric Checks: After N (e.g. 200) frames, the T-TRP 2901 sends, at 2925, a MAC CE “Inference-Quality-Test-Trigger” prompting a Computing UE 2902b to run a known test input.

[0408] At 2930, the test inference data may be transmitted from the T-TRP 2901, through the pipeline, to the UE_B 2902b. At 2935, the test inference data may be transmitted from the UE_B 2902b, through the pipeline, to the UE_C 2902c, which may provide a final output of the inference quality test.

[0409] The UE_C 2902c, at 2940, may transmit results, and T-TRP 2901, at 2945, computes the metric. If stable, the T-TRP 2901, at 2950, sends another MAC CE “HARQ Ignorance Update” increasing Ignorance to, say, 50%.

[0410] For example, if computed metric (e.g., measured BLEU) is greater than the threshold value for the metric (e.g., threshold value for BLEU) , then the T-TRP 2901, at 2950, transmit, to the UE_B 2902b and / or UE_C 2902c, the “MAC CE “HARQ Ignorance Update” with updated HARQ Ignorance rate (50%) .

[0411] If the computed metric is good (e.g., measured BLEU is similar to the threshold value for BLEU) , then the T-TRP 2901, at 2955, does nothing.

[0412] If the computed / measured metric worsens, the T-TRP 2901 reduces Ignorance to 10%.

[0413] For example, if computed metric (e.g., measured BLEU) is lower than the threshold value for the metric (e.g., threshold value for BLEU) , then the T-TRP 2901, at 2960, may inspect channel quality of the UE_B 2902b and UE_C 2902c. For the purpose of illustration, at 2965, the T-TRP 2901 may be updated that the channel quality of the UE_B 2902b is low. Then, at 2970, the T-TRP 2901 may (selectively) transmit, to the UE_B 2902b, the “MAC CE “HARQ Ignorance Update” with reduced HARQ Ignorance rate (10%) .

[0414] PHY Integration: At PHY level, DCI may carry a short-term “HARQ Ignorance Override” bit for immediate per-TB adjustments if needed urgently between intervals.

[0415] As noted above, “HARQ Ignorance Override” bit (or override HARQ ignorance flag) may be an override ignorance flag indicating that the HARQ retransmission is fully activated regardless of the HARQ ignorance rate. Put another way, the “HARQ Ignorance Override” bit (or override HARQ ignorance flag) may be used to cancel the HARQ ignorance for a sensitive TB that may require error-free transmission.

[0416] Alternatives for some implementations: ● The T-TRP 2901 may start with a lower Ignorance (10%) and ramp up only after multiple stable checks. ● The T-TRP 2901 may define step increments or decrements of Ignorance in the I-QoS profile (e.g., + / -5%per  stable / unstable result) . ● If degradation is detected, the T-TRP 2901 may selectively adjust the ignorance rate only for UEs with poor  channel quality, thereby providing more granular control.

[0417] In another implementation, the cooperation between MCS and HARQ mechanisms and their cooperative adjustment is provided.

[0418] In particular, FIG. 30A is provided to illustrate example HARQ and MCS interactive adjustments 3050 in accordance with implementations of the present disclosure. The example HARQ and MCS interactive adjustments 3050 may be carried out in an example wireless network 3000 illustrated in FIG. 30B. The wireless network 3000 may be similar to the wireless network illustrated above or elsewhere in the present disclosure (e.g., wireless networks 700, 1100, 1300, 2200, etc. ) and therefore details regarding the wireless network 3000 are omitted here.

[0419] This implementation focuses on coordinated adaptation of HARQ ignorance levels and MCS settings.

[0420] At 3052, periodic metric check may be performed in a manner similar to that described above or elsewhere in the present disclosure (e.g., some or all steps in the periodic metric check procedure 1500, steps 1625-1640 in the method 1600, steps 2925-2945 in the method 2900, etc. ) .

[0421] At 3054, the T-TRP may compare the measured / computed inference metrics with the metric threshold values that may be configured in the I-QoS profile.

[0422] If the measured inference metric values are equal to or greater than the metric threshold values, then no adjustments may be needed as indicated at 3056.

[0423] Otherwise, the T-TRP 3001 may perform HARQ / MCS adjustment as follows.

[0424] At 3058, the T-TRP 3001 may identify if the measured inference metric values are significantly lower than the metric threshold values (e.g., much out of range) . Put another way, the T-TRP 3001 may identify whether there are large and sudden drops in inference quality.

[0425] Rather than always adjusting HARQ retransmission policies, the network 3000 may first attempt a finer-grained adaptation by selecting a more robust MCS. For example, if inference metrics (e.g., accuracy or a specific AI quality metric) slightly deteriorate, the T-TRP 3001, at 3060, may choose to lower the code rate or switch to a more robust modulation scheme (e.g., from 16QAM to QPSK) without changing HARQ ignorance parameters. By doing so, the system reduces the error probability directly, avoiding extra retransmissions. This will generally imply less overhead than using a fix MCS and only allowing extra retransmissions, because each retransmission introduces extra delay.

[0426] If a large and sudden drop in inference quality is detected, the T-TRP 3001, at 3062 and 3064, may combine MCS adjustments with reduced HARQ ignorance, triggering more retransmissions and stronger coding simultaneously. Such signaling may be carried out via dedicated MAC Control Elements (CEs) and DCI fields standardized for "Error Tolerant Inference" scenarios in future releases, ensuring interoperability and consistent performance gains.

[0427] Put another way, at 3062, the T-TRP 3001 may lower the HARQ ignorance rate, thereby switching to higher HARQ rate. Regarding MCS adjustments, at 3064, the T-TRP 3001 may switch to a lower modulation and lower coding rate.

[0428] In another implementation, the data-aware control mechanism of HARQ ignorance is provided.

[0429] In particular, an example data-aware HARQ ignorance control using partial cyclic redundancy check (CRC) is illustrated in FIG. 31.

[0430] In this implementation, the HARQ process is enhanced by using multiple partial CRC checks within a single transport block (TB) . Instead of relying on a single CRC for the entire TB, the TB is divided into segments, each with its own CRC.

[0431] Specifically, the TB 3100 is divided into TB parts 3101, 3102, 3103, 3104, and 3105 with CRCs 3111, 3112, 3113, 3114, and 3115, respectively. Similarly, the TB 3150 is divided into TB parts 3151, 3152, 3153, 3154, and 3155 with CRCs 3161, 3162, 3163, 3164, and 3165, respectively.

[0432] If the number of erroneous segments exceeds a predefined threshold, the network (e.g., T-TRP) concludes that the TB quality is too low and thus will not apply HARQ ignorance. Instead, it requests a standard HARQ retransmission to ensure data integrity.

[0433] For example, in the case of the TB 3100, the T-TRP may conclude the TB quality is too low, because 80%or more of the TB’s CRC checks failed (4 CRCs 3111, 3113, 3114, and 3115 indicate “fail” ) . Accordingly, the T-TRP may activate standard HARQ retransmissions to ensure data integrity, instead of applying HARQ ignorance.

[0434] Conversely, if only a few segments fail CRC checks-below a specified lower threshold-the HARQ controller may apply HARQ ignorance and skip certain retransmissions, accepting minor errors to meet latency constraints.

[0435] For example, in the case of the TB 3150, the T-TRP may conclude the TB quality is good, because 80%or more of the TB’s CRC checks succeeded (4 CRCs 3161, 3162, 3164, and 3165 indicate “pass” ) . Accordingly, the T-TRP may apply HARQ ignorance and deactivate (some) HARQ retransmissions.

[0436] This data-driven approach allows the system to strike a balance between error tolerance and link reliability. For future networks (e.g., future 3GPP networks) , new MAC / RLC procedures and parameters may be introduced to signal the acceptable thresholds for partial CRC-based decisions.

[0437] In another implementation, signaling design (for example, DCI and / or MAC CE fields) for task-specific error tolerance is provided.

[0438] At the PHY / MAC layer, the T-TRP introduces a “inference tolerated error level (ITEL) ” field in a downlink signaling, for example, DCI. Each ITEL value corresponds to a predefined set of MCS / HARQ configurations and error tolerances, stored locally at the T-TRP and the UEs. Additionally, an “Inference Quality” in a downlink signaling, for example, a MAC CE may specify permissible error ratios for each logical channel associated with a particular inference task. For example: ● ITEL=2: Looser tolerance (1%error) , suitable for tasks proven to handle minor inaccuracies, possibly allowing  higher-order modulation (e.g., 64 quadrature amplitude modulation (64QAM) ) and a moderate coding rate. ● ITEL=1: Stricter tolerance (0.5%error) , applied when metrics like Perplexity show degradation; using lower- order modulation (e.g., quadrature phase shift keying (QPSK) ) or stronger FEC for more reliability.

[0439] If metric checks (inference-quality measurement procedures as described in following implementations) indicate quality degradation, the T-TRP sends DCIs for that inference task to switch from ITEL=2 (loose tolerance) to ITEL=1 (tighter tolerance) . This ensures that even as conditions fluctuate-due to time-varying wireless environments or varying Compute UE availability-the inference remains robust. The dynamic DCI and MAC CE configuration ensures ongoing alignment of wireless parameters with inference-level targets, allowing distributed HAPU benefits to persist without raising the complexity of specialized cooling or power setups.

[0440] FIG. 32 provides an example signaling design of DCI 3200 including the ITEL field along with other fields, as shown in the figure.

[0441] Formalizing a new I-QoS class in RAN signaling is illustrated below or elsewhere in the present disclosure.

[0442] This part extends the 3GPP QoS framework by defining a new Inference-QoS class in RAN signaling. An RRC IE specifies the following fields in the I-QoS class: ● QoS Class Identifier: Indicates that it is an inference-related QoS. ● Metric_Type &Interval: e.g., Metric_Type=Perplexity, Interval=100 frames. ● Baseline_Error_Tolerance &Metric_Threshold: e.g., 1%error allowed, Perplexity ≤ X. ● Adjustment_Policies: Detailed instructions for HARQ / MCS adjustments upon metric deviations. ● Task_Class_Indicator &Resource_Priority_Level: Identifies the inference scenario and how important it is  relative to others.

[0443] Additional details regarding the above fields are provided above or elsewhere in the present disclosure.

[0444] When a new inference session starts, a T-TRP assigns the I-QoS class. This ensures all future signaling (RRC, MAC CEs, DCIs) referencing this I-QoS class may easily interpret how to adjust parameters based on periodic metric results. Future networks may define a set of common I-QoS classes for popular AI workloads, allowing network operators to quickly deploy inference services without custom configurations.

[0445] Additional Details for some implementations: ● Operators or standard bodies may predefine several I-QoS classes covering common scenarios like LLM Q&A or 3D vision tasks. ● As AI models evolve, new metrics or classes may be introduced, maintaining backward compatibility and easy  integration.

[0446] FIG. 33 illustrates an example I-QoS class definition in RRC signaling, consistent with the above illustration.

[0447] Standardizing HARQ ignorance Fields is illustrated below or elsewhere in the present disclosure.

[0448] In some implementations, New Fields may be added in RRC. For example, “HARQ_Ignorance_Max, ” “HARQ_Ignorance_Step, ” “Metric_Type_Code, ” and “Ignorance_Trigger_Thresholds” are added to the I-QoS RRC IE. This allows any metric defined by the operator or future standards to be integrated seamlessly.

[0449] Regarding MAC CE Formats, a standard MAC CE “HARQ Ignorance Update CE” may be defined with a field that indicates Ignorance Index referencing a known table. Another MAC CE “HARQ Ignorance Reset CE” may instantly restore Ignorance to zero upon severe metric degradation. These formats are metric-agnostic; they only need to know if metric conditions are stable or not.

[0450] Regarding DCI Extensions, a DCI field, e.g., a single bit or small field indicating short-term HARQ Ignorance override per TB, may be introduced. This ensures even if a metric check is not due yet, immediate channel changes may prompt instant HARQ Ignorance adjustments. This may also further extend the functionality by occasionally overriding regular HARQ ignorance rate for some special important TB to transmit sensitive data that may be preserved at all costs.

[0451] In future networks, new metrics arise (e.g., a complexity-based metric for next-generation AI models) , and accordingly operators may assign a new Metric_Type_Code in RRC IE. UEs from different vendors interpret these standard fields uniformly.

[0452] In some implementations: ● Vendor-specific extensions may define private HARQ Ignorance modes beyond standard tiers. ● A NAS-level configuration allowing network-wide default HARQ Ignorance profiles for certain task classes  before RRC refinement.

[0453] FIG. 34 illustrates example HARQ Ignorance Signaling Fields, consistent with the above illustration and implementations of the present disclosure.

[0454] Standardizing and negotiations are illustrated below or elsewhere in the present disclosure.

[0455] Negotiation and Capability Exchange Mechanisms: ● During initial call setup or network attachment, the UE and T-TRP may negotiate which metrics and error  tolerances are supported. This may be done via RRC procedures or possibly extended NAS-level messages. ● If a certain metric type is not supported by the UE, the T-TRP selects a fallback I-QoS profile with a simpler  metric.

[0456] Backward Compatibility: ● If a UE does not support Error-Tolerant Inference IEs, the T-TRP detects this capability absence via RRC  Capability Exchange and defaults to URLLC or eMBB profiles without HARQ Ignorance adjustments.

[0457] FIG. 35 illustrates, in a signaling flow diagram, an example method 3500 for UE capability negotiation, in accordance with implementations of the present disclosure.

[0458] The wireless network system in which the method 3500 is performed may include a T-TRP 3501 and a UE 3502. The T-TRP 3501 may be similar to the T-TRPs (e.g., T-TRPs 601 and 701) illustrated above or elsewhere in the present disclosure, and the UE 3502 may be similar to the computing UEs (e.g., computing UEs 612, 614, 712, 714, 722, 724) illustrated above or elsewhere in the present disclosure. Accordingly, details regarding the T-TRP 3501 and the UE 3502 are omitted here.

[0459] At 3510, the T-TRP 3501 may transmit, to the UE 3502, a signal including an inquiry related to UE capability as to whether the UE 3502 supports noisy inference.

[0460] At 3520, the UE 3502 transmits, to the T-TRP 3501, the UE capability response indicating as to whether the UE 3502 supports noisy inference.

[0461] If the UE capability response indicates that the UE 3502 supports noisy inference, the T-TRP 3501, at 3530, may transmit, to the UE 3502, RRC reconfiguration with the I-QoS profile and HARQ ignorance / MCS parameters.

[0462] On the other hand, if the UE capability response indicates that the UE 3502 does not noisy inference, the T-TRP 3501, at 3540, may transmit, to the UE 3502, RRC fallback configuration with URLLC or eMBB profiles without HARQ ignorance adjustments.

[0463] FIG. 36 illustrates, in a signal flow diagram, an example method 3600 for error-tolerant inference in a wireless system, in accordance with implementations of the present disclosure.

[0464] The example method 3600 is comprised of steps 3610 to 3655. It shall be understood that not all of these steps are needed in the method 3600. As one example, in some implementations, the method 3600 may only include steps 3625 to 3635. In other words, one or more of steps 3610 to 3620, and 3640 to 3655 may be optional.

[0465] It should be understood that, in some implementations, the order of one or more steps 3610 to 3655 may be changed, however the general concept may be maintained.

[0466] The wireless network system in which the method 3600 is performed may include a network device 3601, one or more computing devices 3602, and a requesting device 3603. The network device 3601 may include a base station, a TRP, or an access node (e.g., RAN node) as illustrated above or elsewhere in the present disclosure. For example, in some implementations, the network device 3601 shown in FIG. 36 may be similar to some or all of the T-TRPs illustrated above or in the present disclosure (e.g., T-TRPs shown in FIGs. 6-13, 15-16, 22, 25, 30A, 30B, 35) . The computing devices 3602 and requesting device 3603 may include UEs, user devices, communication electronic devices, or any other terminal devices or apparatuses illustrated above or elsewhere in the present disclosure. For example, in some implementations, the computing devices 3602 shown in FIG. 36 may be similar to some or all of the computing UEs shown in FIGs. 6-13, 15-16, 22, 25, 30A, 30B, 35. For example, in some implementations, the requesting device 3603 shown in FIG. 36 may be similar to some or all of the requesting UEs shown in FIGs. 8-10 In some implementations, the requesting device 3603 may not be part of the wireless network system in which the method 3600 is performed, for example, if the method 3600 does not include steps 3610 and 3640.

[0467] In some implementations, the network device 3601 may be a device (e.g., base station) that coordinates an inference task. In some implementations, there are a plurality of computing devices 3602 (e.g., UEs) in the wireless network system in which the method 3600 is performed. In such implementations, each computing device 3602 may collectively perform the inference task, and may communicate to each other in a manner described above or elsewhere in the present disclosure. In some implementations, the requesting device 3603 may be considered a device (e.g., UE) that requests an inference task.

[0468] It should be noted that the expression “first information” and other similar expression are used in a manner consistent with usage of those expressions elsewhere in the present disclosure. The ordinal number “first” in the expression “first information” is used simply for ease of reference. For example, “first” is used to distinguish the “first information” from other expressions containing “information” in the present disclosure, but is not used to limit priorities, importance, or transmission order of the “first information” .

[0469] At 3610, in some implementations, the requesting device 3603 may transmit, to the network device 3601, an inference task request. The inference task request may include at least part of first information configuring one or more parameters associated with inference quality of an interference task.

[0470] In some implementations, the one or more parameters associated with inference quality of the interference task indicate at least one of: a transmission error rate to be satisfied during the inference task; a throughput or latency range to be satisfied during the inference task; a type of the inference task; a metric used for evaluation of the inference quality of the inference task; a metric threshold value used for the evaluation of the inference quality of the inference task; an inference quality evaluation interval for the inference task; priority of the inference task; one or more communication parameters used for transmission of data associated with the interference task; or one or more adjustment policies, for parameter adjustment or resource allocation, based on the evaluation of the inference quality of the inference task.

[0471] In some implementations, the metric used for the evaluation of the inference quality of the inference task includes at least one of: a perplexity metric measuring uncertainty associated with the inference task, an accuracy metric based on accuracy associated with the inference task, a metric based on Δ-probability associated with the inference task, or a metric based on Kullback–Leibler (KL) divergence associated with the inference task.

[0472] In some implementations, the one or more communication parameters may include at least one of: at least one parameter associated with a hybrid automatic repeat request (HARQ) protocol, or at least one parameter associated with a modulation and coding scheme (MCS) . In some implementations, the at least one parameter associated with the HARQ protocol may include at least one of: a HARQ ignorance rate indicative of a portion that HARQ retransmission is deactivated, or an override ignorance flag indicating that the HARQ retransmission is fully activated regardless of the HARQ ignorance rate.

[0473] In some implementations, among the above-noted one or more parameters, the inference task request transmitted by the requesting device 3603 may not include the transmission error rate to be satisfied during the inference task, and / or the one or more adjustment policies, for parameter adjustment or resource allocation, based on the evaluation of the inference quality of the inference task.

[0474] At 3615, in some implementations, the network device 3601 may determine some of the above-noted one or more parameters associated with inference quality of the interference task. In some implementations, the network device 3601 may determine at least one of: the transmission error rate to be satisfied during the inference task, or the throughput or latency range to be satisfied during the inference task. For example, the network device 3601 may determine the transmission error rate and / or the throughput or latency range, based on at least one of the metric used for the evaluation of the inference quality of the inference task, a topology of computing devices performing at least part of the inference task, or a number of the computing devices performing at least part of the inference task.

[0475] At 3620, in some implementations, the network device 3601 may transmit, to the one or more computing devices 3602, the first information. In some implementations, the network device 3601 may transmit the first information to each of the computing devices 3602. In some implementations, the first information may be transmitted via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) .

[0476] At 3625, the network device 3601 may transmit, to the one or more computing devices 3602, a request for at least a portion of the inference task. In some implementations where a plurality of computing devices 3602 collectively performs the inference task, the network device 3601 may transmit, to each computing device 3602, a respective request for the inference task (or a request for the respective portion of the inference task) . In some other implementations where a plurality of computing devices 3602 collectively performs the inference task, the network device 3601 may transmit the request to only one computing device 3602 that performs the first portion of the inference task.

[0477] In some implementations, the request for the inference task is generated (e.g., by the network device 3601) based on the inference task request that the network device 3601 received from the requesting device 3603 at 3610.

[0478] At 3630, the one or more computing devices 3602 may perform the inference task. This may involve receiving an input for at least a portion of an inference task in accordance with the first information and performing the at least a portion of the inference task in accordance with the first information.

[0479] More specifically, as noted above, in some implementations, a plurality of computing devices 3602 may collectively perform the inference task. In such implementations, each computing device 3602 may receive an input for a respective portion of the inference task in accordance with the first information. Each computing device 3602 may receive the input from the network device 3601, a preceding computing device 3602 that performs a preceding portion of the inference task, or some other device. Upon receiving the input for the respective portion of the inference task, each computing device 3602 may perform the respective portion of the inference task in accordance with the first information. Then, each computing device 3602 may transmit an output of the respective portion of the inference task in accordance with first information. Each computing device 3602 (except the computing device 3602 that performs the last portion of the inference task) may transmit the output of the respective portion of the inference task to another computing device 3602 that performs a subsequent portion of the inference task.

[0480] In some implementations where a plurality of computing devices 3602 collectively performs the inference task, each computing device 3602 may be included in a respective inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task. In some other implementations where a plurality of computing devices 3602 collectively performs the inference task, more than one computing device 3602 may be included in a same inference layer of the AI / ML model associated with the inference task. Put another way, at least one inference layer of the AI / ML model associated with the inference task may involve more than one computing device 3602.

[0481] In some other implementations, there is only one computing device that receives an input for the inference task and performs the entire inference task. In such implementations, the computing device 3601 may receive the input for the inference task from the network device 3601 or some other device.

[0482] At 3635, the computing device 3602 (e.g., the computing device 3602 that performs the last portion of the inference task or the computing device 3602 that performs the entire inference task) may transmit, to the network device 3601, an output of the inference task in accordance with the first information.

[0483] At 3640, in some implementations, the network device 3601 may transmit, to the requesting device 3603, the output of the inference task in accordance with the first information.

[0484] At 3645, in some implementations, the network device 3601 may evaluate the inference quality of the inference task. The evaluation of the inference quality of the inference task may be based on the output of the inference task performed at steps 3630.

[0485] In some implementations, the inference quality of the inference task may be evaluated using at least one of: a predefined input of the inference task, or a predefined error-free output of the inference task.

[0486] In some implementations, the inference quality of the inference task may be evaluated periodically based on the inference quality evaluation interval.

[0487] At 3650, in some implementations, the network device 3601 may adjust the one or more parameters and / or may reallocate resources to the interference task, based on at least one of: the evaluated inference quality of the inference task, or the one or more adjustment policies.

[0488] In some implementations, the one or more parameters that are adjusted may include at least one of: the transmission error rate, the throughput or latency range, or the one or more communication parameters. For example, in some implementations where the one or more communication parameters include at least one parameter associated with a MCS, the at least one parameter associated with the MCS may be adjusted based on at least one of: the transmission error rate, or the throughput or latency range.

[0489] At 3655, in some implementations, when the one or more parameters are adjusted, the network device 3601 may transmit, to each of the one or more computing devices 3602, the adjusted one or more parameters via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) .

[0490] In summary, the possible method provided in the present disclosure and corresponding benefits are shown in the following Table 1. Table 1

[0491] It is noted that the method described in the present disclosure may also be applied to Other Wireless Tech. For example, the method may extend to non-3GPP networks (e.g., Wi-Fi) , possibly allowing home-based HAPU nodes for distributed inference with flexible error tolerance. The I-QoS and metric-based adaptation principles may be adapted to Wi-Fi or satellite IoT systems, extending inference-driven optimization beyond cellular domains. These protocol designs may be adapted to Wi-Fi or satellite-based systems, extending metric-driven error handling (HARQ / MCS) strategies beyond cellular networks.

[0492] In some implementations, the method described in the present disclosure may be Integrated with model compression: Combining Error-Tolerant Inference and compressed models further reduces data volume, making wireless transmission even less demanding, and enhancing distributed HAPU viability.

[0493] Furthermore, although the T-TRP is mentioned as the handler of the distributed noisy inference systems, it is also possible to have such distributed inference systems under NTN-TRP and other similar network elements.

[0494] In the present disclosure, the terms “a” or “an” are defined to mean “at least one” , that is, these terms do not exclude a plural number of items, unless stated otherwise.

[0495] In the present disclosure, the word “a” or “an” when used in conjunction with the term “comprising” or “including” in the claims and / or the specification may mean “one” , but it is also consistent with the meaning of “one or more” , “at least one” , and “one or more than one” unless the content clearly dictates otherwise. Similarly, the word “another” may mean at least a second or more unless the content clearly dictates otherwise.

[0496] In the present disclosure, terms such as “substantially” , “generally” and “about” , which modify a value, condition or characteristic of a feature of an example implementation, should be understood to mean that the value, condition or characteristic is defined within tolerances that are acceptable for the proper operation of the example implementation for its intended application.

[0497] In the present disclosure, unless stated otherwise, the terms “connected” and “coupled” , and derivatives and variants thereof, refer herein to any structural or functional connection or coupling, either direct or indirect, between two or more elements. For example, the connection or coupling between the elements may be acoustical, mechanical, optical, electrical, thermal, logical, or any combination thereof.

[0498] In some implementations, the connection or coupling between the elements may be considered one or more intermediate elements, apparatus, or devices.

[0499] In the present disclosure, expressions such as “match” , “matching” and “matched” , including variants and derivatives thereof, are intended to refer herein to a condition in which two or more elements are either the same or within some predetermined tolerance of each other. That is, these terms are meant to encompass not only “exactly” or “identically” matching the two elements but also “substantially” , “approximately” or “subjectively” matching the two or more elements, as well as providing a higher or best match among a plurality of matching possibilities.

[0500] In the present disclosure, the expression “based on” is intended to mean “based at least partly on” , that is, this expression may mean “based solely on” or “based partially on” , and so should not be interpreted in a limited manner. More particularly, the expression “based on” may also be understood as meaning “depending on” , “representative of” , “indicative of” , “associated with” or similar expressions.

[0501] In the present disclosure, the terms "system" and "network" may be used interchangeably in different implementations of the present disclosure. "At least one" means one or more, and "a plurality of" means two or more. The term "and / or" describes an association relationship of associated objects, and indicates that three relationships may exist. For example, A and / or B may indicate the following three cases: Only A exists, both A and B exist, and only B exists, where A and B may be singular or plural. The character " / " indicates an "or" relationship between associated objects. "At least one of the following items (pieces) " or a similar expression thereof indicates any combination of these items, including a single item (piece) or any combination of a plurality of items (pieces) . For example, "at least one of A, B, or C" includes: only A; only B; only C; A and B; A and C; B and C; or A, B, and C, and "at least one of A, B, and C" may also be understood as including: only A; only B; only C; A and B; A and C; B and C; or A, B, and C. In addition, unless otherwise specified, ordinal numbers such as "first" and "second" in implementations of the present disclosure are used to distinguish between a plurality of objects, and are not used to limit a sequence, a time sequence, priorities, or importance of the plurality of objects.

[0502] A person skilled in the art should understand that implementations of the present disclosure may be provided as a method, an apparatus (or system) , computer-readable storage medium, or a computer program product. Therefore, the present disclosure may use a form of a hardware-only implementation, a software-only implementation, or an implementation with a combination of software and hardware. Moreover, the present disclosure may use a form of a computer program product that is implemented on one or more computer-usable storage media (including but not limited to a disk memory, an optical memory, and the like) that include computer-usable program code.

[0503] The present disclosure describes with reference to the flowcharts and / or block diagrams of the method, the apparatus, the device (system) , and / or the computer program product according to the present disclosure. It should be understood that computer program instructions may be used to implement each process and / or each block in the flowcharts and / or the block diagrams and a combination of a process and / or a block in the flowcharts and / or the block diagrams. The computer program instructions may be provided for a general-purpose computer, a dedicated computer, an embedded processor, or a processor of another programmable data processing device and enable a machine to execute the instructions. When executed by any computer or the processor of a programmable data processing device, the instructions cause the apparatus to implement specific functions as described in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams. The computer program instructions may alternatively be stored in a computer-readable memory that may indicate a computer or another programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate an artifact that includes an instruction apparatus. The instruction apparatus implements a specific function in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams.

[0504] The computer program instructions may alternatively be loaded onto a computer or another programmable data processing device, so that a series of operations and steps are performed on the computer or the another programmable device, so that computer-implemented processing is generated. Therefore, the instructions executed on the computer or on another programmable device provide steps for implementing specific functions as described in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams.

[0505] It is clear that a person skilled in the art may make various modifications and variations to the present disclosure without departing from the scope of this disclosure. This disclosure is intended to cover these modifications and variations of the present disclosure provided that they fall within the scope of protection defined by the following claims and their equivalent technologies.

Claims

1.A method for error-tolerant inference in a wireless system, comprising:transmitting a request for an inference task; andreceiving an output of the inference task in accordance with first information, the first information configuring one or more parameters associated with inference quality of the interference task.2.The method of claim 1, wherein the one or more parameters associated with inference quality of the interference task indicate at least one of:a transmission error rate to be satisfied during the inference task;a throughput or latency range to be satisfied during the inference task;a type of the inference task;a metric used for evaluation of the inference quality of the inference task;a metric threshold value used for the evaluation of the inference quality of the inference task;an inference quality evaluation interval for the inference task;priority of the inference task;one or more communication parameters used for transmission of data associated with the interference task; orone or more adjustment policies, for parameter adjustment or resource allocation, based on the evaluation of the inference quality of the inference task.3.The method of claim 2, further comprising:based on at least one of the metric used for the evaluation of the inference quality of the inference task, a topology of computing devices performing at least part of the inference task, or a number of the computing devices performing at least part of the inference task, determining at least one of:the transmission error rate, orthe throughput or latency range.4.The method of claim 2 or 3, wherein the one or more communication parameters include at least one of:at least one parameter associated with a hybrid automatic repeat request (HARQ) protocol, orat least one parameter associated with a modulation and coding scheme (MCS) .5.The method of claim 4, wherein the at least one parameter associated with the HARQ protocol includes at least one of:a HARQ ignorance rate indicative of a portion that HARQ retransmission is deactivated; oran override ignorance flag indicating that the HARQ retransmission is fully activated regardless of the HARQ ignorance rate.6.The method of any one of claims 2 to 5, wherein the metric used for the evaluation of the inference quality of the inference task includes at least one of:a perplexity metric measuring uncertainty associated with the inference task;an accuracy metric based on accuracy associated with the inference task;a metric based on Δ-probability associated with the inference task; ora metric based on Kullback–Leibler (KL) divergence associated with the inference task.7.The method of claim 6, wherein the inference quality of the inference task is evaluated using at least one of:a predefined input of the inference task; ora predefined error-free output of the inference task.8.The method of any one of claims 2 to 7, further comprising:evaluating the inference quality of the inference task; andadjusting the one or more parameters and / or reallocating resources to the interference task, based on at least one of:the evaluated inference quality of the inference task, orthe one or more adjustment policies.9.The method of claim 8, wherein adjusting the one or more parameters includes:adjusting at least one of the transmission error rate, the throughput or latency range, or the one or more communication parameters.10.The method of claim 9 when dependent from claim 4, wherein the at least one parameter associated with the MCS is adjusted based on at least one of:the transmission error rate, orthe throughput or latency range.11.The method of any one of claims 8 to 10, wherein when the one or more parameters are adjusted, the method further comprises:transmitting the adjusted one or more parameters via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) .12.The method of any one of claims 8 to 11, wherein the inference quality of the inference task is evaluated periodically based on the inference quality evaluation interval.13.The method of any one of claims 1 to 12, wherein the inference task is collectively performed by a plurality of computing devices, each computing device performing a respective portion of the inference task.14.The method of claim 13, further comprising:transmitting, to each computing device, the first information.15.The method of claim 14, wherein the first information is transmitted via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) .16.The method of any one of claims 13 to 15, wherein transmitting the request for the inference task includes:transmitting, to each computing device, a respective request for the inference task.17.The method of any one of claims 13 to 16, wherein each computing device is included in a respective inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task.18.The method of any one of claims 13 to 16, wherein more than one of the plurality of computing devices are included in a same inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task.19.The method of any one of claims 13 to 18, wherein, for each computing device, at least one of:the computing device receives an input for the respective portion of the inference task in accordance with the first information; andthe computing device transmits an output of respective portion of the inference task in accordance with the first information.20.The method of any one of claims 12 to 19, wherein the plurality of computing devices include one or more user equipments (UEs) .21.The method of any one of claims 1 to 20, further comprising:receiving, from a requesting device, an inference task request including at least part of the first information; andtransmitting, to the requesting device, the output of the inference task in accordance with the first information.22.The method of claim 21, wherein the request for the inference task is generated based on the received inference task request.23.A network device comprising:at least one processor coupled with a computer-readable medium having stored thereon, computer-executable instructions that, when executed, cause the network device to perform the method of any one of claims 1 to 22.24.A method for error-tolerant inference in a wireless system, comprising:receiving an input for at least a portion of an inference task in accordance with first information, the first information configuring one or more parameters associated with inference quality of the interference task;performing the at least a portion of the inference task in accordance with the first information; andtransmitting an output of the at least a portion of the inference task in accordance with the first information.25.The method of claim 24, wherein the one or more parameters associated with inference quality of the interference task indicate at least one of:a transmission error rate to be satisfied during the inference task;a throughput or latency range to be satisfied during the inference task;a type of the inference task;a metric used for evaluation of the inference quality of the inference task;a metric threshold value used for the evaluation of the inference quality of the inference task;an inference quality evaluation interval for the inference task;priority of the inference task;one or more communication parameters used for transmission of data associated with the interference task; orone or more adjustment policies, for parameter adjustment or resource allocation, based on the evaluation of the inference quality of the inference task.26.The method of claim 25, wherein at least one of the transmission error rate or the throughput or latency range is determined based on at least one of:the metric used for the evaluation of the inference quality of the inference task;a topology of computing devices performing at least part of the inference task; ora number of the computing devices performing at least part of the inference task.27.The method of claim 25 or 26, wherein the one or more communication parameters include at least one of:at least one parameter associated with a hybrid automatic repeat request (HARQ) protocol, orat least one parameter associated with a modulation and coding scheme (MCS) .28.The method of claim 27, wherein the at least one parameter associated with the HARQ protocol includes at least one of:a HARQ ignorance rate indicative of a portion that HARQ retransmission is deactivated; oran override ignorance flag indicating that the HARQ retransmission is fully activated regardless of the HARQ ignorance rate.29.The method of any one of claims 25 to 28, wherein the metric used for the evaluation of the inference quality of the inference task includes at least one of:a perplexity metric measuring uncertainty associated with the inference task;an accuracy metric based on accuracy associated with the inference task;a metric based on Δ-probability associated with the inference task; ora metric based on Kullback–Leibler (KL) divergence associated with the inference task.30.The method of claim 29, wherein the inference quality of the inference task is evaluated using at least one of:a predefined input of the inference task; ora predefined error-free output of the inference task.31.The method of any one of claims 28 to 30, wherein the inference quality of the inference task is evaluated, and wherein the one or more parameters are adjusted and / or resources to the interference task are reallocated based on at least one of:the evaluated inference quality of the inference task, orthe one or more adjustment policies.32.The method of claim 31, wherein the one or more parameters that are adjusted include at least one of the transmission error rate, the throughput or latency range, or the one or more communication parameters.33.The method of claim 32 when dependent from claim 27, wherein the at least one parameter associated with the MCS is adjusted based on at least one of:the transmission error rate, orthe throughput or latency range.34.The method of any one of claims 31 to 33, wherein when the one or more parameters are adjusted, the method further comprises:receiving the adjusted one or more parameters via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) .35.The method of any one of claims 31 to 34, wherein the inference quality of the inference task is evaluated periodically based on the inference quality evaluation interval.36.The method of any one of claims 24 to 35, wherein the inference task is collectively performed by a plurality of computing devices, each computing device performing a respective portion of the inference task.37.The method of claim 36, wherein performing the at least a portion of the inference task includes performing the respective portion of the inference task.38.The method of any one of claims 36 to 37, further comprising:receiving, from a network device coordinating the inference task, the first information.39.The method of claim 38, wherein the network device coordinating the inference task is a base station.40.The method of claim 38 or 39, wherein the first information is received via at least one of radio resource control (RRC) signaling, media access control –control element (MAC-CE) , or downlink control information (DCI) .41.The method of any one of claims 38 to 40, further comprising:receiving, from the network device coordinating the inference task, a request for the at least a portion of the inference task.42.The method of any one of claims 36 to 41, wherein each computing device is included in a respective inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task.43.The method of any one of claims 36 to 41, wherein more than one of the plurality of computing devices are included in a same inference layer of an artificial intelligence or machine learning (AI / ML) model associated with the inference task.44.The method of any one of claims 36 to 43, wherein the plurality of computing devices include one or more user equipments (UEs) .45.A computing device comprising:at least one processor coupled with a computer-readable medium having stored thereon, computer-executable instructions that, when executed, cause the computing device to perform the method of any one of claims 24 to 44.46.A method for error-tolerant inference in a wireless system, comprising:transmitting an inference task request including at least part of first information, the first information configuring one or more parameters associated with inference quality of an interference task; andreceiving an output of the inference task in accordance with the first information.47.The method of claim 46, wherein the one or more parameters associated with inference quality of the interference task indicate at least one of:a throughput or latency range to be satisfied during the inference task;a type of the inference task;a metric used for evaluation of the inference quality of the inference task;a metric threshold value used for the evaluation of the inference quality of the inference task;an inference quality evaluation interval for the inference task;priority of the inference task; orone or more communication parameters used for transmission of data associated with the interference task.48.The method of claim 46 or 47, wherein the inference task request is transmitted to a network device coordinating the inference task, and the output of the inference task is received from the network device coordinating the inference task.49.A requesting device comprising:at least one processor coupled with a computer-readable medium having stored thereon, computer-executable instructions that, when executed, cause the requesting device to perform the method of any one of claims 46 to 48.50.A non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions that, when executed by a processor of a device, cause the device to perform the method of any one of claims 1 to 22, 24 to 44, and 46 to 48.51.An apparatus, configured to perform the method of any one of claims 1 to 22 or any one of claims 24 to 44 or any one of claims 46 to 48.52.An apparatus comprising:a communication unit configured to:transmit a request for an inference task; andreceive an output of the inference task in accordance with first information, the first information configuring one or more parameters associated with inference quality of the interference task.53.An apparatus comprising:a communication unit configured to:receive an input for at least a portion of an inference task in accordance with first information, the first information configuring one or more parameters associated with inference quality of the interference task, andtransmit an output of the at least a portion of the inference task in accordance with the first information; anda processing unit configured to:perform the at least a portion of the inference task in accordance with the first information.54.An apparatus comprising:a communication unit configured to:transmit an inference task request including at least part of first information, the first information configuring one or more parameters associated with inference quality of an interference task; andreceive an output of the inference task in accordance with the first information.55.An apparatus comprising:one or more processors; andan interface circuit configured to:transmit a request for an inference task; andreceive an output of the inference task in accordance with first information, the first information configuring one or more parameters associated with inference quality of the interference task.56.An apparatus comprising:one or more processors; andan interface circuit configured to:receive an input for at least a portion of an inference task in accordance with first information, the first information configuring one or more parameters associated with inference quality of the interference task, andtransmit an output of the at least a portion of the inference task in accordance with the first information; andwherein the one or more processors are configured to:perform the at least a portion of the inference task in accordance with the first information.57.An apparatus comprising:one or more processors; andan interface circuit configured to:transmit an inference task request including at least part of first information, the first information configuring one or more parameters associated with inference quality of an interference task; andreceive an output of the inference task in accordance with the first information.58.The apparatus of any one of claims 55 to 57, wherein the interface circuit comprises one or more transceivers.59.A communication system, wherein the communication system comprises a first apparatus configured to perform the method of any one of claims 1 to 22, a second apparatus configured to perform the method of any one of claims 24 to 44, and a third apparatus configured to perform the method of any one of claims 46 to 48.