Improvements in and relating to management of AI / ML in a wireless network

The AIOU dynamically orchestrates AI/ML inference tasks across network entities, addressing inefficiencies by decomposing, allocating, and synchronizing subtasks, thereby optimizing resource use and reducing latency while maintaining accuracy.

GB2639747APending Publication Date: 2025-10-01SAMSUNG ELECTRONICS CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
GB2025000606
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2025-01-16
Publication Date
2025-10-01

AI Technical Summary

Technical Problem

Existing wireless networks lack mechanisms to dynamically allocate and orchestrate AI/ML inference operations across network entities, leading to inefficiencies in resource usage, latency, and computational constraints, especially in diverse and dynamic environments.

Method used

An apparatus and method for dynamic inference control and orchestration, utilizing an Advanced Inference Orchestration Unit (AIOU) that decomposes tasks into subtasks, allocates them across network entities like UE and RAN/CN, and synchronizes their execution, with early-exit termination options based on real-time conditions.

Benefits of technology

This approach optimizes resource allocation, reduces latency, and maintains prediction accuracy by adaptively managing ML inference processes, enhancing energy efficiency and computational performance across varying network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An orchestration entity receives an input from a first network entity and transmits an output to a second network entity, wherein the orchestration entity determines on a real-time basis where in a telecommunication network, and when, an AI / ML inference task is performed. The orchestrator may comprise a task decomposition entity that divides AI / ML inference tasks into subtasks and optimises each subtask for execution at specific network entities based on the computational demands and nature of the subtasks. The orchestrator may comprise a dynamic inference location allocation entity that determines a suitable network entity for executing each subtask. The orchestrator may comprise a subtask synchronisation manager to coordinate the execution of the subtasks distributed across the network to ensure that the subtasks are synchronised in terms of timing, sequence, and consistency, maintaining the integrity and efficiency of the final inference outcome. Subtasks may run in parallel or in sequence. Early-exit termination of the inference task may occur if a certain time has elapsed, a satisfactory inference result is achieved, or the output is inaccurate, or based on the local status of the entity performing the inference.
Need to check novelty before this filing date? Find Prior Art

Description

The present invention relates to improvements in the dynamic managements of an Artificial Intelligence / Machine Learning (AI / ML) interface across different wireless network entities. More specifically, the present invention relates to the field of machine learning applied to a network function that enables adaptive inference using early-exit machine learning frameworks for optimizing computation and communication resources in various networking environments. Sixth Generation, 6G, networks promise revolutionary advancements over their predecessors, such as 3G, 4G, 5G. Industry and academia envision unparalleled data speeds, unprecedentedly low latencies, superior reliability and the ability to handle a myriad of devices ranging from high-end computational devices to Internet of Things (loT) sensors. These advancements are propelling the communications industry towards an ecosystem where Artificial Intelligence (Al) and Machine Learning (ML) are deeply intertwined with communication protocols and services. In such a scenario, Al-driven operations, especially inference, become pivotal. Inference is the process by which trained Al and ML models provide predictions, classifications, and other outcomes based on new data inputs. These operations, while transformative, are computationally intensive and vary in their processing requirements based on the complexity of the task and the data involved. Prior art wireless networks were predominantly architected with a focus on data transmission, often with a deterministic set of operations in mind. These operations follow predictable patterns of sending and receiving data. However, with the proliferation of Al and ML in various applications - from advanced image recognition in augmented reality to real-time data analytics in autonomous vehicles - the dynamics of network operations have shifted. The challenge is no longer just about transmitting data efficiently but also about making intelligent decisions in real-time, based on the data. Moreover, the locations where these inferences occur - be it the User Equipment (UE), Radio Access Network (RAN), or Core Network (CN) - have their unique advantages and constraints. For instance, while inferences at the UE may reduce network latency, they may be constrained by the device's computational capability or battery life. On the other hand, inferences in the CN might benefit from superior computational resources but could introduce latency due to data travel time. Prior art methodologies and infrastructures often exhibit rigidity in dynamically allocating and orchestrating these inference operations based on predefined threshold conditions and device capabilities. There is a notable absence of mechanisms that could perceive the holistic network environment, understand the requirements and constraints of an AI / ML task, and dynamically control and orchestrate the inference for optimal outcomes. To this end, the integration of Al and ML into wireless networks is not merely a matter of offloading computational tasks but is about weaving intelligence into the very fabric of network operations. Furthermore, ensuring such dynamic operations align with established standards like those from the 3rd Generation Partnership Project (3GPP) is paramount. 3GPP, being a leading body in the establishment of network standards, ensures that any new mechanism seamlessly integrates with existing infrastructures, maintaining interoperability, security, and efficiency. In this light, the need for a dynamic control and orchestration mechanism tailored for inference operations, and compliant with 3GPP, becomes not just beneficial but crucial for the evolution and efficiency of future wireless networks. The topic of inference handling is part of the upcoming Release 19 work item on AI / ML air interface of 3GPP working groups. The following is drawn from RAN#102 (December 11-15, 2023), in particular, [RP-234039] New WID on Artificial Intelligence (AI)ZMachine Learning (ML) for NR Air Interface: “In this study, 3GPP RAN groups will work on the following objectives: Provide specification support for the following aspects: AI / ML general framework for one-sided AI / ML models within the realm of what has been studied in the FS_NR_AIML_Air project [RAN2]: o Signalling and protocol aspects of Life Cycle Management (LCM) enabling functionality and model (if justified) selection, activation, deactivation, switching, fallback ■ Identification related signalling is part of the above objective o Necessary signalling / mechanism(s) for LCM to facilitate model training, inference, performance monitoring, data collection (except for the purpose of CN / OAM / OTT collection of UE-sided model training data) for both UE-sided and NW-sided models o Signalling mechanism of applicable functionalities / models Beam management - DL Tx beam prediction for both UE-sided model and NW-sided model, encompassing [RAN1 / RAN2]: o Spatial-domain DL Tx beam prediction for Set A of beams based on measurement results of Set B of beams (“BM-Case1”) o Temporal DL Tx beam prediction for Set A of beams based on the historic measurement results of Set B of beams (“BM-Case2") o Specify necessary signalling / mechanism(s) to facilitate LCM operations specific to the Beam Management use cases, if any o Enabling method(s) to ensure consistency between training and inference regarding NW-side additional conditions (if identified) for inference at UE NOTE: Strive for common framework design to support both BM-Case1 and BM-Case2 Positioning accuracy enhancements, encompassing [RAN1 / RAN2 / RAN3]: o Direct AI / ML positioning: ■ (1st priority) Case 1: UE-based positioning with UE-side model, direct AI / ML positioning ■ (2nd priority) Case 2b: UE-assisted / LMF-based positioning with LMF-side model, direct AI / ML positioning ■ (1st priority) Case 3b: NG-RAN node assisted positioning with LMF-side model, direct AI / ML positioning o AI / ML assisted positioning ■ (2nd priority) Case 2a: UE-assisted / LMF-based positioning with UE-side model, AI / ML assisted positioning ■ (1st priority) Case 3a: NG-RAN node assisted positioning with gNB-side model, AI / ML assisted positioning o Specify necessary measurements, signalling / mechanism(s) to facilitate LCM operations specific to the Positioning accuracy enhancements use cases, if any o Investigate and specify the necessary signalling of necessary measurement enhancements (if any) o Enabling method(s) to ensure consistency between training and inference regarding NW-side additional conditions (if identified) for inference at UE for relevant positioning sub use cases Core requirements for the above two use cases for AI / ML LCM procedures and UE features [RAN4]: o Specify necessary RAN4 core requirements for the above two use cases. o Specify necessary RAN4 core requirements for LCM procedures including performance monitoring. Study objectives with corresponding checkpoints in RAN#105 (Sept ’24): CSI feedback enhancement [RAN1]: o For CSI compression (two-sided model), further study ways to: ■ Improve trade-off between performance and complexity / overhead • e.g., considering extending the spatial / frequency compression to spatial / temporal / frequency compression, cell / site specific models, CSI compression plus prediction (compared to Rel-18 non-AI / ML based approach), etc. ■ Alleviate / resolve issues related to inter-vendor training collaboration, while addressing other aspects requiring further study / conclusion as captured in the conclusions section of the TR 38.843. o For CSI prediction (one-sided model), further study performance gain over Rel-18 non-AI / ML based approach and associated complexity, while addressing other aspects requiring further study / conclusion as captured in the conclusions section of the TR 38.843 (e.g., cell / site specific model could be considered to improve performance gain). Necessity and details of model Identification concept and procedure in the context of LCM [RAN2 / RAN1] CN / OAM / OTT collection of UE-sided model training data [RAN2 / RAN1]: o For the FS_NR_AIML_Air study use cases, identify the corresponding contents of UE data collection o Analyse the UE data collection mechanisms identified during the FS_NR_AIML_Air (TR 38.843 section 7.2.1.3.2) study along with the implications and limitations of each of the methods Model transfer / delivery [RAN2 / RAN1]: o Determine whether there is a need to consider standardised solutions for transferring / deliverlng AI / ML model(s) considering at least the solutions identified during the FS_NR_AIML_Air study Testability and interoperability [RAN4]: o Finalize the testing framework and procedure for one-sided models and further analyse the various testing options for two-sided models, in collaboration with RAN1, and including at least: ■ Relation to legacy requirements ■ Performance monitoring and LCM aspects considering use-case specifics ■ Generalization aspects ■ Static / non-static scenarios / conditions and propagation conditions for testing (e.g., CDL, field data, etc.) ■ UE processing capability and limitations ■ Post-deployment validation due to model change / drift o RAN5 aspects related to testability and interoperability to be addressed on a request basis" The inference of ML models, especially Deep Neural Networks, DNNs, often demands significant computational resources. Executing them at the UE side could be resourcedraining, affecting the device's energy efficiency, or even not feasible at all given the UEs computational limitations. Conversely, running them entirely on centralized cloud resources could introduce latency, especially when real-time processing is essential. Moreover, the diverse range of applications and the dynamic nature of wireless environments mean that a one-size-fits-all approach to inference execution is sub-optimal. Depending on the application's requirements and the network's current state, the most efficient location for inference might be the UE, the Radio Access Network (RAN), the Core Network (CN), or even a combination of these. Additionally, there may be instances where it is beneficial to pause, stop, resume, or migrate the inference process based on changing conditions or priorities. In this evolving landscape, there is a clear need for a system that can dynamically orchestrate ML inference tasks across various entities of the wireless network, adapting in real-time to the ever-changing conditions and requirements. It is an aim of embodiments of the present invention to address shortcomings in the prior art, whether mentioned herein or not. According to the present invention there is provided an apparatus and method as set forth in the appended claims. Other features of the invention will be apparent from the dependent claims, and the description which follows. According to a first aspect of the present invention, there is provided a method of operating a telecommunication network comprising the steps of: providing an orchestration entity arranged to receive at least one input from a first network entity and to transmit at least one output to a second network entity, wherein the orchestration entity is arranged to determine, on a real-time basis, where in the telecommunication network, and when, an AI / ML inference task is performed. In an embodiment, the orchestration entity comprises a Task Decomposition Entity arranged to dissect complex AI / ML inference tasks into a plurality of subtasks and to optimise each of the plurality of subtasks for execution within one or more specific network entities based on the computational demands and nature of the subtasks. In an embodiment, the orchestration entity comprises a Dynamic Inference Location Allocation entity arranged to determine a suitable network entity for executing one or more of the plurality of subtasks. In an embodiment, the orchestration entity comprises a Subtask Synchronization Manager arranged to coordinate the execution of the plurality of subtasks distributed across various network entities and to ensure that the plurality of subtasks are synchronized in terms of timing, sequence, and consistency, thereby maintaining the integrity and efficiency of the final inference outcome. In an embodiment, the plurality of subtasks are either run in parallel or in sequence. In an embodiment, early-exit termination of the inference task if at least one condition is met. In an embodiment, the at least one condition is one or more of: that a certain time has elapsed; that a satisfactory inference result has been achieved; that the output is inaccurate; and a local status of the entity performing the inference. In an embodiment, the first network entity is the same as the second network entity. According to a second aspect of the present invention, there is provided apparatus arranged to perform the method of the first aspect. Embodiments of the invention focus on Dynamic Inference Control and Orchestration. Embodiments provide a mechanism for dynamic control and orchestration of inference operations across wireless network entities. This mechanism enables real-time adjustments, allowing for optimal resource allocation, reduced latencies, and maintained prediction accuracy, regardless of network conditions. Embodiments provide the ability to adaptively manage and control ML inference processes based on prevailing network conditions and available resources. By allowing the neural networks' operations to be spread, stopped, paused, resumed, or even shifted among different entities - from UEs to RAN and CN -embodiments seek to optimize both energy consumption and processing latency, ensuring that ML tasks are handled most efficiently given the current conditions. A key feature of embodiments is the Advanced Inference Orchestration Unit, AIOU. The AIOU is equipped with capabilities to garner real-time data about network conditions, device computational capacities, battery statuses, latency requirements, and the specific demands of the AI / ML task at hand. Through a blend of advanced analytics and predictive algorithms, the AIOU makes real-time decisions about where and when an inference task should be executed. Further, an inference task may be divided or split into multiple sub tasks and spread across several entities for execution. Embodiments of the invention provide a new set (one or more) of logical network entity(-ies) (and / or function(s)) that are involved in the inference process of a given AI / ML model (or models). In one example, the new entity (and / or function) is referred to as Advanced Inference Orchestration Unit (AIOU) designed to dynamically control and orchestrate inference operations within communications networks. The set of new network entity(-ies) (and / or function(s)) may be co-located (or part of) one (or more) of existing network entity(-ies) (and / or function). In one example, the new network entity (and / or function) is included in the User Equipment (UE) or group of UEs. In another example, one (or more) instances of the AIOU entity (or function) is distributed in several network entities (or functions) and / or the desired UE(s). In another example, the network entity (and / or function) is part of (or included in) a server, cloud, application and / or other internal or external entity. The newly defined network entity (and / or function) AIOU is capable of garnering data from at least one network entity (and / or function). In one example, the AIOU may collect real-time data regarding network conditions, device computational capacities, battery status, latency requirements, and specific demands of the AI / ML task. In one embodiment, the Advanced Inference Orchestration Unit (AIOU) incorporates a Task Decomposition Engine (TDE) that is capable of dissecting complex AI / ML inference tasks into multiple, manageable subtasks. The TDE optimizes each subtask for execution within specific network elements based on the computational demands and nature of the tasks. Additionally, the TDE assigns priority levels to each subtask, considering factors such as urgency, computational intensity, and the overall significance of the task. In another embodiment, the new network (and / or function) AIOU features a Dynamic Inference Location Allocation (DILA) mechanism and determining the most suitable location for executing the subtasks or the whole inference operation, including, but not limited to, UE, Radio Access Network (RAN), or Core Network (CN). In another example, determining the most suitable location, time, network entity (or function), and / or duration (e.g. inference window) for executing these operations. In a further embodiment, the AIOU encompasses a Subtask Synchronization Manager (SSM), which is responsible for coordinating the execution of the subtasks distributed across various network entities, including but not limited to, UE, RAN, and CN. The Subtask Synchronization Manager ensures that the processing of these subtasks is synchronized in terms of timing, sequence, and consistency, thereby maintaining the integrity and efficiency of the final inference outcome. The manager also dynamically adjusts synchronization parameters in response to real-time changes in network latency, device availability, and resource utilization. The aforementioned evolved set of entities and / or functions encompasses comprehensive inference-related actions such as start, stop, activate, deactivate, pause, resume, update, delete, postpone, and notably, early exit termination before a set time, enhancing overall inference efficiency. In one example, decision-making functionalities, such as inference pause, restart, early-exit or termination, are dynamically adjusted based on real-time data pertaining to the inference progression. In certain scenarios, early-exit termination of inference can occur before a predetermined time (e.g. pre-configured or set by at least one network entity (and / or function) or external entity (and / or function)). In another example, the early-exit termination of inference can occur if the acquired data suggests that a satisfactory inference result has been achieved. Optionally, the reasoning behind such an early termination might be specified by, for instance, a designated "cause value." Embodiments of this invention are designed to support the integration of various early-exit termination frameworks for AI / ML inference tasks within network environments. However, while embodiments of the invention accommodate these frameworks, it does not rely on any particular specific implementation and may be considered agnostic in this regard. Alternatively, if the data indicates that the inference output is off-mark (as opposed to expected accuracy metrics), an early-exit termination can be initiated. The cause for such a decision might be labelled under terms like “InaccurateOutput" or “Ear / yExitlnaccuratelnference." Other potential triggers for inference termination and / or early-exit may be associated with the complexity of the inference and / or the encountered latency. In one example, triggers can be related to the local status (e.g. resource, power, etc.) of the entity performing the inference (or involved in the inference operation). The newly defined network entity(-ies) and / or network function(s) introduced herein may be included in (or part of, or co-located with) RAN (e.g. eNB, NG-RAN, gNB, etc.), CN (e.g. AMF, SMF, UPF, etc.), dedicated internal / external entity (or function), a server, database, cloud, Application Function (AF), and / or a UE (e.g. loT UE, NTN UE, loT NTN UE, other UE category, or UE type, etc.) (or a set of UEs), UE(s) with a specific capability, etc. The term AI / ML model may be replaced (or used interchangeably) with the term AI / ML functionality / use-case / configuration / scenario / site. The solutions, methods, embodiments, and / or examples, presented in this invention, may apply to one or more type(s) of communication systems, such as 4G, 4G-Advanced, 5G, 5G-Advanced, and 6G. Moreover, the above may also apply (in full or part or modified) to systems of Non-Terrestrial Networks (NR-NTN and / or loT-NTN), in addition to Terrestrial Networks (TN). New signalling / messages may be required between the UE, the network, and the new set of network entity(-ies) (and / or network function(s)) introduced herein. For example, new RRC and / or NAS signalling / messages / IEs, orX2 orXn or NG, or F1 or E1 signalling / messages / IEs may be required. New system information (and / or new SIBs) may be required. In an embodiment, the exchange of information between the UE, network, and / or newly introduced set of network entity(-ies)(and / or function(s)) are carried on existing RRC and / or NAS signalling / messages / IEs, and / or system information (broadcast periodic and / or on-demand), using exiting and / or newly defined SIBs. Embodiments of this invention are described in terms of 3GPP networks, but may also be used with non-3GPP entities. The execution of steps in all examples, algorithms, procedures, and figures can be in any order, not only in the order shown. Moreover, modification and / or addition of other steps and / or combination of steps is also possible. All examples, and / or Figures, text and description, may apply to the definition of a model, model functionality and / or functionality. Additionally, the terms Model or Functionality can be used interchangeably, and / or jointly (together), for example, Model ID and / or Functionality ID. All examples, Figures, text and description, steps, may apply similarly (or in some cases differently) to the model and / or functionality. For example, all parameters and / or metrics can be used for life-cycle management (LCM) of a given model (or models) and / or a given functionality (or set of functionalities) herein. In summary, embodiment of the invention provide: • A network communication module that receives input data from network nodes or devices and transmits the output. • An early-exit machine learning framework, including multiple exit points and corresponding confidence metrics, enabling adaptive inference based on the confidence of predictions at each exit point. • A confidence threshold determination module that calculates and adjusts the confidence thresholds for each exit point dynamically, based on factors such as network conditions, device constraints, and application requirements. Embodiments of the invention offer several advantages, including: • Faster inference times and reduced latency, improving the performance of Al-driven applications and services in networking environments. • Energy-efficient Al model deployment, particularly suitable for resource-constrained devices and edge computing. • Improved resource allocation and network-level optimization, enhancing the overall performance and efficiency of the network. Although a few preferred embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes and modifications might be made without departing from the scope of the invention, as defined in the appended claims. For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made, by way of example only, to the accompanying diagrammatic drawings in which: Figure 1 shows a schematic ofthe AIOU, showing interconnections; Figure 2 shows a message flow illustrating an embodiment ofthe invention. Herein follows a detailed description of at least one embodiment of the invention, describing the components, operation, and implementation of the new (or existing) network function for adaptive inference that support early-exit ML operations. An aim of an embodiment of this invention is to provide a mechanism for dynamic control and orchestration of inference operations across communication network entities. This mechanism enables real-time adjustments, allowing for optimal resource allocation, reduced latencies, and maintained prediction accuracy, regardless of network conditions. In an embodiment, there is introduces the Advanced Inference Orchestration Unit (AIOU), a novel component designed for optimized AI / ML model inference within wireless networking systems. By dynamically regulating and orchestrating inference operations, the AIOU aims to enhance both the efficiency and efficacy of deploying machine learning models in real-time network scenarios. The AIOU is shown in Figure 1 and illustrates the core internal components: Network communication module, Task Decomposition Engine, Dynamic Inference Allocation, Sub-task Sync Manager and Early Exit ML module. The AIOU comprises the following features: • Integrative Framework: The AIOU is arranged to be seamlessly integrated or co-located with existing network entities or functions, whether within UE, edge / MEC servers, clouds, applications, or other infrastructure components. Its modular architecture allows for easy adaptability and scalability in diverse network environments. • Data Analysis: One of the core strengths of the AIOU is its ability to process real-time data metrics. These include network conditions (Reference Signal Received Power, RSRP, Reference Signal Received Quality, RSRQ, Signal to Interference and Noise Ration, SINR, latency, jitter, etc.), devices conditions (remaining battery, CPU utilization, GPU load, etc), latency constraints, and the specific requirements of the AI / ML task at hand. This continuous stream of data allows the AIOU to make informed decisions dynamically. • Action Management: AIOU is deeply involved in real-time decision-making processes. From initiating inference tasks to pausing, resuming, or even triggering an early-exit termination, the AIOU acts to ensure that each action is tailored to the current network conditionsand requirements. One of the key functionalities of the Advanced Inference Orchestration Unit (AIOU) involves the collaborative operation of its three primary mechanisms: the Task Decomposition Engine (TDE), the Dynamic Inference Location Allocation (DILA), and the Subtask Synchronization Manager. The TDE is responsible for dissecting complex AI / ML inference tasks into several manageable subtasks. This strategic partitioning allows the tasks to be optimized for execution within specific network elements, considering the nature and requirements of each subtask. Upon decomposition by the TDE, the DILA mechanism comes into play. It is designed to distribute these subtasks efficiently across the most optimal locations within the network, such as but not limited to UE(s), RAN, or CN. The DILA's decision-making is founded upon a multilayered decision matrix that integrates a heterogeneous set of network parameters and userspecific demands. This matrix can operate on a combination of implementation solutions, including deterministic algorithms, heuristic methods, machine learning models, and reinforcement learning. The DILA performs several critical operations: 1. Subtask Priority Assignment: Each subtask is assigned a priority level based on factors such as urgency, computational complexity, and overall task significance. For instance, subtasks critical for emergency services might be prioritized over routine tasks. 2. Subtask Decision Logic: DILA's core logic employs a mix of weighted parameters, heuristic thresholds, and potentially trained models to determine the best-suited network entity for each subtask. Considerations include computational capacity (directing resource-intensive subtasks to RAN or CN), latency requirements (favouring entities closer to the data source for real-time tasks), and network bandwidth (allocating data-intensive subtasks to entities capable of handling high bandwidth). 3. Feedback Loop: Post-execution, DILA integrates feedback regarding the efficiency, accuracy, and latency of each subtask's execution. This feedback is crucial in refining future decision-making and potentially retraining the decision models. Concurrently, the Subtask Synchronization Manager coordinates the execution of these distributed subtasks, ensuring that they are processed in a synchronized, timely, and coherent manner. This synchronization is vital to maintain the integrity of the overall inference outcome, especially when subtasks are processed across different network entities. Furthermore, the AIOU integrates several critical functional components, enhancing its capability to efficiently manage and process AI / ML tasks. These components not only facilitate the breakdown of inference tasks into smaller, more manageable subtasks but also ensure effective communication and processing within the network. These components are the following: 1. Network Communication Module: This module is responsible for receiving input data from various network nodes or devices and subsequently transmitting output predictions or interim results. It also boasts data pre-processing capabilities, ensuring the data is in an appropriate format for the early-exit ML structure. 2. Task Decomposition Engine: is responsible for breaking down complex inference tasks into smaller, more manageable subtasks. By doing so, it facilitates the distribution of these subtasks across various network entities as determined by DILA. 3. Dynamic Inference Location Allocation: DILA's primary role is to intelligently allocate subtasks, derived from the TDE, to the most suitable network entities such as UE(s), RAN, or CN. This allocation process takes into consideration an array of critical factors including computational capacity, latency requirements, and network bandwidth. Employing advanced algorithms, such as, but not limited to, heuristic, threshold based, ML, RL, etc. this engine selects the optimal subtask allocation and optimizes each subtask for specific network entities. This optimization considers factors such as computational intensity, required response time, and data sensitivity. 4. Subtask Synchronization Manager: This manager coordinates the execution of subtasks (whether in parallel or sequential) across different network entities, ensuring that the distributed processing does not compromise the integrity or timing of the final inference outcome. It continuously monitors the progress of each subtask, adjusting the synchronization parameters in real-time to account for network latency, device availability, and resource utilization. 5. Early-Exit ML Module: The AIOU is designed to be compatible with various Early-Exit ML frameworks. These frameworks are characterized by multiple decision checkpoints, each with its own set of metrics. This feature of the AIOU allows for flexible inference. Decisions at each checkpoint are made based on the confidence level of the predictions, enabling adaptable responses to different computational scenarios. The network communication module serves as the primary interface between the AIOU framework and the rest of the network, ensuring seamless communication and data exchange. This module is responsible for receiving input data from network nodes or devices, transmitting the output predictions or intermediate results, and handling data-related operations. The following are certain aspects of the network communication module: 1. Data reception and transmission: The network communication module receives input data from various sources, such as network nodes, devices, or sensors, and transmits the output predictions or intermediate results back to the appropriate recipients. 2. Data pre-processing: Before the input data can be processed by the early-exit machine learning framework, it may require pre-processing to ensure it is in a suitable format. The network communication module can include data pre-processing functionalities, such as data normalization, feature extraction, or data augmentation, to convert raw input data into a format compatible with the AIOU. 3. Data buffering and prioritization: The network communication module can implement data buffering and prioritization mechanisms to handle varying data loads, network congestion, or latency requirements. For instance, the module can buffer incoming data temporarily if the AIOU entity is not ready to process it or prioritize processing data with higher importance or stricter latency requirements. 4. Security and privacy: The network communication module should ensure the security and privacy of the data being transmitted, especially in sensitive applications or regulated environments. This can involve implementing encryption and decryption methods, secure authentication, and access control mechanisms to protect the data during transmission and storage. By handling these aspects, the network communication module ensures that input data is appropriately processed and transmitted, allowing the ML framework to focus on adaptive inference and resource optimization. The Task Decomposition Engine (TDE) is arranged to enhance the processing of complex AI / ML tasks within multiple entities of a communications network. This engine fundamentally transforms how inference tasks are managed, allowing for a more distributed and efficient approach to handling computational workloads. The TDE is responsible for dissecting complex AI / ML inference tasks into smaller, manageable subtasks. This decomposition is based on a detailed analysis of the task's intrinsic properties, such as computational requirements, data dependencies, and expected outcomes. The goal is to create subtasks that are individually simpler to process, reducing the computational load on any single network entity and facilitating parallel processing whenever possible. The DILA mechanism is designed to optimize the execution of AI / ML inference tasks within wireless networks. DILA plays a crucial role in ensuring that these tasks are executed efficiently and effectively, with minimal latency and optimal use of network resources. DILA's primary function is to allocate the execution of AI / ML tasks, or their subtasks, generated by the TDE across the most suitable network elements such as UE(s), RAN, or CN. This allocation is based on a sophisticated analysis of various network parameters and the specific demands of each task. DILA takes the inference task or the subtasks generated by the TDE and determines the best network entity for each subtask. It considers several factors including computational capacity, latency requirements, and network bandwidth to make these decisions. DILA can modify and adapt the subtasks generated by the TDE to optimize them for execution across different network entities (e.g., UE(s), RAN, CN). This optimization considers several factors, including the computational intensity of each subtask, the available processing power at each potential location, network bandwidth, and latency requirements. The aim is to allocate each subtask to the network entity where it can be processed most efficiently. At the core of DILA lies a multi-layered decision matrix that integrates a diverse set of network parameters and user-specific requirements. This matrix operates on a combination of implementation solutions such as deterministic algorithms, heuristic methods, machine learning models, and reinforcement learning techniques. To do so, DILA assigns priority levels to each subtask based on factors like urgency, computational complexity, and overall task significance. DILA decision logic mechanism can use a mix of weighted parameters and / or more advanced solution like AI / ML / RL to ascertain the optimal location for each subtask. This could mean directing computationally intensive tasks towards the RAN or CN and those requiring real-time processing closer to the data source. For subtasks involving significant data transfer, DILA allocates them to network entities that can manage such bandwidth demands efficiently. Furthermore, DILA is designed to be dynamically adaptable. It can reevaluate and adjust the allocation of subtasks in response to changing network conditions, such as fluctuations in bandwidth, alterations in device availability, or shifts in computational load across the network. This dynamic adaptability ensures that the decomposition and distribution of tasks remain optimal overtime, even as network conditions evolve. Post-execution, DILA incorporates feedback regarding the execution of each subtask. This includes aspects such as efficiency, accuracy, and latency. The feedback is used to refine the decision-making processes of DILA, enhancing its ability to make more accurate allocations in future tasks. DILA significantly enhances the efficiency of AI / ML task processing in wireless networks. By intelligently allocating subtasks to the most suitable network entities, it ensures that resources are utilized optimally, and latency is minimized. This leads to faster and more accurate AI / ML inference outcomes, crucial for real-time applications and services. Furthermore, DILA's adaptability and learning capabilities mean that it continually improves its allocation strategies. This continuous improvement is vital in dynamic network environments where conditions and requirements can change rapidly. The Subtask Synchronization Manager (SSM) role is to harmonize the execution, and define the workflow of, distributed subtasks created by the TDE. This component ensures that the distributed inference processing of AI / ML tasks across various network entities maintains coherence and efficiency, ultimately leading to accurate and timely inferences. Certain of the the main goals of this functional block are: • Coordinated Execution: The primary function of the (SSM) is to oversee the sequential or simultaneous execution of subtasks across multiple network nodes, such as UE, RAN, and CN. This involves managing the timing and sequence of subtask processing, ensuring that all components of the distributed task are synchronized, and that the final outputs are consolidated in a coherent and timely manner. • Real-Time Monitoring and Adjustment: The manager is equipped with real-time monitoring capabilities, enabling it to track the progress of each subtask across the network. It can dynamically adjust synchronization parameters in response to network latency, varying device availability, or changes in resource utilization to ensure optimal processing efficiency and minimal response time. • Feedback Integration: Post-execution feedback from each subtask is integrated into the manager's decision-making process. This feedback loop allows the manager to continually refine its synchronization strategies, improving the overall efficiency and accuracy of distributed task processing. By ensuring that all subtasks are processed in a coordinated and timely manner, the manager enhances the overall efficiency of the AIOU. This coordination is important for maintaining the accuracy of the final inference output, especially in tasks where the outputs of individual subtasks are interdependent. The manager's ability to monitor and adjust task synchronization in real-time allows the AIOU to be highly scalable and flexible. It can handle a wide range of tasks, from small-scale inferences on individual devices to large-scale computations distributed across the network. Finally, the manager's dynamic adjustment capabilities make the AIOU resilient to variations in network conditions. This resilience is important in maintaining consistent performance in wireless networks, where bandwidth and connectivity can fluctuate. The Early-Exit ML module is an element of the AIOU, specifically engineered to integrate seamlessly with various early-exit machine learning (ML) frameworks. This module brings a crucial dimension of flexibility and adaptability to the AIOU, enhancing its capability to efficiently handle AI / ML tasks across diverse network scenarios. To this end, embodiments of this invention do not mandate any specific early exit mechanism, but provide a solution for any early exit framework to be integrated into known solutions. The module is arranged to be inherently compatible with a range of early-exit ML frameworks, thus not tied to any singular early-exit mechanism. Instead, it offers a versatile solution, enabling seamless integration of diverse early-exit ML frameworks into the existing AIOU structure. This approach ensures that the AIOU remains agnostic to specific early-exit methodologies, thereby accommodating a wide range of ML strategies and enhancing its applicability across different AI / ML scenarios. This inclusivity in design allows the AIOU to adapt and evolve with emerging technologies and methodologies in the field of machine learning, ensuring its long-term relevance and effectiveness. Early-Exit frameworks are characterized by their architecture, which includes multiple decision checkpoints throughout the ML model. Each checkpoint within these frameworks is equipped with its unique set of metrics, designed to assess the confidence or accuracy of the predictions at that stage of the ML process. The core functionality of the Early-Exit ML Module lies in its ability to make real-time decisions at each checkpoint based on the assessed confidence level of the predictions. If a prediction at an early checkpoint meets or exceeds a pre-defined confidence threshold, the inference process can be concluded early. On the contrary, if a prediction is a certain threshold, the inference can be cancelled, due to inaccuracy or wrong inference. This reduces unnecessary computational overhead, saving time and resources. The module allows the adaptation of the inference behaviour, based on the computational complexity and requirements of different scenarios. In situations where high accuracy is paramount, the module can allow the inference process to continue through more checkpoints. Conversely, in scenarios where speed is more critical, or computational resources are limited, the module can make an early exit decision as soon as a sufficient confidence level is reached. By enabling the integration of AIOU with early exits frameworks in the inference process, the module significantly enhances computational efficiency. It allows the AIOU to avoid expending resources on processing steps that do not contribute to improving the outcome. Furthermore, the ability to conclude inference processes early, depending on the confidence level, allows for more flexible allocation of computational resources. This is particularly beneficial in network environments where resources are shared or limited. The Early-Exit ML Module strikes an optimal balance between processing speed and prediction accuracy. This balance is crucial in environments where network conditions and computational demands can vary significantly. A typical workflow of an embodiment involves several interconnected steps, each important for the efficient processing of AI / ML tasks. These interconnected steps may be defined as follows. There then follows a more detailed description, supported by Figures 2a and 2b. 1. Task Request from Network Node: The process initiates when a network node (such as UE, RAN, CN, or any other network entity ) sends a request for AI / ML task processing to the AIOU. This request includes the necessary data for the AI / ML task. 2 Reception and Pre-Processing by Network Communication Module: The AIOU's Network Communication Module receives this request. It then undertakes the preprocessing of the data to ensure it is in the appropriate format and state for further processing within the AIOU framework. 3. Task Decomposition by the TDE: The TDE within the AIOU takes over, breaking down the complex AI / ML task into smaller, manageable subtasks. 4. Subtask Allocation: The DILA mechanism then assesses each subtask. It assigns priority levels to each subtask based on various factors such as urgency and computational complexity and decides the most suitable network entity (UE, RAN, or CN) for executing these subtasks, based on criteria like computational capacity, latency requirements, network bandwidth, etc. 5. Execution Coordination by Subtask Synchronization Manager: Once the subtasks are allocated, the Subtask Synchronization Manager coordinates their execution across the various network entities. It ensures that all subtasks are processed in sync and that the final outputs are consolidated efficiently and accurately. 6. Early-Exit Decisions: Parallelly, if implemented, the Early-Exit ML Module monitors the confidence levels at various checkpoints for each subtask. If a subtask reaches the required confidence threshold, it can conclude early, and the results are sent back to the AIOU. This step helps in saving computational resources and time. 7. Consolidation and Output Transmission: After all subtasks are completed, the AIOU consolidates the results to form the final output of the AI / ML task. This consolidated output is then transmitted back to the initiating network node. 8. Feedback Reception: The network node, upon receiving the output, provides feedback regarding the execution, including aspects such as accuracy and latency. 9. Feedback Integration for Future Improvement: Finally, the AIOU integrates this feedback into its system. This integration aids in refining the decision-making processes of DILA and TDE for future tasks, ensuring continuous improvement in the AIOU's performance. Error! Reference source not found, and 2b (which follows from Figure 2a) show an example flow chart of a possible setup message exchange in relation to the creation / execution of a an ML inference Workload over multiple network devices using an AIOU, forming an embodiment of the invention. The figures shows message exchange between the network entities involved on the ML procedure and the AIOU. In this example, the subscription request and response procedure is as follows: • Step S1 - Initialization of ML Inference Workload: The Entity or Application (Client) that wants to perform an inference task, initiates the process by sending a request for an ML inference workload to the AIOU. This message is referred to as ML_infernce_workload, but can have any other suitable name. This message includes detailed description of the necessary data and specific requirements. These can include but are not limited to: o Nature of the Task: A clear description of the ML inference task, outlining what the client aims to achieve. This may include the type of ML model (e g., image recognition, data analysis, natural language processing). o Specific Use Case: Details about the specific application or use case for which the inference is required, providing context to the AIOU. o Input Data: The actual data on which the inference is to be performed. This could include datasets, images, text, sensor data, etc. o Data Format: Specification of the format of the input data (e g., CSV, JSON, image files) to ensure the AIOU can process it effectively. o Data Characteristics: Information about the size, complexity, and nature of the data, which can influence how the workload is handled. o Model Specification: Identification of the specific ML model or algorithm to be used, if the client has a preference. o Accuracy Requirements: Any requirements regarding the accuracy or precision of the model’s output. o Latency Sensitivity: Information about how time-sensitive the task is, which can impact where and how quickly the task needs to be processed. o Computational Constraints: Any limitations regarding computational resources or power, especially relevant if the task is to be run on devices with limited capabilities like mobile devices. o Preferred Network Entities: If the client has any preferences regarding where the task should be executed (e.g., preferring edge computing for real-time analysis). o Bandwidth Considerations: Constraints or preferences regarding the amount of network bandwidth that can be used. o Data Sensitivity: Any confidentiality or privacy concerns related to the input data. o Compliance Requirements: Any regulatory compliance requirements that need to be adhered to during the inference process. o Others. • Step S2 - Reception and Preprocessing: The AIOU Network Communication Module receives the request and prepares the data, for example, conducts preprocessing, including data normalization and formatting, to ensure compatibility with the ML model and AIOU's processing capabilities. The result is then sent to the TDE module. • Step 3 - Subtask decomposition: the TDE analyses the workload and decomposes it into smaller, manageable subtasks suited for distribution across multiple network devices. The TDE sends the details of each subtask, including computational requirements and data dependencies, to the DILA. • Step 4 - Subtask Allocation Across Network Devices: the AIOU DILA assesses each subtask and identifies the most suitable network device (e.g., UE, edge servers, RAN, or CN) for its executions using state of the art AI / ML solutions, or any other solution that is based on factors like computational capacity, latency requirements, and network bandwidth of each device. The different inference subtasks are then sent in step S5 to the distinct network components, using the message referred to ML_infernce_subtask_execution or any other suitable name. The content of the message can be, but is not limited to: o Description: Detailed information about each subtask allocated to the specific network entity. This includes the nature of the subtask, its computational requirements, and the expected outcome. o Data: Any relevant data or parameters needed for the subtask's execution. o Parameters: Specific execution parameters tailored to the capabilities and current load of the network entity. o Timing: Instructions regarding the timing of the subtask execution, especially important if the subtask is time-sensitive or needs to be synchronized with other tasks. o Resource Allocation: Details about the amount and type of resources (such as CPU, memory, bandwidth) that should be allocated to the subtask. o Efficiency Recommendations: Suggestions or guidelines for efficient resource utilization, considering the overall network status and the specific capabilities of the network entity. o Latency Requirements: Information on expected or acceptable latency levels for the subtask's processing. o Performance Metrics: Expected performance metrics or standards that the subtask should meet, which could be specific to the type of AI / MLtask. o Data Security: Instructions on handling sensitive data, including encryption and privacy measures, if applicable o Error Reporting: Protocols for reporting and handling errors or unexpected issues during the subtask's execution. o Fallback Procedures: Contingency measures or fallback procedures in case the network entity is unable to execute the subtask as planned. o Others • Step S6 (optional) - Early exit criteria: If an early exit criteria is defined, the EE module sends to the network entities a message, defining the specific early exit points corresponding to the specific subtask. The message, ML_inference_early_exit_criteria, can include, but is not limited to, information fields such as: o Subtask Identifier: A unique identifier for the subtask to which the early exit criteria apply. o Associated Task Description: Brief description of the subtask for clarity and reference. o Defined Early Exit Checkpoints: Specific checkpoints within the subtask where an early exit decision can be made. o Criteria at Each Point: Detailed description of the criteria to be met at each checkpoint for an early exit decision. o Threshold Levels: The required confidence levels at each early exit point that must be reached or exceeded. o Metric Specifications: Details of the metrics used to measure confidence (e.g., probability scores, error rates). o Latency Considerations: Instructions on how latency should be factored into the early exit decision. o Resource Utilization Limits: Guidelines on resource usage limits that, if reached, may trigger an early exit. o Data Reporting: Instructions on how to handle and report data if an early exit is triggered. o Result Submission Format: Specification of the format in which results should be submitted upon early exit o Others. • Step S7 - Coordination and Synchronization of Subtask Execution: The AIOU Subtask Synchronization Manager ensures synchronized and efficient execution of subtasks across the allocated network devices by sending messages that content periodic synchronization commands, timing adjustments, and execution status updates, among other possible commands. These messages are referred to ML_infernce_sync_execution, and the inference starts after the reception of the first message of such type. The content these type of synchronization messages can provide are the following: o Synchronisation Protocols: Establishes synchronisation protocols to ensure that subtasks executed across different network entities are processed in a harmonious and synchronized fashion, such as precise timing and sequencing instructions, synchronisation checkpoints, and coordination protocols. o Real-Time Monitoring and Adjustment: The Subtask Synchronization Manager actively monitors the progress of each subtask, ensuring they are adhering to the planned execution timeline, requesting real-time updates on subtask status, monitoring data, and progress reports. o Dynamic Synchronization Adjustments: Based on real-time data, the manager may make dynamic adjustments to the synchronization parameters to account for any delays, resource constraints, or unexpected network conditions. The adjustments can include adjusted synchronization instructions, revised timing or sequencing parameters, and contingency measures. o Consistency and Dependency Management: Ensuring that subtasks with dependencies are processed in the correct order and that their outputs are correctly integrated. Done by verifying the dependency fulfilment, instructions for handling output integration, and consistency checks. • Step S8 - Execution and Monitoring: The designated network devices execute the allocated subtasks, potentially utilizing Early-Exit ML Modules for efficiency if available. The AIOU continuously monitors the execution, adjusting strategies as needed for optimal performance. These entities send execution progress, early-exit decisions, and interim results periodically, using a message referred to as ML_infernce_sync_reporting. The content of this message, while not limited to, would typically include the following elements: o Progress Update: Real-time updates on the current status of the subtask execution, including completion percentage or stages reached. o Timina Information: Elapsed time and expected time to completion, providing insight into the pace ofthe subtask execution. o Early-Exit Decisions: Notifications if an early-exit decision has been made for a subtask, indicating that the necessary confidence threshold has been met and further processing is not required. o Preliminary Outputs: Partial results or findings from the subtask, particularly important if the task was concluded early or is part of a larger, ongoing analysis. o Data Samples: Samples of processed data or intermediate calculations, if relevant for monitoring and quality assurance. o Computational Load: Information on the computational resources used for the subtask, such as CPU and memory usage. o Network Bandwidth Consumption: Data on the amount of network bandwidth utilized during the subtask execution. o Accuracy Indicators: Metrics or indicators related to the accuracy or efficacy of the subtask, particularly relevant for AI / ML tasks. o Latency Metrics: Measurements of any latency experienced during the execution of the subtask. o Error Logs: Details of any errors or issues encountered during the subtask execution. o Anomaly Detection: Reports of any anomalies or unexpected behaviours observed, which could be critical for adaptive adjustments. o Others. • Step S9 (Optional) - Early Exit Trigger. If any network entity performing an inference subtask triggers early termination based on the criteria set by the Early-Exit ML Module, a notification is sent to the AIOU's Subtask Synchronization Manager. The message is referred to as ML_infemce_early_exit, and includes an EarlyExitlnaccuratelnference indicator when applicable,, and contains information such as, but not limited to: o Early Exit Notification: Indicates that an early exit has been triggered for a specific subtask. o Reason for Early Exit: Details on why the early exit was initiated, such as achieving sufficient confidence levels in the results. This field contains the "EarlyExitlnaccuratelnference" indicator, which flags if the early exit was due to the detection of inaccurate results or deviations from expected performance metrics, triggering a proactive termination to prevent further resource consumption on potentially flawed processing. o Subtask Results: The current results or findings of the subtask up to the point of early exit. o Resource Utilization: Information on the resources consumed by the subtask until the point of early exit. o Time of Exit: The timestamp indicating when the early exit was triggered, o Others • Step S10 - Consolidation of Results: Once a result is obtained, either by early exit or because the whole inference process has finished, the AIOU Subtask Synchronization Manager collects and consolidates the results of all subtasks into a final output and sends a message that we refer to as ML_infernce_subtask_result or any other suitable name to the communication module. An example of the type of content of the message is the following: o Completion Status: Information on which subtasks have been completed or terminated early. o Subtask Results: Compiled results from all completed or early-exited subtasks. o Resource Summary: Summary of resources utilized across all subtasks, aiding in efficiency analysis. o Synchronization Report: A report on the synchronization status of subtasks, highlighting any discrepancies or delay o Others • Step S11- Final Output:. The finalized ML inference results, ready for delivery to the client. This is done using a ML_inference_result. An example of the type of content of the message is the following: o Final Inference Results: The conclusive results of the entire ML inference workload. o Performance Metrics: Detailed metrics on the performance, accuracy, and efficiency of the inference process. o Resource Utilization Report: A comprehensive report on the resources utilized during the entire process. o Feedback Request: A solicitation for feedback on the results and the overall process. o Others. • Step S12 - Delivery and Feedback: The AIOU delivers the final output to the client or application that initiated the request. This client acknowledges the reception and provides feedback on the results and performance, using a message labelled as ML_infernce_ack. The AIOU integrates this feedback to refine and improve future workload processing. The efficacy of AIOU is distinctly evident in its capability to pause and migrate inference tasks between network entities based on real-time challenges such as changing network conditions, computational needs, or task exigencies. Below are some illustrative scenarios that shed light on AIOU's flexible orchestration: 1. Autonomous Vehicle Navigation: o Scenario: An autonomous vehicle (AV) utilizes its onboard Al system (integrated within the UE) for navigation, processing live data streams that include vehicle telemetry, road status, and real-time traffic dynamics. o AIOU Action: Initially, the Al model processes inference locally within the AV. But as the AV manoeuvres into an area with intense traffic and intricate intersection patterns, the onboard computation starts to strain. AIOU intelligently pauses the onboard analysis and transitions the task to a proximate edge server, specifically optimized for such demanding traffic scenarios. Once the AV exits this challenging zone, the AIOU may decide to return the task to the vehicle's native system. 2. Remote Health Monitoring: o Scenario: A patient, from the comfort of their home, wears a health monitoring apparatus that persistently measures vital signs and employs predictive algorithms to identify potential health anomalies. o AIOU Action: Upon noticing an irregularity, the device's UE embarks on a detailed analysis. However, given a depleting battery, AIOU halts the local inference and delegates the task to an edge server within the hospital infrastructure for an in-depth assessment. After this detailed analysis, the task can either revert to the UE for ongoing surveillance or remain on the edge, based on battery conservation strategies. 3. AR-based Retail Experience: o Scenario: A consumer dons AR glasses and immerses himself or herself in a virtual retail environment, experimenting with "trying on" various outfits. The AR application, rooted in the UE, offers instant rendering and feedback. o AIOU Action: As the consumer toggles a feature that simulates varying ambient light effects on the attire, the computational load surges. Predicting a probable latency or performance decline, AIOU proactively pauses the UE-centric inference. It then shifts the resource-intensive lighting simulation module to an edge server within the retail facility. When the user reverts to basic interactions, AIOU seamlessly transfers the processing back to the UE, ensuring fluidity in experience. At least some of the example embodiments described herein may be constructed, partially or wholly, using dedicated special-purpose hardware. Terms such as ‘component’, 'module' or 'unit' used herein may include, but are not limited to, a hardware device, such as circuitry in the form of discrete or integrated components, a Field Programmable Gate Array (FPGA) or Application Specific Integrated Circuit (ASIC), which performs certain tasks or provides the associated functionality. In some embodiments, the described elements may be configured to reside on a tangible, persistent, addressable storage medium and may be configured to execute on one or more processors. These functional elements may in some embodiments include, by way of example, components, such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. Although the example embodiments have been described with reference to the components, modules and units discussed herein, such functional elements may be combined into fewer elements or separated into additional elements. Various combinations of optional features have been described herein, and it will be appreciated that described features may be combined in any suitable combination. In particular, the features of any one example embodiment may be combined with features of any other embodiment, as appropriate, except where such combinations are mutually exclusive. Throughout this specification, the term “comprising” or “comprises” means including the component(s) specified but not to the exclusion of the presence of others. Attention is directed to all papers and documents which are filed concurrently with or previous to this specification in connection with this application and which are open to public inspection with this specification, and the contents of all such papers and documents are incorporated herein by reference. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification (including any accompanying claims, abstract and drawings) may be replaced by alternative features serving the same, equivalent or similar purpose, unless expressly stated otherwise. Thus, unless expressly stated otherwise, each feature disclosed is one example only of a generic series of equivalent or similar features. The invention is not restricted to the details of the foregoing embodiment(s). The invention extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.

Claims

1. A method of operating a telecommunication network comprising the steps of:providing an orchestration entity arranged to receive at least one input from a first network entity and to transmit at least one output to a second network entity, wherein the orchestration entity is arranged to determine, on a real-time basis, where in the telecommunication network, and when, an AI / ML inference task is performed.

2. The method of claim 1 wherein the orchestration entity comprises a Task Decomposition Entity arranged to dissect complex AI / ML inference tasks into a plurality of subtasks and to optimise each of the plurality of subtasks for execution within one or more specific network entities based on the computational demands and nature of the subtasks.

3. The method of claim 2 wherein the orchestration entity comprises a Dynamic Inference Location Allocation entity arranged to determine a suitable network entity for executing one or more of the plurality of subtasks.

4. The method of claim 3 wherein the orchestration entity comprises a Subtask Synchronization Manager arranged to coordinate the execution of the plurality of subtasks distributed across various network entities and to ensure that the plurality of subtasks are synchronized in terms of timing, sequence, and consistency, thereby maintaining the integrity and efficiency of the final inference outcome.

5. The method of claim 4 wherein the plurality of subtasks are either run in parallel or in sequence.

6. The method of any preceding claim wherein early-exit termination of the inference task if at least one condition is met.

7. The method of claim 6 wherein the at least one condition is one or more of: that a certain time has elapsed; that a satisfactory inference result has been achieved; that the output is inaccurate; and a local status of the entity performing the inference.

8. The method of any preceding claim wherein the first network entity is the same as the second network entity.

9. Apparatus arranged to perform the method of any preceding claim.

Citation Information

Patent Citations

  • A method for DNN task unloading

    CN112822264B

  • Machine learning inference task deployment method for throughput optimization based on cloud-edge collaboration

    CN113315669A

  • AI on-demand service method based on 6G network

    CN116801219A

  • Fluid Client Server Partitioning of Machines Learning, AI Software, and Applications

    US20200106829A1

  • Subtask Assignment for an Artificial Intelligence Task

    US20210319281A1