Devices and methods for split-ai in-network computing in a mobile network
The control plane entity in 3GPP mobile networks optimizes neural network distribution across user equipment and network nodes by selecting rules based on system capabilities and conditions, addressing inefficiencies in existing mobile networks and enhancing performance indicators.
Patent Information
- Application Number
- PCT/EP2024/064593
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-28
- Publication Date
- 2025-12-04
AI Technical Summary
Existing mobile networks face challenges in efficiently distributing neural network layers across user equipment, access nodes, and user plane functions due to varying hardware capabilities, privacy concerns, and dynamic network conditions, which affect key performance indicators such as latency, energy consumption, and traffic volume.
A control plane entity dynamically selects neural network distribution rules based on system capabilities and real-time network conditions to optimize the allocation of neural network layers across user equipment, access nodes, and user plane functions, leveraging In-Network Computing (INC) in 3GPP mobile networks.
This approach enhances the dynamicity and efficiency of neural network distribution, allowing mobile network operators to tailor and optimize split-AI realization, minimizing inference duration, network traffic, and UE energy consumption while ensuring privacy and compliance with network conditions.
Smart Images

Figure EP2024064593_04122025_PF_FP_ABST
Abstract
Description
[0001] DEVICES AND METHODS FOR SPLIT-AI IN-NETWORK COMPUTING IN A MOBILE NETWORK
[0002] TECHNICAL FIELD
[0003] The present disclosure relates to wireless communications. More specifically, the present disclosure relates to devices and methods for supporting Split-AI, leveraging In-Network Computing, in a mobile network, in particular a 3GPP mobile communication system.
[0004] BACKGROUND
[0005] In-Network Computing (INC) is a concept for carrying out computations within network nodes (e.g. routers, switches), that are typically only used to forward the traffic. With INC, the nodes are capable of computing on network packets, on top of transmitting them. In the context of mobile networks, it means that the Access Nodes (AN) and User Plane Functions (UPFs) compute on the packets of a network flow. In-Network Computing is considered as an enabler for integrating native computing as a key feature into 6G networks.
[0006] A Neural Network (NN) is a specific Machine Learning (ML) / Artificial Intelligence (Al) technique, whose learning technique is inspired by the human brain. NNs are mostly used for tasks such as object detection, image recognition, machine translation, or speech recognition. NNs consist of a plurality of processing layers, including one input layer, one output layer, and one or more hidden layers. Each layer consists of a set of neurons, a neuron implementing a so-called activation function. The activation function (and hence the neuron) receives input values, which are transformed to an output by computing the activation function. Well-known and typically used activation functions for NNs are tanh (hyperbolic tangent), ReLU (Rectified Linear Unit), sigmoid, softmax, or binary step. The output of a neuron from layer n is passed forward and serves as an input of a neuron from layer n+1.
[0007] Split-AI a refers to a vertical splitting of an NN between its layers. In a communication network, the execution of the different layers may be split up between a user device, e.g. UE and an application server (which is referred to as split-AI without INC support in the following). If not only the UE and the application server do the computation, but as well the networking equipment (i.e. with the support of INC), then the NN execution can be split between UE, server, and the network nodes, e.g. routers and switches. In the case of mobile networks, the execution may be split between UE, server, access node(s), and user plane functions. Herein this is referred to as split-AI with INC support or INC-assisted split-AI. The layers, at which a NN is split, and how the different processing layers are allocated to the network nodes (UE, AN, UPF(s), server), impacts several Key Performance Indicators (KPIs), which are relevant to different stakeholders (end user, application, network). These KPIs, which may be differently affected by the split, comprise, in particular, end-to-end latency (also referred to as inference latency), UE energy consumption, the overall traffic volume, and the overall energy consumption.
[0008] As will be appreciated, not necessarily are all combinations / allocations of NN layers to compute nodes viable. This can have various reasons, the most obvious being described in the following. In case of INC-assisted split-AI, the MNO is carrying out computing tasks for the application. Hence, agreements may exist, covering financial issues as well as clarifying any possible privacy concerns (for handling the user’s application data). The agreements between MNO and application provider may also cover different scopes when it comes to split-AI. For example, the application provider may only agree that the MNO assists in finding an allocation of NN execution between UE and server. That is, no computation is carried out in the network, all NN layer execution resides in the application (either at UE or server or both). However, the MNOs knowledge of the underlying network conditions is used to determine how the NN layer execution is split between UE and server. On the other hand, if there is a stronger agreement between MNO and application provider, the network may be actively involved in NN layer execution. Furthermore, privacy constraints may relate to the application, the user, or the device settings. Specific privacy constraints may lead to requirements such as “the first layer may be executed locally at the UE”, which would prevent personal data leaving the personal device. After a first step of processing, it may be hard or impossible to reconstruct the original data, thus ensuring the users privacy. Moreover, not all nodes in the network (AN, UPF(s)) may support In- Network Computing. Further, even if they do have INC support, they are not necessarily capable of executing certain neurons or layers (e.g. because the activation function is unknown to them). The same applies to the UE, it may just not be powerful enough to be involved in the Al process.
[0009] SUMMARY
[0010] It is an objective of the present disclosure to provide improved devices and methods for supporting Split-AI In-Network Computing in a mobile network, in particular a 3GPP mobile communication system, such as a 5G or 6G network.
[0011] The foregoing and other objectives are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.
[0012] According to a first aspect a control plane entity for supporting the distribution, i.e. splitting of a neural network, NN, across a plurality of communicate and compute nodes of a mobile network, including a user equipment, UE, one or more access nodes, one or more user plane functions, and / or an application server, wherein the control plane entity is configured to dynamically select one of a plurality of selectable NN distribution rules, e.g. a list of selectable NN distribution rules based on one or more of the following: a NN distribution rule selection policy; capability information about the system, e.g. hardware capabilities each of the plurality of communicate and compute nodes; and / or monitoring data indicative of one or more dynamic conditions of the mobile network. Thus, the control plane entity according to the first aspect may determine based on different input data the NN distribution rule most suited for the current state of the mobile network. For instance, the monitoring data indicative of one or more dynamic conditions of the mobile network may comprise load information about the current communication load and / or computation load of each of the plurality of communicate and compute nodes. This allows the control plane entity according to the first aspect to select the NN distribution rule most suited for the current communication and / or computation load of the plurality of communicate and compute nodes.
[0013] Thus, an improved control plane entity is provided for supporting the distribution, i.e. splitting of a neural network, NN, across a plurality of communicate and compute nodes of a mobile network. More specifically, the control plane entity according to the first aspect allows providing a high degree of dynamicity (in terms of distributing the NN), according to the current situation, needs, and state. It further provides the capability to the mobile network operator to tailor and optimize the split-AI realization in its network.
[0014] In a further possible implementation form, the control plane entity according to the first aspect is further configured to compute a NN distribution, i.e. allocation across the plurality of communicate and compute nodes of the mobile network based on the selected NN distribution rule (i.e. the NN distribution rule selected from the plurality of selectable NN distribution rules) or to provide information about the selected NN distribution rule to a further control plane entity for computing a distribution, i.e. allocation across the plurality of communicate and compute nodes based on the selected NN distribution rule by the further control plane entity. Thus, the control plane entity according to the first aspect enables a specific description of what to optimize when implementing split-AI for a specific user, device, or application.
[0015] In a further possible implementation form, the control plane entity or the further control plane entity is configured to compute the NN distribution, i.e. allocation across the plurality of communicate and compute nodes based on the selected NN distribution rule by determining a NN distribution optimizing one or more key performance indicators defined by the selected NN distribution rule. This allows the control plane entity according to the first aspect to optimize the NN distribution for one or more key performance indicators defined by the mobile network operator or a third party providing the NN. Different possible key performance indicators will be described further below.
[0016] In a further possible implementation form, the control plane entity is further configured to provide, based on the computed NN distribution, the NN across the plurality of communicate and compute nodes of the mobile network or to provide information about the computed NN distribution to a further control plane entity for distributing the NN across the plurality of communicate and compute nodes of the mobile network by the further control plane entity based on the computed NN distribution. This allows to efficiently distribute the NN across the plurality of communicate and compute nodes of the mobile network.
[0017] In a further possible implementation form, the NN comprises a plurality of NN processing layers and the control plane entity according to the first aspect or the further control plane entity is configured to distribute the NN across the plurality of communicate and compute nodes based on the computed NN distribution by allocating a plurality of subsets of the plurality of NN processing layers to the plurality of communicate and compute nodes of the mobile network. This allows to efficiently distribute subsets of the plurality of NN processing layers across the plurality of communicate and compute nodes of the mobile network.
[0018] In a further possible implementation form, the mobile network is a 5G mobile network, wherein the control plane entity is a Policy Control Function, PCF, entity of the 5G mobile network and the further control plane entity is a Session Management Function, SMF, entity of the 5G mobile network. This allows to seamlessly integrate the control plane entity according to the first aspect or the further control plane entity in a 5G mobile network.
[0019] In a further possible implementation form, the control plane entity according to the first aspect is configured to receive the NN distribution rule selection policy from a management plane entity, in particular a Network Management System, NMS, of the mobile network. This allows for a well-defined configuration of the control plane entity according to the first aspect by the NMS of the mobile network.
[0020] In a further possible implementation form, the control plane entity is configured to receive the monitoring data from a user plane function of the mobile network. This allows to the control plane entity according to the first aspect to receive current monitoring data about the current load state of the mobile network.
[0021] In a further possible implementation form, the plurality of selectable NN distribution rules comprise one or more of the following: a NN distribution rule minimizing inference duration; a NN distribution rule minimizing network traffic volume; and / or a NN distribution rule minimizing UE energy consumption. This allows the control plane entity according to the first aspect to select a NN distribution rule optimizing one or more key performance indicators defined by the mobile network operator or a third party providing the NN.
[0022] According to a second aspect a method is provided of operating a control plane entity for supporting the distribution, i.e. splitting of a neural network, NN, application across a plurality of communicate and compute nodes of a mobile network, including a user equipment, UE, one or more access nodes, one or more user plane functions, and / or an application server, wherein the method comprises dynamically selecting one of a plurality of selectable NN distribution rules based on one or more of the following: a NN distribution rule selection policy; capability information, i.e. system capability information about each of the plurality of communicate and compute nodes; and / or monitoring data indicative of one or more dynamic conditions of the mobile network, e.g. monitoring data about the current communication load and / or computation load of each of the plurality of communicate and compute nodes.
[0023] The method according to the second aspect can be performed by the control plane entity according to the first aspect. Thus, further features of the method according to the second aspect result directly from the functionality of the control plane entity according to the first aspect as well as its different implementation forms described above and below. The method according to the second aspect allows determining based on different input data the NN distribution rule most suited for the current state of the mobile network.
[0024] According to a third aspect a control plane entity is provided for distributing, i.e. splitting a neural network, NN, application across a plurality of communicate and compute nodes of a mobile network, including a user equipment, UE, one or more access nodes, one or more user plane functions, and / or an application server. The control plane entity according to the third aspect is configured to obtain information about a selected NN distribution rule and to compute a NN distribution, i.e. allocation across the plurality of communicate and compute nodes of the mobile network based on the selected NN distribution rule. Thus, the control plane entity according to the third aspect allows computing a NN distribution based on the selected NN distribution rule which may be most suitable for the current state of the mobile network.
[0025] In a further possible implementation form, the control plane entity according to the third aspect is configured to compute the NN distribution, i.e. allocation across the plurality of communicate and compute nodes of the mobile network based on the selected NN distribution rule, capability information, i.e. information about the system capability of each of the plurality of communicate and compute nodes and / or monitoring data indicative of one or more dynamic conditions of the mobile network, in particular load information about the current communication load and / or computation load of each of the plurality of communicate and compute nodes of the mobile network. This allows the control plane entity according to the third aspect to determine the NN distribution most suited for the current communication and / or computation load of the plurality of communicate and compute nodes
[0026] In a further possible implementation form, the control plane entity according to the third aspect is configured to compute the NN distribution, i.e. allocation based on the selected NN distribution rule by determining a NN distribution optimizing one or more key performance indicators defined by the selected NN distribution rule. This allows the control plane entity according to the third aspect to determine a NN distribution based on the NN distribution rule that is optimized for one or more key performance indicators.
[0027] In a further possible implementation form, the control plane entity according to the third aspect is further configured to distribute, based on the computed NN distribution, the NN across the plurality of nodes of the mobile network or to provide information about the computed NN distribution to a further control plane entity for distributing the NN across the plurality of nodes of the mobile network by the further control plane entity, based on the computed NN distribution. This allows the control plane entity according to the third aspect to efficiently distribute the NN across the plurality of nodes of the mobile network.
[0028] In a further possible implementation form, the NN comprises a plurality of NN processing layers and the control plane entity according to the third aspect is configured to distribute the NN across the plurality of nodes based on the computed NN distribution by allocating a plurality of subsets of the plurality of NN processing layers to the plurality of nodes of the mobile network. This allows the control plane entity according to the third aspect to efficiently distribute the NN processing layers across the plurality of nodes of the mobile network. In a further possible implementation form, the mobile network is a 5G mobile network, wherein the control plane entity is a Session Management Function, SMF, entity of the 5G mobile network.
[0029] In a further possible implementation form, the control plane entity according to the third aspect is configured to obtain the information about the selected NN distribution rule by receiving the information about the selected NN distribution rule from a further control plane entity, e.g. a further control plane entity according to the first aspect.
[0030] In a further possible implementation form, the further control plane entity is a Policy Control Function, PCF, entity of the 5G mobile network.
[0031] According to a fourth aspect a method is provided for operating a control plane entity for distributing, i.e. splitting a neural network, NN, application across a plurality of communicate and compute nodes of a mobile network, including a user equipment, UE, one or more access nodes, one or more user plane functions, and / or an application server. The method according to the fourth aspect comprises the steps of: obtaining information about a selected NN distribution rule; and computing a NN distribution, i.e. allocation of the NN across the plurality of communicate and compute nodes of the mobile network based on the selected NN distribution rule.
[0032] The method according to the fourth aspect can be performed by the control plane entity according to the third aspect. Thus, further features of the method according to the fourth aspect result directly from the functionality of the control plane entity according to the third aspect as well as its different implementation forms described above and below.
[0033] According to a fifth aspect, a computer program product is provided, comprising a computer-readable storage medium for storing a program code which causes a computer or a processor to perform the method according to the second aspect or the method according to the fourth aspect, when the program code is executed by the computer or the processor.
[0034] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0035] BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In the following, embodiments of the present disclosure are described in more detail with reference to the attached figures and drawings, in which:
[0037] Fig. 1 is a schematic diagram illustrating an example of the input, internal logic, and output of a neuron of a neural network;
[0038] Fig. 2 is a schematic diagram illustrating an example of a fully connected neural network with different layer characteristics;
[0039] Fig. 3 is a schematic diagram illustrating components of a mobile network, including a control plane entity according to an example for distributing a neural network across a plurality of communicate and compute nodes of the mobile network;
[0040] Figs. 4 and 5 are schematic diagram illustrating further details of a control plane entity according to an example for distributing a neural network across a plurality of communicate and compute nodes of the mobile network; Fig. 6 is a schematic diagram illustrating the architecture of a mobile network, including a control plane entity according to an example for distributing a neural network across a plurality of communicate and compute nodes of the mobile network;
[0041] Fig. 7 is a schematic diagram illustrating components of a mobile network, including a control plane entity according to a further example for distributing a neural network across a plurality of communicate and compute nodes of the mobile network;
[0042] Figs. 8a-c are signalling diagrams illustrating different stages of operations and interactions of a control plane entity according to an example with other network entities of a mobile network for distributing a neural network across a plurality of communicate and compute nodes of the mobile network;
[0043] Fig. 9 is a flow diagram illustrating a method for operating a control plane entity according to an example for supporting the distribution of a neural network across a plurality of communicate and compute nodes of a mobile network; and
[0044] Fig. 10 is a flow diagram illustrating a method for operating a control plane entity according to an example for distributing of a neural network across a plurality of communicate and compute nodes of a mobile network.
[0045] In the following, identical reference signs refer to identical or at least functionally equivalent features.
[0046] DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] In the following description, reference is made to the accompanying figures, which form part of the disclosure, and which show, by way of illustration, specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that embodiments of the present disclosure may be used in other aspects and comprise structural or logical changes not depicted in the figures. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims.
[0048] For instance, it is to be understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or a plurality of specific method steps are described, a corresponding device may include one or a plurality of units, e.g. functional units, to perform the described one or plurality of method steps (e.g. one unit performing the one or plurality of steps, or a plurality of units each performing one or more of the plurality of steps), even if such one or more units are not explicitly described or illustrated in the figures. Moreover, if a specific apparatus is described based on one or a plurality of units, e.g. functional units, a corresponding method may include one step to perform the functionality of the one or plurality of units (e.g. one step performing the functionality of the one or plurality of units, or a plurality of steps each performing the functionality of one or more of the plurality of units), even if such one or plurality of steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically noted otherwise.
[0049] Before describing detailed embodiments of devices and methods for supporting Split-AI, leveraging In-Network Computing, in a mobile network, in particular a 3GPP mobile network, some acronyms / abbreviations and definitions are introduced as well as some technical background helpful for understanding the detailed embodiments disclosed herein.
[0050] Access Node AN
[0051] Application Function AF Artificial Intelligence Al
[0052] Communication and Compute Flow CC Flow
[0053] Control Plane CP
[0054] In-Network Computing INC
[0055] Intelligent User Plane IUP
[0056] Key Performance Indicator KPI
[0057] Machine Learning ML
[0058] Management Plane MP
[0059] Mobile Network Operator MNO
[0060] Network Function NF
[0061] Network Management System NMS
[0062] Neural Network NN
[0063] Policy Control Function PCF
[0064] Session Management Function SMF
[0065] User Equipment UE
[0066] User Plane UP
[0067] User Plane F unction UPF
[0068] As used herein, a communicate and compute is an entity with capability of sub-NN computation. That is, communicate compute nodes can execute allocated sub-NNs, i.e. a set of layers of an NN. Communicate and compute nodes may include the end-devices (e.g. and UE and an application server), as well as networking devices on the connecting path between the UE and the server, e.g. AN and UPF(s) in the case of a mobile communication system.
[0069] A Communication and Compute flow (or CC Flow) differs from conventional flows in the sense of providing computation on top of communication. The packets of a CC Flow are not only forwarded in the network (UP) entities, but their pay load data is potentially modified by carrying out computations on them.
[0070] An allocated sub-NN refers to parts (i.e. layers) of an NN that are allocated for execution to a communicate and compute node. A sub-NN can comprise 0 layers (nothing is allocated to that node), all N layers of the NN (the whole NN is allocated to that node), or an arbitrary number between 1 and N-l of consecutive layers (a certain part of the NN is allocated to that node).
[0071] A Microservice is a unit of a set of computations, which is carried out on a CC Flow.
[0072] Figure 1 denotes inputs, weights, internal logic, and output of the neuron 10 of a NN. It shall be used in the following to explain the key terms (weights, bias, activation function). The input values (denoted as x0, x1;x2, xd) are coming from the neurons of the previous layer, i.e. x0, x1;x2, ■ ■ ■, xdare the outputs of the neurons from the preceding layer. To each input, a different weight is given, denoted as w0, w1;w2, ..., wd. All weighted inputs (x;* w;) are linearly combined and an additional constant is added, referred to as bias (bt). The resulting linear combination is then fed into the activation function (f), the core of the neuron 10. The output of the neuron 10 is either used as an input for neurons of the next layer, or (in case the denoted neuron 10 is the output layer), it is the final output, i.e. the predicted value (inference goal of the NN).
[0073] Figure 2 illustrates an exemplary NN 20 for image recognition, which may be split across the nodes of a mobile network by the control plane entities described in the following. By way of example, the photo of a tree is the input of the NN 20 shown in figure 2. Several intermediate / hidden NN processing layers 20a.d follow, before it is finally returned what type of object could be recognized from the photo. As will be appreciated the intermediary NN layers 20a-e can differ with respect to the complexity and / or the data size. The complexity of a layer 20a-e is usually determined by the number of neurons 10 contained in a layer and the complexity of the neurons’ activation functions. With respect to the data size some layers of the NN 20 reduce the data size (i.e. they receive more input than they produce output), while other layers may increase the data size (i.e. they receive less input than they produce output).
[0074] Figure 3 is a schematic diagram illustrating components of a mobile network 100, including a control plane entity 110a according to an embodiment for supporting the distribution of a neural network, such as the example NN 20 of figure 2, across a plurality of communicate and compute nodes of the mobile network 100. These communicate and compute nodes of the mobile network 100 may include a UE 140, an access node 150, one or more UP entities 130, and an application server 190a, as illustrated in figures 6 and 8. The control plane entity 110a may be implemented as a standalone control plane entity 110a or together with a further control plane entity 110b as functional components of a control plane entity 110, as illustrated in figure 3 and will be described in more detail in the following. In the embodiment shown in figure 3, the mobile network 100 is a 6G 3GPP mobile network 100, which includes in addition to the control plane entities 110, 110a, b a network management system 120 and one or more user plane entities 130, in particular UP functions 130.
[0075] As will be appreciated, the amount of possible NN splits and NN layer allocations to the communicate and compute nodes in a mobile network is a complex problem. Even for a simple case of splitting a NN 20 with only four layers 20a-e and allocating the NN subsets to four possible, pre-defined compute nodes (e.g. UE 140, AN 150, UPF 130, server 190a), this already leads to 35 different options for how the layers can be allocated to the communicate and compute nodes. The number of possible allocations increases exponentially with the number of compute node options and the number of NN layers. This issue is addressed by the control plane entity 110a and the control plane entity 110b according to embodiments disclosed herein or the control plane entity 110 comprising the control plane entity 110a and the control plane entity 110b.
[0076] As will be described in more detail below under further reference to figure 4, the control plane entity 110a or the control plane entity 110 comprising the control plane entity 110a is configured to select one of a plurality of NN distribution rules 402a-m based on one or more of the following: a NN distribution rule selection policy 401 : capability information about each of the plurality of communicate and compute nodes; and / or monitoring data indicative of one or more dynamic conditions of the mobile network 100.
[0077] Thus, as illustrated by stages 1 and 2 in figure 4, the NN distribution rule selection policy 401 is used for selecting an appropriate NN distribute rule from a set of available, i.e. selectable NN distribution rules 402a-m (referred to as split-AI rules 402a-m in figure 4). In a stage 2 of figure 4, the appropriate rules is selected from the set of split-AI rules 402a-m for determining how to split the NN 20 and how to allocate the layers to the compute nodes, whereby the selection of the appropriate split-AI rule is performed based on the policy 401.
[0078] In an embodiment, the policy 401 may be provided by the network operator. That is, the CP entity 110a is provided with the policy by means of being programmed by the Management Plane (MP) 120. This allows to consider the operator’s preferences when it comes to carrying out split-AI. In an embodiment, additional inputs associated with the underlying conditions of the mobile network 100 and / or the system capabilities of the communicate and compute nodes may be provided to the control plane entity 110a and / or the control plane entity 110b for selecting an appropriate NN distribute rule 402a-m in accordance with the NN distribution rule selection policy 401 and / or determining a NN distribution across the plurality of communicate and compute nodes of the mobile network 100 based on the selected NN distribution rule 402a-m. More specifically, the underlying conditions refers to monitoring data from the UP, including the current UL / DL volume, the compute load at the communicate and compute nodes in the UP, and any other relevant dynamic information. The underlying conditions may further include UE information. That is, the compute capabilities (e.g. CPU, RAM) of the device and its current compute pressure. Also, the energy efficiency as well as the current battery load of the device can be relevant information for deciding about the proper split. Also, the CP can leverage information regarding the application’s requirements when determining the split. This can include possible privacy constraints or the required level of inference accuracy or requirements on the end-to-end inference latency. As compared to the underlying conditions, the information about the system’s capabilities may be more static. It mostly refers to information relating to the UP topology, including the compute capabilities of the UP nodes, RAM / CPU / GPU equipment, capabilities for executing certain activation functions (and hence NN layers), energy efficiency information, and so on.
[0079] Figure 5 illustrates the workflow implemented by the control plane entity 110a and / or the control plane entity 110b for determining the appropriate split- Al rule. The control plane entity 110a and / or the control plane entity 110b: (0.1 ) are programmed by the MP 120 with a split-AI policy 401, representing the (iii) operator’s preferences for split-AI; (0.2) are equipped with a set of split-AI rules 402a-m (e.g. minimize inference duration, minimize network traffic volume, etc.); and (0.3) are equipped with a mapping, determining which app-ID shall use which NN 20. Once these initial steps have been performed (once, initially to the system), the steps implemented by the control plane entity 110a are as follows. In a step 1, the control plane entity 110a of the mobile communication system 100 is receiving a request for split-AI operation (including a NN-ID to execute). In a step 2, the control plane entity 110a of the mobile communication system 100 acquires information relating to the underlying conditions (i). In a step 3, the control plane entity 110a of the mobile communication system 100 acquires information relating to the system capabilities (ii). In a step 4, the control plane entity 110a of the mobile communication system 100 applies the operator-specific policy 401 to determine which rule 402a-m to select. In a step 5, the control plane entity 110b of the mobile communication system 100 enforces and executes the selected rule (NN splitting, layer allocation, execution), as will be described in more detail below.
[0080] Figure 6 illustrates an architecture of the mobile network 100, including the control plane entities 110a, 110b or the combined control plane entity 110 according to an embodiment, which in figure 6 is referred to as the Split Control Entity (SCE). As already mentioned, the functionality of the SCE disclosed herein may be realized as one CP NF 110 or may be composed of a set of CP NF s 110a,b. How the Split Control Entity 110; 110a,b is realizing the split, depends on the policy 401, which is defined and provided to the SCE 110; 110a, b by the Network Management System (NMS) 120, residing in the MP of the mobile communication system 110. Once the SCE 110; 110a, b has decided about the split and the allocation, it communicates to the UP Entities and other instances (e.g. application server 190a connected via the DN 190, UE 140) which sub-NNs they should execute.
[0081] As already mentioned above, figure 3 shows an embodiment, wherein the CP entities 110a, b are implemented in a 6G mobile network 100. In this embodiment, the CP entities 110a, b are CP Network Functions (NFs) 110a, b. The two core entities 110a, b (the policy and the set of split-AI rules) may be realized by means of dedicated NFs in the CP. For instance, NF1 110a hosts the policy, which is used to derive the split-AI rule 402a-m to apply. Then, NF1 110a notifies NF2 110b, which is in charge of executing the selected split-AI rule 402a-m, i.e. finding the proper split and allocation of NN layers to the communicate and compute nodes of the 6G mobile network 100. By means of programming, the 6G MP, specifically the Network Management System (NMS) 120, provides or configures NF1 110a with the policy 401 to use to derive the proper split-AI rule 402a-m. The 6G MP may further provide or configured NF2 110b with all necessary information to properly execute the NN split. For instance, information on the UP topology, including the compute capabilities of the UP nodes. NF2 110b further acquires information from the 6G UP, which includes dynamic, real-time statistics, such as the current UL / DL load of UP nodes, the current compute load, split-AI related performance metrics, and the like. This information may be beneficial for NF2 110b when deciding about the proper split, UP path setup, and NN layer to compute node allocation. Furthermore, NF2 110b may configure the UP nodes for split-AI, i.e., it enforces split-AI in the UP, which then executes the NN split.
[0082] Figure 7 illustrates a variant of the embodiment shown in figure 3 for a 5G mobile network 100. In this embodiment, the 5G CP entities 110a, b in charge for hosting the policy 401 may be implemented by the Policy and Control Function (PCF) 110a, which is provided or configured by means of programming by the 5G MP 120 with the operator-defined policy 401 for split- AI. Once the PCF 110a has determined the split-AI rule 402a-m to apply for a given flow, it notifies the control plane entity 110b implemented by the Session Management Function (SMF) 110b, which is then executing the split-AI rule 402a-m. In other words, the SMF 110b defines where to split the NN 20 and how to allocate the resulting sub-NNs to the communicate and compute nodes of the 5G mobile network 100. The SMF 110b may be also in charge of notifying the UP entities about the sub-NNs to execute.
[0083] Figures 8a-c show signaling diagrams illustrating several processing steps and message exchanges for the 5G embodiment shown in figure 7 for setting up a CC flow with split-AI capability during three different stages, including an initial setup stage shown in figure 8a, a second constant monitoring stage shown in figure 8b, and a UE PDU session stage shown in figure 8c.
[0084] In step 0.1 of figure 8a, the Application Function (AF) 170 provides information to the 5G system (to the 5G MP 180), namely information relating to the NN 20. This information may include the NN architecture, i.e. the number layers 20a-e, the neurons 10 on each layer 20a-e, the activation function the neurons 10 execute, as well as the weights of input and the bias of the neuron 10.
[0085] In step 0.2 of figure 8a, the 5G MP 180 provides the NN information (obtained via the AF 170) to the SMF 110b (5G CP NF). In this way, the information provided by the application provider (via the AF 170), is available at the SMF 110b, which, as described above, may implemented the control plane entity 110b for splitting the provided NN 20.
[0086] In step 0.3 of figure 8a, the MP 180 prepares the UP entities 130 for split-AI operation execution. This includes programming the entities for supporting the execution of sub-NNs.
[0087] In step 0.4 of figure 8a, further information is provided from the MP 180 to the SMF implementing the control plane entity 110b, namely information about the system capabilities. On top of conventional UP topology information (location of UP entities, links between UP entities, UL / DL bandwidth, etc.), further information that is relevant for split-AI operations is provided. This may include the available compute capabilities of the UP entities, their support of executing specific activation functions, of information such as the energy efficiency of nodes.
[0088] In step 0.5 of figure 8a, the network operator via the MP entity 180 programs the PCF implementing the CP entity 110a with the split-AI policy 401, i.e. the MP 180 provides the split-AI policy 401 to the PCF 110a.
[0089] As will be appreciated, steps 0.1 to 0.5 of figure 8a are initial steps that may have to be executed only once for setting up the system so that afterwards only updates may become necessary.
[0090] In step 0.6 of figure 8b, the UP entities 130 provide dynamic monitoring information to the SMF 110b, i.e. information relating to the underlying conditions. This may include typical statistics, already provided as of today (current UL / DL traffic volume, delay, and packet loss rate), but may be extended towards metrics relating to computation. That is, the current compute load (e.g. percentage utilization of CPU / RAM / GPU), current energy usage / energy efficiency, or split-AI specific reporting (e.g. current duration for executing certain layers, statistics on the actual output size, etc.). As will be appreciated, step 0.6 may be performed constantly so that the UP statistics are permanently collected and sent to the SMF 110b.
[0091] In step 1 of figure 8c, the E 140 initiates a PDU session establishment for a CC Flow with split-AI capability. The PDU Session Establishment Request (PER) contains various information, also relating to the split-AI operations.
[0092] In step 2a of figure 8c, the SMF 110b receives the PER from the UE 140. In order to determine which split-AI rule 402a-m to apply for the requested CC Flow, it requests the split-AI rule to apply by forwarding relevant information to the PCF 110a.
[0093] With the information provided by the SMF 110b, the PCF 110a may use in step 3 of figure 8c its internal policy 401 (provided by the MP 180 in step 0.5 of figure 8a) to select the split-AI rule 402a-m to apply for the requested flow.
[0094] In step 2b of figure 8c, the PCF 110a notifies the SMF 110b about the split-AI rule 402a-m to be applied (as a response to step 2a of figure 8c).
[0095] In step 4 of figure 8c, the SMF 110b determines which NN 20 to use for the application ID in the PER. Then, given the information provided earlier, i.e. the (static) system capabilities provided in step 0.4, the (dynamic) underlying conditions provided in step 0.6, and the flow-related info provided through PER in step 1), the SMF 110b splits the NN 20 according to the split-AI rule 402a-m provided by the PCF 110a in step 2b.
[0096] In step 5 of figure 8c, the CC Flow is being established, that is a flow between the UE 140 and the DN 190 is set up, so to allow for split-AI operations.
[0097] Figure 9 is a flow diagram illustrating a method 900 for operating the control plane entity 110a or 110 according to an embodiment for supporting the distribution of the NN 20 across a plurality of communicate and compute nodes of the mobile network 100, including the UE 140, one or more access nodes 150, one or more user plane functions 130, and / or an application server 190a. The method 900 comprises a step 901 of selecting one of the plurality of selectable NN distribution rules 402a-m based on one or more of the following : the NN distribution rule selection policy 401; capability information about each of the plurality of communicate and compute nodes of the mobile network 100; and / or monitoring data indicative of one or more dynamic conditions of the mobile network 100.
[0098] The method 900 shown in figure 9 can be performed by the control plane entity 110a or 110 according to an embodiment. Thus, further features of the method 900 shown in figure 9 result directly from the functionality of the control plane entity 110a or 110 as well as the different embodiments thereof described above and below.
[0099] Figure 10 is a flow diagram illustrating a method 1000 for operating the control plane entity 110b or 110 according to an embodiment for distributing the NN 20 across a plurality of communicate and compute nodes of the mobile network 100, including the UE 140, one or more access nodes 150, one or more user plane functions 130, and / or an application server 190a. The method 1000 comprises a step 1001 of obtaining information about the selected NN distribution rule 402a-m, i.e. which of the plurality of selectable NN distribution rules has been selected. Moreover, the method 1000 comprises a step 1003 of computing, i.e. determining a NN distribution across the plurality of communicate and compute nodes of the mobile network 100 based on the selected NN distribution rule 402a-m. The method 1000 shown in figure 10 can be performed by the control plane entity 110b or 110 according to an embodiment. Thus, further features of the method 1000 shown in figure 10 result directly from the functionality of the control plane entity 110b or 110 as well as the different embodiments thereof described above and below.
[0100] The person skilled in the art will understand that the "blocks" ("units") of the various figures (method and apparatus) represent or describe functionalities of embodiments of the present disclosure (rather than necessarily individual "units" in hardware or software) and thus describe equally functions or features of apparatus embodiments as well as method embodiments (unit = step).
[0101] In the several embodiments provided in the present application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described embodiment of an apparatus is merely exemplary. For example, the unit division is merely a logical function division and may be another division in an actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented by using some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.
[0102] The units described as separate parts may or may not be physically separate, and parts displayed as units may or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.
[0103] In addition, functional units in the embodiments of the disclosure may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units may be integrated into one unit.
Claims
CLAIMS1. A control plane entity (110; 110a) for supporting the distribution of a neural network, NN, (20) across a plurality of communicate and compute nodes of a mobile network (100), including a user equipment, UE, (140), one or more access nodes (150), one or more user plane functions (130), and / or an application server (190a), wherein the control plane entity (110; 110a) is configured to select one of a plurality of NN distribution rules (402a-m) based on one or more of the following: a NN distribution rule selection policy (4019; capability information about each of the plurality of communicate and compute nodes; and / or monitoring data indicative of one or more dynamic conditions of the mobile network (100).
2. The control plane entity (110; 110a) of claim 1, wherein the control plane entity (110; 110a) is further configured to compute a NN distribution based on the selected NN distribution rule (402a-m) or to provide information about the selected NN distribution rule (402a-m) to a further control plane entity (110b) for computing a distribution based on the selected NN distribution rule (402a-m) by the further control plane entity (110b).
3. The control plane entity (110; 110a) of claim 2, wherein the control plane entity (110) or the further control plane entity (110b) is configured to compute a NN distribution based on the selected NN distribution rule (402a-m) by determining a NN distribution optimizing one or more key performance indicators defined by the selected NN distribution rule (402a-m).
4. The control plane entity (110; 110a) of claim 2 or 3, wherein the control plane entity (110) is further configured to provide the NN (20) across the plurality of communicate and compute nodes of the mobile network (100) according to the computed NN distribution or to provide information about the computed NN distribution to a further control plane entity (110b) for providing the NN (20) across the plurality of communicate and compute nodes of the mobile network (100) by the further control plane entity (110b) according to the computed NN distribution.
5. The control plane entity (110; 110a) of any one of the preceding claims, wherein the NN (20) comprises a plurality of NN processing layers (20a-e) and wherein the control plane entity (110) or the further control plane entity (110b) is configured to distribute the NN (20) across the plurality of communicate and compute nodes based on the computed NN distribution by allocating a plurality of subsets of the plurality of NN processing layers (20a-e) to the plurality of communicate and compute nodes of the mobile network (100).
6. The control plane entity (110; 110a) of any one of the preceding claims, wherein the control plane entity (110; 110a) is configured to receive the NN distribution rule selection policy (401) from a management plane entity (120) of the mobile network (100).
7. The control plane entity (110; 110a) of any one of the preceding claims, wherein the control plane entity (110; 110a) is configured to receive the monitoring data from a user plane function (130) of the mobile network (100).
8. The control plane entity (110; 110a) of any one of the preceding claims, wherein the plurality of NN distribution rules (402a-m) comprise one or more of the following: a NN distribution rule minimizing inference duration; a NN distribution rule minimizing network traffic volume; and / or a NN distribution rule minimizing UE energy consumption.
9. A method (900) of operating a control plane entity (110; 110a) for supporting the distribution of a neural network, NN, (20) across a plurality of communicate and compute nodes of a mobile network (100), including a user equipment, UE, (140), one or more access nodes (150), one or more user plane functions (130), and / or an application server (190a), whereinthe method (900) comprises selecting (901) one of a plurality of NN distribution rules (402a-m) based on one or more of the following : a NN distribution rule selection policy (401); capability information about each of the plurality of communicate and compute nodes; and / or monitoring data indicative of one or more dynamic conditions of the mobile network (100).
10. A control plane entity (110b) for distributing a neural network, NN, (20) across a plurality of communicate and compute nodes of a mobile network (100), including a user equipment, UE, (140), one or more access nodes (150), one or more user plane functions (130), and / or an application server (190a), wherein the control plane entity (110b) is configured to: obtain information about a selected NN distribution rule (402a-m); and compute a NN distribution based on the selected NN distribution rule (402a-m).
11. The control plane entity (110b) of claim 10, wherein the control plane entity (110b) is configured to compute the NN distribution based on the selected NN distribution rule (402a-m), capability information about each of the plurality of communicate and compute nodes of the mobile network (100) and / or monitoring data indicative of one or more dynamic conditions of the mobile network (100).
12. The control plane entity ( 110b) of claim 10 or 11 , wherein the control plane entity ( 110b) is configured to compute the NN distribution based on the selected NN distribution rule (402a-m) by determining a NN distribution optimizing one or more key performance indicators defined by the selected NN distribution rule (402a-m).
13. The control plane entity (110b) of any one of claims 10 to 12, wherein the control plane entity (110b) is further configured to provide, based on the computed NN distribution, the NN (20) across the plurality of communicate and compute nodes of the mobile network (100) or to provide information about the computed NN distribution to a further control plane entity for providing the NN (20) across the plurality of communicate and compute nodes of the mobile network (100) by the further control plane entity, based on the computed NN distribution.
14. The control plane entity (110b) of any one of claims 10 to 13, wherein the NN (20) comprises a plurality of NN processing layers (20a-e) and wherein the control plane entity (110b) is configured to distribute the NN (20) across the plurality of communicate and compute nodes of the mobile network (100) based on the computed NN distribution by allocating a plurality of subsets of the plurality of NN processing layers (20a-e) to the plurality of communicate and compute nodes of the mobile network (100).
15. The control plane entity (110b) of any one of claims 10 to 14, wherein the control plane entity (110b) is configured to obtain the information about the selected NN distribution rule (402a-m) by receiving the information about the selected NN distribution rule (402a-m) from a further control plane entity (110a).
16. A method (1000) for operating a control plane entity (110b) for distributing a neural network, NN, (20) across a plurality of communicate and compute nodes of a mobile network (100), including a user equipment, UE, (140), one or more access nodes (150), one or more user plane functions (130), and / or an application server (190a), wherein the method (1000) comprises: obtaining (1001) information about a selected NN distribution rule (402a-m); and computing (1003) a NN distribution based on the selected NN distribution rule (402a-m).
17. A computer program product comprising a computer-readable storage medium for storing program code which causes a computer or a processor to perform the method (900) of claim 9 or the method (1000) of claim 16, when the program code is executed by the computer or the processor.
Citation Information
Patent Citations
Providing distributed ai models in communication networks and related nodes / devices
US20230412513A1