Automatic architecture adaptation for distributed artificial intelligence
The automatic discovery and registration of DAI agents and controllers within a network repository function streamline DAI infrastructure management, addressing inefficiencies in dynamic network environments by enabling rapid adaptation and reduced complexity.
Patent Information
- Application Number
- PCT/EP2024/065572
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2025-12-11
AI Technical Summary
Existing distributed artificial intelligence (DAI) infrastructure management in networks is complex and lacks mechanisms for real-time adaptation to dynamic changes, leading to prolonged adaptation cycles and inefficiencies in managing DAI groups.
Implementing a network entity with enhanced functionalities for automatic discovery and registration of DAI agents and controllers, utilizing a network repository function to exchange profiles and manage DAI groups dynamically, reducing management complexity through proactive and reactive group joining.
Facilitates seamless and efficient adaptation of DAI infrastructure to dynamic changes, reducing management complexity and enabling faster response to network dynamics.
Smart Images

Figure EP2024065572_11122025_PF_FP_ABST
Abstract
Description
[0001] AUTOMATIC ARCHITECTURE ADAPTATION FOR DISTRIBUTED ARTIFICIAL INTELLIGENCE
[0002] FIELD OF THE INVENTION
[0003] This invention relates to distributed artificial intelligence, in particular in service-based architecture networks with dynamic architectures.
[0004] BACKGROUND
[0005] Distributed artificial intelligence (DAI) is used to refer to a framework of several artificial intelligence (Al) agents (or Al models), where each Al agent can perform operations such as autonomously monitoring system variables, training and inferring its models, taking an action on managed or controlled network entities and exchanging monitoring information with other distributed Al agents.
[0006] Some network operations, such as dynamic task migration (where tasks may have very stringent deadlines) among user plane functions (UPFs) may require a very fast network decision on the operation. In such scenarios, distributed Al can be beneficial due to its ability to take fast and local decisions.
[0007] In a distributed Al architecture, there can be multiple distributed Al (DAI) groups. Figure la shows a high-level architecture of a DAI group 100. Each Al agent 105, 106, 107 in the DAI group 100 is deployed as a part of an existing network entity (NE). In this example, each agent 105, 106, 107 is deployed as part of an existing network function instance 101, 102 and 103 respectively. Figure Ib shows anAI controller 108 forthe three Al agents l05, 106, 107 in the DAI group 100. In this example, the controller 108 is deployed as part of an existing NE 104. As shown in Figures la and lb, each DAI group 100 comprises two types of entities: two or more Al agents 105, 106, 107 and an Al controller 108.
[0008] The Al controller in each DAI group controls and manages the Al agents in its own group. The Al controller and each Al agent as a part of an existing network entity or network function instance are deployed at the user plane (UP), control plane (CP) or management plane (MP), or across different network planes, in a hybrid manner. An example of hybrid DAI group deployment is when the session management function (SMF), which belongs to the CP, is enhanced with Al controller functionality, while its UPFs, which belongs to the UP, are enhanced with Al agent functionality. Moreover, depending on the implementation, an Al controller could additionally be enhanced with the core functionalities of an Al agent, or a network entity could be enhanced both with the core functionalities of an Al agent and the core functionalities of an Al controller. The Al agents in a single DAI group perform the same Al task. Non-limiting examples of Al tasks include UP path reconfiguration and dynamic task migration.
[0009] The key features of a DAI agent are outlined below, and depicted in Figure 2, which is a high-level diagram of the key functions of each distributed Al agent 201, 202, and their respective functions 203, 204 showing exchange of monitoring information between the neighboring DAI agents 201 and 202 (in the same DAI group) at 205.
[0010] In this example, each Al agent is empowered with the following functions: 1 ) data collection and sharing function that enables the Al agent to collect data not only from its managed NE, but also share minimal but essential information with other Al agents in its DAI group to learn cooperatively, 2) intelligence function, that stores the current status of its Al model, e.g. a neural network for a deep learning-based approach or a table for a non-deep learning approach, 3) monitoring function, constantly running in the background, that monitors the important Al / machine leaning (ML) KPIs, e.g. average reward and loss functions, meaning that the Al agent raises a warning if the AI / ML KPIs deviate by more than a pre-defined threshold value. Finally, each Al agent also includes 4) an action function, that enables it to autonomously perform an action on a managed or controlled network entity based on the decision of the intelligence function. Each of these functions can be implemented in different ways, use case by use case.
[0011] The life-cycle phases of an Al agent on a specific Al task generally comprise the following: 1) training, 2) deployment of trained models and inference, and 3) reporting and maintenance.
[0012] As mentioned above, the Al controller is responsible for the control and management of Al agents in its group. The Al controller is also aware of the topology information (e.g. Al agent connectivity), Al agent capabilities (i.e. Al algorithm settings) and information relating to infrastructure changes. An example of an infrastructure change is when a new network function (NF) is activated or an existing NF is terminated, requiring updates to the connectivity of the DAI group. The Al controller can additionally monitor system key performance indicators (KPIs), for example the percentage of tasks violating their deadlines in task migration.
[0013] In a mobile network, the Al agent / AI model can be implemented as a sub functionality of a NE, which may be a network function (NF). For example, in of scenario of dynamic task migration among UPFs, the Al agent can be implemented at a UPF. An Al controller hosted at the SMF or as a separate NF can manage and control one or multiple DAI groups each composed of multiple distributed Al capable UPFs. Analytics generation by the Al agents can be used to dynamically migrate tasks among UPFs, when UPF task deadlines could be violated.
[0014] Figure 3 shows an example of a distributed Al architecture in the use case of dynamic task migration. In this exemplary network 300, there are two SMFs, SMF1 301 and SMF2 302. There are multiple UPFs 303-318. In this example, UPF 308 is under high load and it is desirable to migrate its task to another UPF. In this example, the optimal destination UPF is UPF 312.
[0015] From 5G, the mobile network utilizes network function visualization (NFV) technology and / or service-based architecture (SBA). The network functions (NF s) are dynamically deployed considering the changing load of the network traffic. Therefore, DAI infrastructure (i.e., available Al controller and Al agents) which are deployed as part of an NE, and the mapping between Al controller and Al agents, may change dynamically.
[0016] Three scenarios are listed below with reference to Figure4a, 4b and 4c where the DAI infrastructure can change dynamically:
[0017] 1. Network scaling: Deployment of new Al agent-capable NFs or Al controller;
[0018] 2. Energy saving: Some of the Al agents / AI controller can be switched off in certain time period;
[0019] 3. Load balancing: Al controller of an Al agent could change depends on the control / management load of the Al controller.
[0020] In the example show in Figure 4a, the network comprises SMF1 401 which acts as an Al controller, and UPFs 402-406. SMF1 401 controls two DAI groups 407 and 408. Groups 407 and 408 are associated with a first Al task (task 1). A new UPFx (Al agent) 406 is deployed and joins the existing Al groups for task 1.
[0021] In the example shown in Figure 4b, the network comprises two SMFs which acts as Al controllers, SMF1 451 and SMF2452, UPFs 453-456 and SMFs 457-460. Initially, SMF1 451 controls groups 461 and 462 and SMF2452 controls group 463. Groups 461 and 462 are associated with a first task (task 1 ) and group 463 is associated with a second task (task 2). When SMF2452 is turned off, control of Al task 2 is moved to SMF 1 451. In the example of Figure 4c, the network comprises two SMFs which acts as Al controllers, SMF1 471 and SMF2472, UPFs 473-476 and SMFs 477-480. Initially, SMF1 471 controls groups 481, 482 and group 483. Groups 481 and 482 are associated with a first task (task 1) and group 483 is associated with a second task (task 2). In this example, the control of Al task 2 is taken over by SMF2472 since SMF1 471 is overloaded.
[0022] Figure 5 show an example of a known management plane (MP) based approach of the DAI infrastructure update. In this approach, it is assumed that operation and maintenance (OAM) 501 maintains a mapping between all Al controllers 502 and Al agents 503. Since Al agent resources may be shared across multiple Al tasks (e.g., in different time scale, and / or considering CPU resources), an Al agent may be managed by multiple Al controllers for efficient resource usage.
[0023] The following steps can be foreseen:
[0024] 1. The OAM configures the Al controller in the network on the Al agent(s) to be managed by this Al controller. In case a new Al agent capable NF is added / removed (trigger 1) (e.g., for maintenance or for energy saving)
[0025] 2a. The OAM may reconfigure an Al controller (e.g., adjust the Al agent managed by an Al controller) and / or configure a new Al controller (e.g., since the configured Al controller has already reached the load limits). For the reconfiguration of an Al controller, the OAM needs to determine which and when to configure (so that the ongoing distributed Al tasks should not be affected)
[0026] 2b. The OAM (re)configurates the determined Al controller. In case the Al agent resource availability changes (trigger 2). 3a. The OAM needs to know which Al controller to notify the Al agent resource updates (i.e., avoid the signalling overhead to send the information to the irrelevant Al controllers)
[0027] 3b. The OAM notifies the determined Al controller on the Al agent resource updates.
[0028] Further, the 3GPP specification (TS 23.501 [2]) in Clause 6.2.6.2 defines the NF profile maintained in the network repository function (NRF), that captures NF instance ID, NF type, NF capacity information etc. However, this NF profile gives no indication of distributed Al capability. Network data analytics function (NWDAF) profile (TS 23.288 [4] ) defined from Rel. 18 only supports “FL server” or “FL client” types in a horizontal Federate Learning (FL) case. Furthermore, there is a lack of mechanism that maintains a real-time mapping of Al controllers and Al agents.
[0029] FL is also a distributed learning paradigm. FL involves training multiple AI / ML models contained in local nodes without exchanging data samples. In centralized FL, a central server serves to orchestrate several steps of the learning algorithms and coordinate all the participating nodes during the process, including aggregating the parameters of the distributed models into a single global model. In decentralized FL, the distributed nodes coordinate among themselves to obtain the global AI / ML model. Further, 3GPP has also introduced enhancements in the 5G core to support centralized FL among several NWDAF instances (3GPP TS 23 ,288[4]). However, there is no cooperation (i.e., monitoring data exchange) among the distributed nodes during the inference phase.
[0030] The OAM based procedure above has disadvantages including the following. It has high management complexity for DAI infrastructure updates. OAM needs to maintain complex mapping relationship between the all the Al controllers and Al agents. Each DAI infrastructure change (for example adding / removing an Al agent, or a change of the resource availability of Al agent) would require OAM to make network management decision (when, which to (re)configure / notify) and take related actions (configuration / notify the Al controller). There is also a long DAI groups adaptation cycle in case of DAI infrastructure change. Since the MP reaction time (i.e., notify the DAI infrastructure changes) is normally in the scale of a few seconds (compared to the ms time scale in the CP), the DAI groups adaptation in case of DAI infrastructure change would require at least a few seconds. Meanwhile, the MP based solution does not support run time group adaptation, since MP is not aware of the DAI groups neither the states of the DAI groups. This further increases the DAI group adaptation cycle.
[0031] It is desirable to develop an improved method for supporting distributed Al operations in the case of a dynamic DAI infrastructure.
[0032] SUMMARY OF THE INVENTION
[0033] According to a first aspect, there is provided a network entity for a service-based architecture network, the network entity being configured to implement an agent for performing distributed artificial intelligence for one or more tasks, each task being associated with one or more distributed artificial intelligence groups comprising two or more agents in the network, wherein the network entity is configured to send a registration request to a network repository function of the network, the registration request comprising a profile indicating a distributed artificial intelligence capability of the agent.
[0034] This can enable the automatic discovery of the appropriate agents by the Al controller and therefore may reduce the management complexity for distributed artificial intelligence infrastructure changes.
[0035] The profile may comprise one or more of: an agent identifier, a service area, an indication of a type of supported distributed artificial intelligence tasks, neighborhood information for neighbouring network entities and agent load or remaining capacity of the agent. The profile can be used for discovery and selection of an appropriate agent for performing an distributed artificial intelligence task. This may allow the network repository function to discovery an appropriate agent according to the discovery request. This may also allow the agent to be correctly discovered and selected by a controller and subsequently join a suitable distributed artificial intelligence group.
[0036] The agent may be configured to discover a controller by sending a discovery request to the network repository function, the controller being configured to manage one or more distributed artificial intelligence groups. The automated discovery controllers can allow an agent to identify by itself the appropriate controller and therefore may result in reduced management complexity to operator distributed artificial intelligence in the control plane.
[0037] The discovery request may comprise one or more of: a type of network entity hosting the controller, an area of interest, a specified controller capability and one or more supported distributed artificial intelligence task types or analytics identities. This may allow the agent to discover the appropriate controller according to its distributed artificial intelligence capabilities for intended interactions and / or to join a suitable distributed artificial intelligence group.
[0038] The agent may be configured to receive a discovery response from the network repository function, the discovery response comprising one or more of the following: one or more endpoint address(es) of one or more discovered controller(s) or discovered instance(s) hosting the controllers) and a respective profile for the or each discovered controller. This may allow the agent to interact with a discovered controller and join a suitable distributed artificial intelligence group.
[0039] The agent may be configured to select the controller from multiple discovered controllers for further interaction, where the selection is based on respective profiles for each of the multiple discovered controllers received in the discovery response. This may allow the agent to select an appropriate controller according to its distributed artificial intelligence capabilities for followup interactions. The further interaction may comprise one or more of the following: notifying the controller that a new agent is deployed, requesting distributed artificial intelligence group information and requesting to join a distributed artificial intelligence group. This may allow the network to adapt to dynamic infrastructure changes.
[0040] The agent may be implemented as a sub- functionality of the network entity. The agent may be implemented as a subfunctionality of a network function. This may allow the agent to scale together with the deployed network entity and to be implemented as part of a service-based architecture network.
[0041] According to another aspect, there is provided a network entity for a service-based architecture network, the network entity being configured to implement a controller for managing distributed artificial intelligence for one or more tasks, each task being associated with one or more distributed artificial intelligence groups comprising two or more agents in the network, wherein the network entity is configured to send a registration request to a network repository function of the network, the registration request comprising a profile indicating a distributed artificial intelligence controller functionality.
[0042] This can enable the automatic discovery of the appropriate controller by the Al agents and therefore may reduce the management complexity for distributed artificial intelligence infrastructure changes.
[0043] The controller may be responsible for the control and management of agents in its distributed artificial intelligence group. The controller may configure, reconfigure or update the computation resources and functionalities of agents regarding a specific Al task and trigger the different phases for the life-cycle management of the Al agents on a specific Al task.
[0044] The profile may comprise one or more of: a controller identifier, a service area, a controller capability and one or more distributed artificial intelligence groups managed by the controller. This may allow the network repository function to discover an appropriate controller according to the discovery request. This may also allow the agent to discover and select an appropriate controller and subsequently join a suitable distributed artificial intelligence group.
[0045] The controller may be configured to control and manage membership of one or more distributed artificial intelligence groups throughout different phases of a life-cycle of an agent on the associated distributed artificial intelligence task. The life-cycle phases of an Al agent on a specific Al task may comprise the following: 1) training, 2) deployment of trained models and inference, and 3) reporting and maintenance. This enables an agent to join a distributed artificial intelligence group in the right time and may reduce the distributed artificial intelligence group adaptation cycle in case of distributed artificial intelligence infrastructure change.
[0046] The controller may be configured to discover one or more agents in the network by sending a discovery request to the network repository function, wherein the discovery request comprises one or more of: a type of network entity hosting the agent, a service area, a distributed artificial intelligence capability of the agent, an indication of a type of supported distributed artificial intelligence tasks, agent load and remaining agent capacity. The automated discovery of appropriate agents by the controller itself may result in reduced management complexity to operator distributed artificial intelligence in the control plane.
[0047] The controller may be configured to receive a discovery request response from the network repository function, the discovery request response comprising one or more of: endpoint address(es) of one or more discovered network entities hosting respective agents and respective profile(s) of the discovered agent(s). This may allow the controller to interact with discovered agents suitable for a distributed artificial intelligence task. The controller may be configured to receive a discovery request response from the network repository function comprising a respective profile for each of multiple discovered agents and select one or more agents of the multiple discovered agents to join a distributed artificial intelligence group in dependence on the respective profile for each of the multiple discovered agents. This may allow the controller to select agents with appropriate distributed artificial intelligence capabilities.
[0048] The controller may be configured according to one of the following: (i) the controller is a dedicated controller for a distributed artificial intelligence group; (ii) the controller is a shared controller across distributed artificial intelligence groups of the same network entity type; and (iii) the controller is a shared controller across distributed artificial intelligence groups of different network entity types. This may allow for different deployment options in different network scenarios. Other scenarios are also possible (for example, different network architecture, different network infrastructure and / or different network traffic pattern).
[0049] The controller may be configured to maintain distributed artificial intelligence group information for one or more distributed artificial intelligence groups that it is configured to control. This may allow for flexible scaling of the controller independent of the agents.
[0050] The controller may be configured to include one or more discovered agents in a distributed artificial intelligence group based on the respective profile(s) of the discovered agent(s), agent resource availability and the maintained distributed artificial intelligence group information. This may allow for automatic and seamless distributed artificial intelligence group updates in the case of dynamic distributed artificial intelligence infrastructure.
[0051] The network entity may be a network function. The network entity may be a service-based entity. The network entity may be a hardware-based network apparatus or software-based.
[0052] The service-based architecture may be a 3GPP 5G network. The agent(s) may be deployed as sub- functionality of one or more of a user plane function, a control plane function, an access management function, a session management function, a Service Communication Proxy, a policy control function, etc. These examples are not exhaustive. The agent may be deployed as a sub-functionality of any 5G network function. This may allow the approach to be used in telecommunications applications. The approach described herein may also be used for other service-based architectures with similar advantages.
[0053] According to a further aspect, there is a provided a network repository function configured to receive a registration request comprising a profile indicating a distributed artificial intelligence capability of the agent and / or a profile indicating a distributed artificial intelligence controller functionality. The network repository function may be further configured to store one or more of the profiles. The network repository function may be configured to send the profile indicating a distributed artificial intelligence capability of the agent to one or more controllers in the network. The network repository function may be configured to send the profile indicating a distributed artificial intelligence controller functionality to one or more agents in the network. The profile may be sent in response to a discovery request from the controller or agent. This may allow for automatic discovery of distributed artificial intelligence infrastructure changes via the network repository function (i.e. via the control plane).
[0054] In some examples, a network entity may host multiple agents. For example, multiple agents implemented by a network entity may perform multiple tasks in parallel, or one agent could perform a different task in a different stime slot to another agent.
[0055] An agent may also be referred to as an artificial intelligence model. An agent could be configured by a controller with one or more Al models. Each model may be used for a specific AI / ML task. Non-limiting examples of Al tasks include UP path reconfiguration and dynamic task migration.
[0056] According to another aspect, there is provided a method for implementation at a network entity for a service-based architecture network, the network entity being configured to implement an agent for performing distributed artificial intelligence for one or more tasks, each task being associated with one or more distributed artificial intelligence groups comprising two or more agents in the network, the method comprising sending a registration request to a network repository function of the network, the registration request comprising a profile indicating a distributed artificial intelligence capability of the agent.
[0057] According to a further aspect, there is provided a method for implementation at a network entity for a service-based architecture network, the network entity being configured to implement a controller for managing distributed artificial intelligence for one or more tasks, each task being associated with one or more distributed artificial intelligence groups comprising two or more agents in the network, the method comprising sending a registration request to a network repository function of the network, the registration request comprising a profile indicating a distributed artificial intelligence controller functionality.
[0058] These methods may reduce the management complexity for distributed artificial intelligence infrastructure changes. The methods may be computer-implemented methods.
[0059] According to a further aspect, there is provided one or more computer programs for instructing a computer comprising one or more processors to implement the method above.
[0060] According to a further aspect there is provided a data carrier storing in non-transitory form the one or more computer programs above.
[0061] BRIEF DESCRIPTION OF THE FIGURES
[0062] Figure la shows a high-level architecture of a DAI group.
[0063] Figure lb shows an Al controller for the three Al agents in a DAI group.
[0064] Figure 2 is a high-level diagram of the key functions of each distributed Al agent and their respective functions, showing exchange of monitoring information between the neighboring DAI agents in the same DAI group.
[0065] Figure 3 schematically illustrates an example of a DAI architecture in the use case of dynamic task migration.
[0066] Figures 4a-4c show examples of some different scenarios with dynamic DAI infrastructures.
[0067] Figure 5 shows a communication flow for a known management plane-based approach for a DAI infrastructure update.
[0068] Figure 6 shows an overview of a procedure for registration, discovery, selection and joining a DAI group.
[0069] Figure 7 schematically illustrates an exemplary deployment of a DAI infrastructure in 3GPP.
[0070] Figure 8 schematically illustrates DAI groups and DAI tasks and Analytics ID in 3GPP.
[0071] Figure 9 shows an exemplary communication flow for a registration procedure in a 3GPP 5G system. Figure 10 shows an exemplary communication flow for a discovery procedure in a 3GPP 5G system.
[0072] Figure 11 shows an exemplary communication flow for a DAI group information inquiry procedure in a 3GPP 5G system.
[0073] Figure 12a shows an exemplary communication flow for reactive DAI group joining in a 3GPP 5G system.
[0074] Figure 12b shows an exemplary communication flow for proactive DAI group joining in a 3GPP 5G system.
[0075] Figure 13a schematically illustrates an arrangement where there is a dedicated Al controller per DAI group.
[0076] Figure 13b schematically illustrates an arrangement where there is a shared Al controller across DAI groups of the same network function type.
[0077] Figures 14a and 14b show an example of a network entity configured to implement an agent and a network entity configured to implement a controller respectively.
[0078] DETAILED DESCRIPTION
[0079] Embodiments of the present invention may conveniently support distributed Al operation in the case of dynamic DAI infrastructure, allowing the network to be aware of the DAI infrastructure changes and how to adapt a distributed Al architecture (for example, DAI groups) to the DAI infrastructure changes.
[0080] Additional functionalities are defined for the Al controller or Al controller-hosting NEs and Al agent-hosting NEs (each agent comprising one or more Al models). The signaling messages between the NEs and the NRF (i.e., registration / registration update request, discovery request / response) include new information elements including a profile. The NE profile stored in the NRF includes additional information on the Al agent and / or Al controller.
[0081] The method generally follows two a step approach: automatic Al controller and / or Al agent discovery and seamless DAI group update, with enhanced functionalities at the NEs and the CP.
[0082] Figure 6 shows a general overview of the process. The procedure at the CP comprises two steps, with four phases, A)-D).
[0083] In step 1 , the following phases are performed:
[0084] Phase A) 601 : Registration of Al controller and Al agent in the NRF ;
[0085] Phase B) 602: Discovery and selection of Al agent by Al Controller;
[0086] Phase C) 603: Discovery and selection of Al controller by Al agents.
[0087] In this example, phases B) and C) are performed in parallel. In other implementations, they may not be performed in parallel, as appropriate.
[0088] In step 2, the following phase is performed:
[0089] Phase D) 604: Joining a DAI Group by Al agents, which may be performed proactively or reactively, as will be described in further detail later. In the approach described herein, the DAI infrastructure components have enhanced functionalities. At the Al controller or hosting NE of the Al controller, enhanced functionalities include registration of Al controller functionality / services, Al agent discovery and selection, and DAI group management and control. At the NE hosting the Al agent, enhanced functionalities include registration of Al agent sub-functionality / services, Al controller discovery and selection and DAI group discovery and selection. At the NRF, there is enhanced service discovery for Al agents and Al controllers.
[0090] These functionalities and the operation of the abovementioned phases will now be described in more detail.
[0091] Figure 7 schematically illustrates an exemplary deployment of DAI infrastructure in a 3GPP 5G network 700. The network comprises a plurality of network entities which are each part of the DAI infrastructure. The NEs may be NFs, which may be software-based, such as those described below. The NEs may alternatively be network apparatus (hardware-based). The NEs may be NE instances, including NF instances.
[0092] Network 700 comprises a network slice selection function (NSSF) 701, network exposure function (NEF) 702, NRF 703, PCF 704, unified data management (UDM) 705, application function (AF) 706, edge application server discovery function (EASDF) 707, network slice-specific authentication and authorization function (NSSAAF) 708 and authentication server function (AUSF) 709. The network comprises an AMF 710 with AMF instances 710a-d which each implement an Al agent. The network comprises a SMF 711 with SMF instances 71 la-d which each implement an Al agent. The network also comprises a service communication proxy (SCP) 712, network slice admission control function (NSACF) 713, Al controller 714, a user equipment (UE) device 715 , an access network (AN) 716 (optionally a radio access network (RAN)), UPF 717 and data network (DN) 718.
[0093] The controller can be a dedicated network function or a subfunction of a normal communication network function (for example, part of a SMF). In this example, Al controller 714 is deployed as a dedicated NE connected to 5GC using a service-based interface (SBI). Al agents are deployed as a sub-functionality of the AMF / SMF / PCF. This Al controller 714 may control multiple DAI groups (for example, one or more AMF groups, and / or SMF groups, and / or PCF groups. The DAI group may be assigned by the controller based on the DAI tasks.
[0094] In the examples described herein, Al controller functionalities may include one or more of Al agent discovery, DAI group management and Al agent configuration and control (functionality, neighbourhood, etc.) throughout the AI / ML life-cycle. More specifically, the Al controller can configure, reconfigure or update the computation resources and functionalities of Al agents regarding a specific Al task and / or trigger the different phases for the life-cycle management of the Al agents on a specific Al task.
[0095] In this example, the AMF instances 710a-d and the SMF instances 71 la-d each implement an Al agent. The AMF instances 710a-d form a first DAI group 719 and the SMF instances 71 la-d form a second DAI group 720. The PCF 704 also implements an agent.
[0096] In this example, there is a shared Al controller 714 across DAI groups of different NE types.
[0097] The approach may also be used with other deployments of DAI infrastructure in 3GPP 5G, such as that shown in Figure 3. In that example, an Al controller is deployed as a sub- functionality of SMF and Al agents are deployed as sub- functionality of UPFs. Similar infrastructure also applies to 3GPP 5G management plane architecture or future SBA-based mobile network (for example, 6G) architectures. Figure 8 schematically illustrates DAI groups and DAI tasks and analytics ID in a 3GPP network 800. Analytics ID is used in 3GPP (see 3GPP specification TS 23.288) to refer to a type of data analytics task with certain inputs and expected analytics output.
[0098] An Al controller may control and manage multiple DAI tasks. Each task is performed by one or more DAI groups. The NEs in a DAI group perform the same task in a distributed manner. In Figure 8, the Al controller 801 (here an SMF) controls and manages two DAI tasks: analytics ID 1 (task 1 ) and analytics ID 2 (task 2). For example, task 1 may be a UPF processing task migration and task 2 may be a packet data unit (PDU) session processing load balancing. The network comprises multiple NEs: UPFs 802-805 and SMFs 806-809.
[0099] In this example, the Al agents in a single DAI group perform the same Al task. For each DAI task, the Al controller may decide to split the involved Al agents for the DAI task into multiple DAI groups comprising two or more NEs (for example, group 1 810 and group 2 811 for the task analytics ID 1 for the network of Figure 8) considering the efficiency of the DAI algorithm. Task analytics ID 2 is performed by the agents in group 3 812 (i.e. by SMFs 806-809).
[0100] As mentioned above, phase A) comprises the registration of one or more Al controllers and Al agents in the NRF. The registration phase will now be described in more detail.
[0101] The NE (for example, NF) service registration and registration update processes are enhanced to include an additional profile to support the registration of the Al controller and Al agent. The communication flow of Figure 9 shows an exemplary registration procedure in 3GPP 5G network. Similar steps also apply to registration updates, as well as the registration in other systems.
[0102] To register or update the registration of an Al agent in the NRF, the additional profile of the NE hosting the Al agent comprises a profile indicating a distributed artificial intelligence capability of the agent.
[0103] The profile of the NE hosting the agent may indicate one or more of the following: Al agent ID (which may be unique within a hosting NE instance), service area (for example, represented as a list of tracking area ID, which can be the same as the hosting NE) and Al agent capability (DAI capability, type of supported distributed Al tasks / analytics ID). DAI capability can be true or false, true means it supports required functionalities for distributed Al (e.g. as specified in clause 2.1 of the 3GPP specification (TS 23.501 [2]). The profile may also comprise neighborhood information (for example, a list of address of neighbor NF instances with Al agent capability), agent load or remaining capacity (for example, a list of available resources, e.g., computation resources), available time period (for example, 2-4 every day, weekend / workday, daytime / night) and responsible Al controllers) (for example, a list of [Al controller ID, (time period), (analytics ID)]).
[0104] The Al controller can also be registered or updated in the NRF. The network entity hosting the controller is configured to send a registration request to the NRF comprising a profile indicating a distributed artificial intelligence controller functionality.
[0105] The profile of the NE hosting the Al controller may indicate one or more of the following: Al controller ID (which may be unique in a network / network slice / a plane), service area (for example, represented as a list of tracking area ID, can be the same as the hosting NF), Al controller capability (distributed Al controller functionality, for example, as defined in clause 2.1 of the 3GPP specification, supported distributed Al tasks / analytics ID) and managed distributed Al tasks and / or groups (for example, a list of [distributed Al task ID and / or Al group ID, analytics ID, type of Al agents hosting NF]). In some implementations, the NE hosting the Al controller can also host an Al agent at the same time. The Al agent can also optionally be registered as distributed Al service of the hosting NE in the NRF. In some implementations, the Al controller can be registered as an Al controller service of the hosting NE or a dedicated NE in the NRF. When the agent is deployed as a service in a NE, it can be registered as a service of the hosting NE. In this case the NE profile can be the service profile of the hosting NE instead of the NE profile.
[0106] As shown in Figure 9, a network entity (in this case a NF service consumer 901) hosting the agent and / or controller can send a registration request to the NRF. In this example, as shown at 903, the registration request is in the form of a Nnrf_NFManagement_NFRegister_request message. The profile for the agent or controller is stored at the NRF 902, as shown at 904. The NRF then sends a registration response to the agent or the controller. In this example, as shown at 905, the registration request is in the form of a Nnrf_NFManagement_NFRegister_response message.
[0107] Phase B), comprising the discovery and selection of one or more Al agents by the Al controller, will now be described in more detail.
[0108] Figure 10 illustrates an exemplary discovery procedure. A network entity (in this case a NF service consumer 901) hosting the agent or controller can send a discovery request to the NRF 902. In this example, as shown at 1003, the registration request is in the form of an Nnrf_NFDiscovery_Request message. The NRF then authorizes NF service discovery, as shown at 1004. The NRF then sends a discovery response to the agent or the controller. In this example, as shown at 1005, the registration request is in the form of an Nnrf_NFRequest_Response message.
[0109] The Al controller can discover one or more Al agents via the NRF. In this case, the discovery request may comprise one or more of the following: type of the hosting NE (e.g., SMF, UPF, etc.), area of Interest (e.g., represented as a list of tracking area ID), distributed Al capability, supported Al task types or analytics IDs and remaining agent capacity / agent load.
[0110] The discovery response may comprise one or more of the following: endpoint address(es) of the discovered NE instances hosting the Al agents and Al agent profile / AI agent service profile.
[0111] The Al controller may further select the Al agents from multiple discovered hosting NE instances considering the NE profile(s) in the discovery response. For instance, the Al controller may not include the supported Al task types / analytics IDs in the discovery request. Then the NRF may return multiple Al agents with the hosting NF profile indicating the supported analytics ID. The Al controller can select the Al agents to join a certain DAI group based on the supported analytics ID of that Al agent.
[0112] Phase C) comprises the discovery and selection of one or more Al controllers by an Al agent.
[0113] The Al agent can discover one or more Al controllers via the NRF. The discovery request may comprise one or more of the following: type of hosting NF or distributed Al controller NF type, area of Interest, distributed Al controller capability and supported Al task types or analytics IDs.
[0114] The discovery response may comprise one or more of the following: endpoint address(es) of the discovered Al controller NE instance(s) / discovered NE instance(s) hosting the Al controllers) and Al controller profile / AI controller service profile.
[0115] Al agents may further select the Al controller(s) from multiple discovered Al controller NF instance(s) / multiple discovery NF instance(s)for further interactions considering the NF profile(s) in the discovery response, in the same way as the Al controller case. The further interaction between the agent and the controller may comprises one or more of the following: notifying the Al controller that a new Al agent is deployed now, sending a request for DAI group information (available group, group active / idle, phase in Al lifecycle if active) and asking to join a DAI group
[0116] Figure 11 illustrates DAI Group information management at the Al controller, with an exemplary procedure for group information inquiry.
[0117] The Al controller 1100 maintains the DAI group information (i.e., DAI group membership and Al states information), and responses to the DAI group information inquiry from the NF service consumer 901. The NF service consumer 901 sends a DAI group information request to the controller 1100, as shown at 1101. At the controller, discovery of the related DAI group is then performed, as shown at 1102. The controller 1100 then sends a DAI group information response to the NF service consumer 901, as shown at 1103.
[0118] DAI group information may comprise, per managed DAI group, one or more of the following:
[0119] • Group ID
[0120] • DAI Task ID
[0121] • DAI Task type / Analytics ID
[0122] • Group states (i.e., active / idle, phase in Al lifecycle if active)
[0123] • List of member Al agents, Al agent ID, Al agent states (i.e., active / idle)
[0124] In phase D), the one or more Al agents join a DAI group. Two exemplary options for the group joining will now be described in more detail with reference to Figures 12a and 12b.
[0125] Figure 12a illustrates an exemplary procedure for reactive DAI group joining. An Al agent 1200 is asked to join a DAI group by an Al controller 1101. The following steps can be performed:
[0126] 0: the agent 1200 and the controller 1101 register with the NRF 902, using signalling as shown at 1201 and 1202 respectively. 1 : The Al controller obtains the profile of NF hosting Al agent via NRF discovery, as shown by the signal at 1203.
[0127] 2: The Al controller inquires the discovered Al agent on resource availability (e.g. available timeslot), as shown by the signal at 1204. This step is optional.
[0128] 3: The Al controller decides to include Al agent in a new or existing DAI group, based on the discovered Al agent profile, Al agent resource availability, and local maintained DAI group information, as shown at 1205.
[0129] 4: The Al controller adds the Al agent in the new or existing DAI group considering the states of the DAI group, as shown at 1206.
[0130] Step 4 (DAI group joining) may comprise the following sub steps: i. The Al controller determines to trigger the DAI group joining process for selected Al agent into a selected DAI group, considering the states of the DAI group (active, idle and lifecycle). ii. The Al controller configures the DAI group to include the selected Al agent iii. The Al controller configures the Al agent states (e.g. idle, active)
[0131] Figure 12b illustrates an exemplary procedure for proactive DAI group joining.
[0132] In this implementation, the Al agent 1200 wants to proactively join a DAI group. The following steps can be performed: 1. The Al agent obtains the profile of Al controller via NRF discovery (not shown);
[0133] 2. The Al agent subscribes to / requests DAI group information from the discovered Al controller. The step may comprise three sub steps. In step 2a, a DAI group information request / subscription is performed, as shown at 1251. In step 2b, the related DAI group is discovered at the controller 1101, as shown at 1252. In step 2c, the controller 110 sends a DAI group information response to the agent 1200, as shown at 1253.
[0134] 3. The Al agent decides whether to join a DAI group based on received DAI group information, as shown at 1254.
[0135] 4. The Al agent indicates its willingness to join a specific DAI group to the Al controller with resource availability (e.g., time slot) by sending a joining request to the controller 1101, as shown at 1255.
[0136] 5. The Al controller proceeds to perform Steps 3 and 4 of the reactive joining option. These steps are performed at 1256 in Figure 12b.
[0137] 6. The Al controller responds to the request with the DAI group ID if successful, as shown at 1257.
[0138] Some further exemplary embodiments of the approach will now be described.
[0139] There are several possible ways in which neighbourhood Al agents are discovered by each Al agent-hosting NE.
[0140] In a first option, the OAM configures the NE hosting Al agents. An exemplary neighbor Al agents list may be formed as: <NF instance ID, NF type, FQDN or IP address of NF, Endpoint Address of instance of Al agent service / sub- functionality^
[0141] This option has an advantage of simple implementation and low signaling overhead.
[0142] In a second option, discovery of neighborhood Al agents may be performed via IP Neighbor Discovery Protocols (e.g., Neighbor Discovery Protocol (NDP), Address Resolution Protocol (ARP), Internet Control Message Protocol (ICMP) Router Discovery and Router Redirect protocols) during bootstrap and with NRF assistance.
[0143] An agent may learn the neighbor IP address via NDP. It may then discover the IP address of NEs hosting Al agents via NRF in the same area of interest (Aol) (e.g., indicated as TAI(s)). The agent may then compare neighbor IP address with the IP addresses of discovered NFs in the NRF response. A mapped IP address indicates an identified neighbor Al agent.
[0144] This option has an advantage oflow management complexity and good scalability, since neighbor agent discovery is automatic.
[0145] A further implementation option of Al controller / AI agent discovery will now be described.
[0146] The discovery request can also be implemented as a subscription to one or more of the Al-related NEs in the network. The discovery request is used to discover existing Al controllers and / or Al agents in the network. In such implementations, a subscription can be used to allow agents or controllers to be notified on the addi tion / removal / change of Al controllers / AI agents in the network.
[0147] For example, the Al controller may subscribe to the NRF. The subscription may include the same contents as the discovery request. The notification in response to the subscription may include the same contents as the response to the discovery request. The discovery request can be implemented in combination with the subscription. For example, the Al controller may discover the Al agents in a certain area at first using a discovery request, then subscribe to the changes of the Al agents in this area. A newly deployed Al agent may discover the Al controller using a discovery request, while an Al controller may discover a new deployed Al agent by subscription or notification. The Al controller has several different deployment options. Figures 13a- 13c schematically illustrate some non- limiting examples.
[0148] As shown in Figure 13a, there may be a dedicated Al controller 1301 per DAI group.
[0149] In this example, a DAI group comprises SMF instances 1302-1305 which each host a respective Al agent. One disadvantage of this deployment option is that there may be a large amount of Al controller instances in the network. The proposed solution in this disclosure can be used for the Al controller to automatically discover a new Al agent or a replacement Al agent in the Int DAI group.
[0150] As shown in Figure 13b, Al controller 1301 can be shared across DAI groups of the same NF type. In addition to controlling SMF instances 1302-1305, in this example controller 1301 also controls SMF instances 1306-1308. This deployment option can be used in case of the Al agents with same type of hosting NF are distributed in a large area, or in case the number of Al agents is high. The Al controller can decide to split the Al agents into multiple DAI groups to improve the efficiency / performance of DAI. The proposed solution in this disclosure can be used both for automatic Al controller discovery and automatic Al agent discovery.
[0151] As shown previously in Figure 7, there may be a shared Al controller across DAI groups of different NF types. This deployment option may be advantageous in the case of a dedicated Al controller NF deployed in a certain area, and the Al controller controls and manages different types of Al tasks in its serving area. The proposed solution in this disclosure can be used both for automatic Al controller discovery and automatic Al agent discovery.
[0152] Figures 14a and 14b show examples of a network entity 1400 configured to implement an agent and a network entity 1500 configured to implement a controller respectively. Entities 1401, 1501 are computing entities. Entities 1402, 1502 are command and control entities. These entities are logical entities. In practice they may each be provided by one or more physical devices such as servers and data stores, and the functions of two or more of the entities may be provided by a single physical device. In some implementations, the entities may be cloud-based. Each physical device implementing an entity may comprise a processor and a memory. The devices may also comprise a transceiver for transmitting and receiving data. The memory stores in a non-transient way code that is executable by the processor to implement the respective entity in the manner described herein.
[0153] Therefore, the methods may be deployed in multiple ways, for example in the cloud or alternatively in dedicated hardware.
[0154] Some advantageous effects of embodiments of the present invention will now be described.
[0155] Compared to management plane centralized decisions and mapping maintenance, distributed decisions and mapping maintenance are implemented at each controller per Al / selected Al agents. Rather than requiring additional OAM management efforts, the approach described herein allows for automatic discovery of DAI infrastructure (e.g. agents and controllers) changes via the NRF.
[0156] The OAM procedure of the prior art suffers from high management complexity and long adaptation cycles of DAI groups in case of DAI infrastructure changes. The approach described herein can advantageously reduce the management complexity for DAI infrastructure changes and reduce the DAI group adaptation cycle in case of DAI infrastructure change.
[0157] The closed-loop two step control plane approach (i.e. 1. discovery of DAI infrastructure changes, and 2. DAI group management and control) for DAI group adaptation can result in fast DAI group adaptation to DAI infrastructure changes. For example, compared to the scale of seconds in management plane approaches, the control plane approach described herein may perform this in the scale of milliseconds. This can also enable runtime adaptation of distributed Al operations to dynamic DAI infrastructure (with consideration of DAI group lifecycle).
[0158] The automated discovery and selection of Al agents and Al controllers may also result in reduced management complexity to operator-distributed Al in the control plane (which may be an enabler for utilizing distributed Al in the control plane).
[0159] The mechanisms for passive and active DAI group joining described herein can allow for seamless DAI group update considering deployments updates of Al agents and controllers and also the states of the distributed Al group. This may close the loop of Al operation and Al management (enabler for network automation) and enable the fast adaptation of distributed Al operation under dynamic DAI infrastructure.
[0160] The two-step approach to decouple the Al group / states management from the DAI infrastructure changes can also result in reduced control signaling for DAI group update (group / states update within DAI group and among relevant Al agents only)
[0161] These enhancements are based on existing NRF framework for the automated discovery and selection of Al agents and Al controller, which provides direct compatibility with the 3GPP service-based architecture, and inherits the flexible NF deployment, auto-scaling, and good scaling feature (only relevant NF s) of service-based architectures.
[0162] There is no central or global Al model in the present approach. All Al models are trained and inferred in a distributed manner. The Al controller is tasked with coordinating the different phases of the life-cycle of Al agents on a specific Al task in its DAI group. This includes agents' dynamic deployment. The distributed Al agents can also be authorized to exchange only essential data with others in their DAI group, and not model parameters.
[0163] The main examples described herein are exemplified in a 3GPP 5G network use case. The solution can alternatively be implemented at NEs in any other network type using a service-based architecture.
[0164] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present invention may consist of any such individual feature or combination of features. In view of the foregoing description it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention.
Claims
CLAIMS1. A network entity (704, 710, 710a-d, 711, 711a-d, 802-809, 901, 1200, 1302-1308, 1400) for a service-based architecture network, the network entity being configured to implement an agent for performing distributed artificial intelligence for one or more tasks, each task being associated with one or more distributed artificial intelligence groups (719, 720, 810, 811, 812) comprising two or more agents in the network, wherein the network entity is configured to send a registration request (903) to a network repository function (703, 902) of the network, the registration request comprising a profile indicating a distributed artificial intelligence capability of the agent.
2. The network entity as claimed in claim 1 , wherein the profile comprises one or more of: an agent identifier, a service area, an indication of a type of supported distributed artificial intelligence tasks, neighborhood information for neighbouring network entities and agent load or remaining capacity of the agent.
3. The network entity as claimed in claim 1 or claim 2, wherein the agent is configured to discover a controller (714, 801, 1100, 1301, 1500) by sending a discovery request (1003) to the network repository function, the controller being configured to manage one or more distributed artificial intelligence groups.
4. The network entity as claimed in claim 3, wherein the discovery request comprises one or more of: a type of network entity hosting the controller, an area of interest, a specified controller capability and one or more supported distributed artificial intelligence task types or analytics identities.
5. The network entity as claimed in claim 3 or claim 4, wherein the agent is configured to receive a discovery response (1005) from the network repository function, the discovery response comprising one or more of the following: one or more endpoint address(es) of one or more discovered controller(s) or discovered instance(s) hosting the controller(s) and a respective profile for the or each discovered controller.
6. The network entity as claimed in claim 5, wherein the agent is configured to select the controller from multiple discovered controllers for further interaction, where the selection is based on respective profiles for each of the multiple discovered controllers received in the discovery response.
7. The network entity as claimed in claim 6, wherein the further interaction comprises one or more of the following: notifying the controller that a new agent is deployed, requesting distributed artificial intelligence group information and requesting to join a distributed artificial intelligence group.
8. The network entity as claimed in any preceding claim, wherein the agent is implemented as a sub-functionality of the network entity.
9. A network entity (714, 801, 901, 1100, 1301, 1500) for a service-based architecture network, the network entity being configured to implement a controller for managing distributed artificial intelligence for one or more tasks, each task being associated with one or more distributed artificial intelligence groups (719, 720, 810, 811, 812) comprising two or more agents in the network, wherein the network entity is configured to send a registration request (903) to a network repository function (703, 902) of the network, the registration request comprising a profile indicating a distributed artificial intelligence controller functionality.
10. The network entity as claimed in claim 9, wherein the profile comprises one or more of: a controller identifier, a service area, a controller capability and one or more distributed artificial intelligence groups managed by the controller.
11. The network entity as claimed in claim 9 or claim 10, wherein the controller is configured to control and manage membership of one or more distributed artificial intelligence groups throughout different phases of a life-cycle of an agent on the associated distributed artificial intelligence task.
12. The network entity as claimed in any of claims 9 to 11, wherein the controller is configured to discover one or more agents (704, 710, 710a-d, 711, 711a-d, 802-809, 901, 1200, 1302-1308, 1400) in the network by sending a discovery request to the network repository function, wherein the discovery request comprises one or more of: a type of network entity hosting the agent, a service area, a distributed artificial intelligence capability of the agent, an indication of a type of supported distributed artificial intelligence tasks, agent load and remaining agent capacity.
13. The network entity as claimed in any of claims 9 to 12, wherein the controller is configured to receive a discovery request response from the network repository function, the discovery request response comprising one or more of: endpoint address(es) of one or more discovered network entities hosting respective agents and respective profile(s) of the discovered agent(s).
14. The network entity as claimed in claim 13, wherein the controller is configured to receive a discovery request response from the network repository function comprising a respective profile for each of multiple discovered agents and select one or more agents of the multiple discovered agents to join a distributed artificial intelligence group in dependence on the respective profile for each of the multiple discovered agents.
15. The network entity as claimed in any of claims 9 to 14, wherein the controller is configured according to one of the following:(i) the controller is a dedicated controller for a distributed artificial intelligence group;(ii) the controller is a shared controller across distributed artificial intelligence groups of the same network entity type; and(iii) the controller is a shared controller across distributed artificial intelligence groups of different network entity types.
16. The network entity as claimed in any of claims 9 to 15, wherein the controller is configured to maintain distributed artificial intelligence group information for one or more distributed artificial intelligence groups that it is configured to control.
17. The network entity as claimed in claim 16 as dependent on claim 13 or claim 14, wherein the controller is configured to include one or more discovered agents in a distributed artificial intelligence group based on the respective profile(s) of the discovered agent(s), agent resource availability and the maintained distributed artificial intelligence group information.
18. The network entity as claimed in any preceding claim, wherein the service-based architecture is a 3GPP 5G network and wherein the agent(s) is / are deployed as sub-functionality of one or more of a user plane function, a control plane function, an access management function, a session management function, a service communication proxy and a policy control function.
19. A method for implementation at a network entity for a service-based architecture network, the network entity being configured to implement an agent for performing distributed artificial intelligence for one or more tasks, each task being associated with one or more distributed artificial intelligence groups comprising two or more agents in the network, the methodcomprising sending a registration request to a network repository function of the network, the registration request comprising a profile indicating a distributed artificial intelligence capability of the agent.
20. A method for implementation at a network entity for a service-based architecture network, the network entity being configured to implement a controller for managing distributed artificial intelligence for one or more tasks, each task being associated with one or more distributed artificial intelligence groups comprising two or more agents in the network, the method comprising sending a registration request to a network repository function of the network, the registration request comprising a profile indicating a distributed artificial intelligence controller functionality.
21. One or more computer programs for instructing a computer comprising one or more processors to implement the method of claim 19 or claim 20.
22. A data carrier storing in non-transitory form the one or more computer programs as claimed in claim 21.
Citation Information
Patent Citations
System and methods for supporting artificial intelligence service in a network
US20220060390A1
Communication method, apparatus, and system
US20230224752A1