Systems and methods for effective conflict resolution between intents

The MARL system with a supervisor agent effectively manages conflicting intents in wireless communication networks by generating sub-goals that optimize network parameters, addressing the challenges of resource constraints and dynamic changes in network demands.

WO2025104739A1PCT designated stage expired Publication Date: 2025-05-22TELEFONAKTIEBOLAGET LM ERICSSON (PUBL) +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/IN2024/050776
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-06-11
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing technologies face challenges in effectively managing conflicting intents in wireless communication networks, particularly in resource-constrained environments where optimization for one intent may adversely affect others.

Method used

A multi-agent reinforcement learning (MARL) system with a supervisor agent that generates sub-goals for agents to optimize network parameters, using a fusion layer to combine agent embeddings, global goals, and penalty contexts to adapt to changing priorities and resource availability.

Benefits of technology

The system enables efficient conflict handling and adaptation to dynamic changes in network demands, ensuring optimal fulfillment of multiple intents while maintaining high performance in resource-constrained settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2024050776_22052025_PF_FP_ABST
    Figure IN2024050776_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A method of operating a supervisor agent for collaboration with a plurality of trained agents in a multi-agent reinforcement learning (MARL) system is provided. The supervisor agent receives an embedding (mi t) for each of the respective agents, where the embedding represents a performance of the respective agent. The supervisor agent receives one or more global goals for the MARL system, and a penalty context (p t+1 ) that is based on one or more utility function values associated with the environment. The supervisor agent generates a context representation (c t+1 ) based on the embeddings, the one or more global goals and the penalty context, and applies a goal policy to the context representation to generate a plurality of sub-goals (g N t+1 ), each of which is associated with one of the agents. The supervisor agent provides the sub-goals to the agents to be used as goals for training local policies of the agents.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR EFFECTIVE CONFLICT RESOLUTION BETWEEN INTENTSBACKGROUND[1] The operation of future wireless communication networks, such as 6G networks, is expected to be driven primarily by intents provided by network operators. Such intents may have a defined priority of fulfillment and may consist of one or more objectives. The intent objectives may be defined in the form of targets for single or multiple key performance indicators (KPIs). An "intent" can be defined as a formal specification of all expectations including requirements, goals, and constaints given to a technical system. An intent defines the state or states a systems should attempt to reach, without explicitly defining how to reach those states. In natural language, an intent could take to a form such as "I want a conversational video service, where 80% of the users have a QoE of at least 4.0." Based on that intent, different agents may be instantiated to ensure intent fulfillment.[2] 5G and later wireless networks employ the concept of network slicing to provide customized, isolated, and dedicated network segments to meet the specific requirements of different services, applications, or user groups. A network "slice" refers to a logical network that is a partition or virtualization of a larger network. Each network slice may have its own set of resources and characteristics, and may be designed to meet particular service requirements.[3] There may be multiple intents associated with a single network slice. Each such intent may embody different expectations based on multiple KPIs. Constraints on available radio resources combined with varying demands from customers can lead to conflicting intents for the same network slice. For example, actions that fulfill certain expectations within a network slice may adversely affect the fulfillment of other intents within the slice.[4] In general, an intent management system may include multiple agents, each of which targets, or attempts to optimize or meet, one or more KPIs by controlling one or more network parameters. It is important for such agents to coordinate in order to handle conflicts. Such coordination may be managed by an intent manager.[5] To facilitate such coordination, each intent may be decomposed into more than one objective by the intent manager. For example, within a network slice, there may be multiple objectives, each of which needs one or more parameters to be optimized. In the telecom domain, each such objective measured by a KPI goal could be achieved by one agent that is trained to optimize a respective network parameter. Quite often, the operation of theseagents interacts, such that any change implemented by one agent may affect the KPIs targeted by other agents. Optimization for each objective in isolation, or in some cases sequentially, is possible, but a challenge arises when the objectives conflict with each other and a model of the environment is not available. In such circumstances, the intent manager must be able to autonomously manage multiple conflicting objectives and adapt to the expectations based on priority of individual intents.[6] A classical optimization technique is often not suitable because a model of the environment is not available. Also, the computation for optimization may not be available at run time, as the targets for the goals may change frequently.SUMMARY[7] Some embodiments provide a method of operating a supervisor agent for collaboration with a plurality of trained agents in a multi-agent reinforcement learning (MARL) system. The plurality of agents execute respective local policies on an environment. The method includes receiving, at a fusion layer, an embedding (mi) for each of the respective agents, the embedding representing a performance of the respective agent. The supervisor agent receives, at the fusion layer, one or more global goals for the MARL system, and a penalty context pt+i) that is based on one or more utility function values associated with the environment. The supervisor agent generates, at the fusion layer, a context representation (ct+i) based on the embeddings, the one or more global goals and the penalty context, and applies a goal policy to the context representation to generate a plurality of subgoals (g'v / + / ), each of which is associated with a respective one of the agents. The supervisor agent provides the sub-goals to respective ones of the agents to be used as goals for training the respective local policies of the agents.[8] The embedding for each agent may be generated based on an encoded capability vector for the agent and based on a state-action-goal tuple associated with the agent.[9] The encoded capability vector of the agent may include a set of goals for the agent and a set of probabilities, associated with respective ones of the goals, of the agent achieving the goal within a specified time period.

[0010] The embedding for each agent may be created by merging the encoded capability vector and the state-action-goal tuple of the agent.

[0011] The penalty context may be generated based on one or more utility function values.

[0012] The utility function values may be normalized by maximum values of the utility function values.

[0013] The penalty context may be generated as the output of a neural network that takes the one or more utility function values as inputs. In some embodiments, the penalty context may be generated by a utility function network, wherein an output of the neural network is generic to a form of utility function values.

[0014] The output of the neural network may be generic as to priority of utility functions used to generate the utility function values.

[0015] The fusion layer may include a fully connected neural network.

[0016] The context representation may be used to train the goal policy.

[0017] The MARL system may be generic with respect to a form of the utility function and / or with respect to changes in the environment.

[0018] Some embodiments provide a supervisor agent that is adapted to perform the foregoing operations.

[0019] Some embodiments provide a supervisor agent including a processing circuitry and a memory coupled to the processing circuitry. The memory includes computer program instructions that, when executed by the processing circuitry, cause the supervisor agent to perform operations described above.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 illustrates an MARL environment in which embodiments of the inventive concepts may be advantageously employed.

[0021] Figure 2 illustrates a block diagram for performing ad-hoc teaming to generate sub-goals for agents in a MARL system according to some embodiments.

[0022] Figure 3 illustrates a supervisor according to some embodiments.

[0023] Figure 4 illustrates operations of systems / methods according to some embodiments.

[0024] Figure 5 illustrates a plot of the distributions of UEs used for verification testing of embodiments.

[0025] Figure 6 illustrates simulated performance of the agents in maintaining KPIs in a system according to some embodiments.

[0026] Figure 7 shows an example of a communication system in accordance with some embodiments.

[0027] Figure 8 shows a UE in accordance with some embodiments.

[0028] Figure 9 shows a network node in accordance with some embodiments.

[0029] Figure 10 is a block diagram of a host in accordance with various aspects described herein.

[0030] Figure 11 is a block diagram illustrating a virtualization environment in which functions implemented by some embodiments may be virtualized.

[0031] Figure 12 shows a communication diagram of a host communicating via a network node with a UE over a partially wireless connection in accordance with some embodiments.DETAILED DESCRIPTION

[0032] Some approaches employ a model free technique using Multi-agent Reinforcement Learning (MARL) to solve the problem of managing intents. In a MARL- based system, multiple goal-conditioned Reinforcement Learning (RL) agents are employed to meet predetermined intents. Each of the agents is responsible for tuning a parameter related to the objective defined by the intent.

[0033] In this approach, intents will provide different goals for the multiple RL agents depending on their scope and the agents will perform all necessary monitoring, decision and actuation tasks on the managed entities.

[0034] RL provides a set of techniques that enable an agent to learn to interact with an environment by iteratively exploring and evaluating the outcomes of its actions. The overall goal in RL is to derive policies that map the observed situations perceived by the agent to actions that would maximize the cumulative rewards received by such an agent.

[0035] There are many cases where multiple agents could be deployed in the environment but they need to learn, based on context, either collaboration, competition or both. To learn these, MARL techniques may be used. However, learning the joint policy is not trivial, as agents usually have a local perception of the whole system (partially observable), and, in many cases, direct communication between them is absent. These and other characteristics define environments in which it is difficult for a learning agent to separate the effect of its actions from dynamics caused by other agents or other unobserved components.

[0036] The foremost technique to learn the collaboration between agents is independent Q-learning. However, a problem with this technique is that since each agent will take independent action, it will result in non-stationary environment. To overcome this problem, value function decomposition methods such as QMIX have been proposed. In the QMIXmethod, the agents are trained based on combined value function Qtot, which is obtained by combining the individual agent Q-values Qi using a neural network. In QMIX, the primary goal is to guarantee that decentralized agents’ decisions are consistent with those of a centralized counterpart. That is enforced during the training phase, allowing actuation to be performed in a decentralized fashion (centralized learning with decentralized execution).

[0037] Each agent’s policy (i = 1, . . . , ri) is derived from a Qi function. QMIX adds to that setup a mixing network Q(s, a), with joint states s = (s1, ... , .11) and joint actions a = (a1, ... , a"), from which a joint policy X$) is derived. During training, the loss function directs the learning of ^"towards producing optimal joint actions while directing Qi value-functions towards Q. This strategy induces individual policies X^!) such that X«) = (X^7), ■ ■ • , ")). This enables the system to employ TH for decentralized execution, instead of the centralized x.

[0038] In scenarios such as intent-based service assurance, RL policies are required that can adapt to changes in the goals during the execution phase. As an example, suppose an agent is pursuing a certain level of quality of experience, and as over time the target level changes, the agent should seamlessly continue to pursue the new goal, i.e., the agent generalizes over the domain of goals. One way of achieving such results is via goal- conditioned reinforcement learning. These aspects can be formalized by state-action value functions Q(o, g, a), where o and g come from the KPI domain and reward functions r(o, g) = d(o, g), where d refers to any similarity measure. During training, g values are randomly chosen from the domain at the beginning of every episode, so that the resulting policy is able to generalize over different goals during the execution phase. The simpler way of implementing goal-conditioned RL is by adding the goals g as an additional dimension in the observation space. That effectively changes the observation space. Although there might be limitations to such an approach, it allows to easily employ out-of-the-box RL algorithms in a goal-conditioned setting.

[0039] As an example in the telecom setting, a conflict situation may arise where an Intent Manager has to fulfill intents for different services, such as Conversational Video (CV), Ultra Reliable Low Latency Communication (URLLC) and massive loT (mloT). Each of these may involve optimizing multiple parameters in the network. In the examples described herein, the parameters that can be adjusted by the agents are packet priority and Maximum Bit Rate (MBR). Considering a resource constrained environment, an increase of packet priority might improve the Quality of Experience (QoE) of the CV service but might degrade the packet loss of URLLC service. Similarly, increasing MBR for URLLC mayimprove packet loss and reduce latency for URLLC but may degrade the mloT and CV service. Additionally, the target for each of the services might change frequently and thus the model may need to respond to the change without any additional training cycles. So in summary, the crux of the challenge is to optimize the realization of the intents, some of which may conflict with each other, in a resource constrained network setting.

[0040] To achieve this goal, the three agents learn to "plan to coordinate" to achieve an optimal global trade-off during the training phase. Additionally, during execution phase, none of the agents may need to observe actions or rewards (+ / - towards goal) from another agent. Each agent can see only the aggregated global reward and the current KPI (parameter value) for its own service. Thus, this design may reduce communication bottlenecks during execution.

[0041] An MARL environment in which embodiments of the inventive concepts may be advantageously employed is illustrated in Figure 1. As shown therein, a plurality of lower- level agents 112, 114 are trained to take specific type of actions, i.e. adjusting priority and adjusting MBR, for specific applications or protocols, such as CV, URLLC and mloT. As shown in Figure 1, each agent 112, 114 is implemented by a first multilayer perceptron (MLP) that receives observations ot,i and previous actions ut-i,i . The inputs are processed by a gated recurrent unit (GRU) and an output MLP, which generates a Qi value.

[0042] The agents 112, 114 generate respective joint actions 113, 115, which are provided to a supervisor agent 110. Using rule-based switching, the supervisor agent alternates between the agents through a rule-based switching based on intents 116 to decide which joint action to take on the environment, which in this case is a network emulator 200 that emulates various aspects of a wireless communication network, including base stations (gNBs), core network functions such as the user plane function (UPF) and various services.

[0043] In this scenario, multiple closed loops are interacting while optimizing their particular KPIs by tuning certain parameters. Those characteristics translate to MARL formalism as multiple heterogeneous agents for which the conflicting demands for resources require cooperative behavior.

[0044] In some embodiments, traditional QMIX is used to train agents able to cooperate without communication in decentralized settings. The objective is to arrive at goal- conditioned agents able to adapt to dynamic changes in the intents while also maintaining high performance in an open MARL setup, where agents may join or leave at any time. The RL components defined on top of the network emulator infrastructure are as follows.

[0045] 1) Observation space: Since multiple heterogeneous agents are employed, the local observation spaces are first specified, and later the joint or global observation space.

[0046] Each service type is associated with a single KPI, e.g., while the quality of a conversational video service is measured in terms of QoE, the quality of an URLLC service might be measured in terms of packet losses. Overall, the first dimension of the local observation spaces is always the respective KPI measurement, while the second dimension is the KPI target defined by the intents. To facilitate the emergence of cooperation during training, the global reward G e 91 and the number of UEs n e K using the service are also included as additional dimensions.

[0047] Considering T as a service type index, and KT as the KPI ranges associated with T, the local observation spaces are defined as .S' / e KTX KT X R X N. The specific KPI and their domains are specified in Table 1. For easy notation, consider s = (o7, gj , G, nj) e Sj , where j is an index referring to the service types, and Oj, gj e .

[0048] Table 1 - Service Types, KPIs and Domains# Service type T KPI1 Conversational video (CV) QoE [1, 5]2 URLLC Packet- loss ratio (PLR) [0, 1]3 mloT Packet-loss ratio (PLR) [0, 1

[0049] Altogether, the local observation spaces compose a joint observation space that can be interpreted as the global state of the MARL environment. Such joint observation space is based on the concatenation of all local observation spaces, which is composed of elements s e S.

[0050] 2) Reward function: Analogous to the observation spaces, the reward functions are also divided in local and global. The local reward functions concerns the individual KPIs domains, and are defined to comply with the goal-conditioning setup, i.e., r7(o, g) = \o-g\, where o, g e KT . Such reward functions would output higher values as the current KPI measurement s is further away from the target g in either direction.

[0051] Additionally, as the system includes heterogeneous agents and each reward function may be in a different range, the individual rewards are scaled down into the range [0, 1] by employing a normalization factor Aj that represents the maximum absolute differencebetween Oj and gj. Therefore, given a pair (o7, gj), the reward function for a service j e T is defined as:

[0052] From local rewards functions, local observations and goals, a linear global reward function can be defined. Here, the notion of penalties pj (or preferences) is introduced. If all penalties are equal, all services should be equally served, otherwise, those services with higher penalties should be favored at the expense of the others.

[0053] Penalties are particularly important whenever there are not enough resources to achieve all intents. In such cases, equal penalties lead to equal degradation of the service KPIs. By adding the penalty terms to the global reward function, the desired behavior can be fine tuned by properly choosing different penalty terms for each service.

[0054] All the above components are same for both the packet priority MARL agent and MBR MARL agent. However, their action space is different and is explained below.

[0055] 3) Action spaces: Service assurance for each service type can be pursued via tuning of multiple configuration parameters (control knobs), i.e., for each service there can be multiple agents. In the example herein, two control knobs are used: the MBR and packet priority. Therefore, there are two types of agents, referred to as MBR agents and packet priority agents, both trained via RL. Since the QMIX algorithm was used, both action spaces were enforced to be discrete.

[0056] In the case of MBR agents, the actions consist of decisions for increasing or decreasing the current MBR value, i.e., A = {-1, 1 }. The increment / decrements are fixed and defined from the discretization of the original action space [1, a] into bins of size 0.5 Mbps (where a is the air-link bandwidth). In this setting an increase / decrease action would move to the next / previous bin. After the updated bin bt= bt-i + atis identified, a random value is chosen within its bounds to specify the UE’s MBR.MBRt= rand(7? / ) [3]

[0057] Similarly, in the case of priority agents, the actions consist of decisions for increasing or decreasing the current packet priority of a service, i.e., A = {-1, 1 }. The minimum / maximum priority values are defined to comply with the Network Emulator design, e.g., [1, 100], where a small value denotes high priority.Priorityt= Priorityt-i + at[4]

[0058] MARL training and evaluation

[0059] Each agent is modelled by 2-layer Gated Recurrent Unit (GRU) network with 2 nodes in each layer. During training, QMIX benefits from local and global observations and rewards, i.e., centralized training. The training phase consists of multiple episodes, each lasting for a maximum of H = 30 time steps or until all agents have reached their intents. During the evaluation phase, agents only rely on their local observations to reach their individual goals, i.e., decentralized execution.

[0060] Packet priority and MBR agents have different scopes, as specified by the network emulator used for the experiments. While packet priority is defined at the service level, MBR is defined at the UE level. Although it is scalable to set packet priority for each service, it may not be realistic to set MBR to individual UEs. Therefore, MBR agents take decisions for sets of UEs instead.

[0061] The rationale for grouping UEs depend on the format of the intents supported. Each intent follows a scheme "x% of UE’s to have KPI greater than g". Hence, the UEs from service can be divided into two groups of sizes x% and (100 - x)%. The MBR agents’ actions then affect all the UEs within a group equally, with all of them receiving the same MBR value. If the intent is on 100% of UEs, then the MBR actuation reduces to the service level, analogously to packet priority agents.

[0062] Both priority and MBR agent groups are trained independently to reduce the computational overhead. Additionally, it may be assumed that the agents never actuate at the same time. Given that the agents have different strengths and weaknesses, such an approach requires an additional level of decision-making specifying which agents to activate and deactivate during execution time. In prior approaches, a very simple supervisor agent was used, which chooses the specific agent to activate based on knowledge about how thenetwork emulator works. Example rules governing the supervisor agent are summarized in Table 2.

[0063] Table 2 - Supervisor Agent and Its Rules# Rufe Agent to useA. If all UEs throughput = MBR, MBR agentsB. Otherwise, Priority agents

[0064] The supervisor agent decisions take place every 5 time steps, after which the chosen agent (MBR or Priority) stays active for another round.

[0065] There are a few drawbacks to using a rules-based supervisor agent, however. For example, the supervisor agent works on the basis of heuristics and rules which are designed by a human expert. Also, the rules may depend on various factors e.g. radio environment, which is difficult for humans to comprehend with an increase in scale and complexity.

[0066] Moreover, since the rules are pre-configured, the execution for each system (5 time-steps) has to occur for a pre-determined time period and is not conditioned on the state of the environment. Ideally, this approach is not optimal as the optimal actions of MARL agents depend on the state of the environment prevailing at that point in time.

[0067] Sequential execution of each task, (either priority or MBR actions), increases the total time window for intent fulfilment. Parallel execution is not feasible in since the supervisor cannot deal with the non-stationarity of simultaneous actions.

[0068] Additionally, as the radio environment changes frequently, an agent trained on a specific case might not work satisfactorily on other cases.

[0069] Finally, every intent is associated with utility function which has the information on how important is achieving the intent. Quite often the form of the utility function also changes depending on the requirements of the customer. The approach shown in Figure 1 might not handle the changes in form of utility function.

[0070] Certain aspects of the disclosure and their embodiments may provide solutions to these or other challenges. In an intent manager according to some embodiments, multiple agents are trained using MARL or MR. An ad-hoc teaming approach is integrated with MARL to the limitations described above. Lower-level MARL agents similar to the ones shown in Figure 1 are trained. Then, AHT approaches are used to train a supervisor agent that can create efficient sub-goals for the lower-level agents. With the top-level agent, bettergeneralization capabilities can be achieved with respect to changes in the environment, changes in the form of the utility function, and efficient conflict handling.

[0071] Certain embodiments may provide one or more of the following technical advantage(s). Some embodiments may help agents to achieve better generalization with respect to the form of utility function being used without any retraining. Moreover, some embodiments may help agents to achieving better generalization with respect to changes in the environment.

[0072] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art.

[0073] Some embodiments described herein train the supervisor agent using AHT approaches to generate sub-goals for agents in a MARL system. For this, some embodiments provide an ad-hoc teaming MARL (AT-MARL) system that applies ad-hoc teaming to the pre-trained MARL framework of Priority and MBR systems shown in Figure 1.

[0074] Figure 2 illustrates a block diagram for performing ad-hoc teaming to generate, at a supervisor node, sub-goals for agents in a MARL system. As shown therein, a supervisor agent 110 uses a fusion layer 130 to train a goal policy 124 that generates sub-goals g^i for a plurality of agents 112, 114. The fusion layer receives inputs from the agents 112, 114, as well as penalties pt+ifrom a utility function neural network 126 and global goals 122 determined by a network operator and generates a context representation ct+1 which is used to train the goal policy 7tg124. Each agent 112, 114 includes an encoder block 123 for encoding past performance information and a merger block 125 that merges past performance information with current performance information to form embeddingsthat are provided to the supervisor 110.

[0075] The encoder block 123 may be a fully connected network which reduces the dimensionality of the capability vector (for example, from 6 to 2). The merger block 125 may be a fully connected network which adds the current state tuple (st, at, gt) as three nodes to the merger network in addition to the nodes corresponding to the encoder output. The output of the merger block 125 is also a vector of length 2. For each agent 112, 114, an embedding vector of length 2 is obtained.

[0076] The fusion layer 130 may be a fully connected network with 16 nodes, 12 nodes corresponding to 6 agents (each of length 2), 3 nodes for the global goals (as there are three different intents) and 1 node for the utility context (output of the utility network).

[0077] The utility function network 126 which takes the utility function values is also fully connected. It has three input nodes and one output node. The output of the utility function network is the utility context, expressed as a penalty context pt+i. The penalty context pt+imay be generated based on one or more utility function values Pi ... PN, which may be normalized by maximum values of the utility function values Pn,max. Providing a penalty context pt+ito the supervisor agent 110 that is generated by the utility function network 126 based on normalized utility function values Pi ... PN may enable the supervisor agent 110 to handle any form of utility function, which is advantageous in a real-time network where customer expectations may vary.

[0078] The output of the fusion layer 130 is a context vector of length 4. This is sent as state input to the RL agent which executes the goal policy 7tg124.

[0079] The operation of each block is described in greater detail below.

[0080] The approach consists of training a goal policy (within the supervisor agent) that generates the sub-goals. To train the policy, two inputs are required, namely (i) each individual agent’s capabilities and (ii) each individual agent’s current performance.

[0081] At first, the individual agent’s capabilities are input in the form of a vector at each time instant t. The capability vector Ftfor agent I at time t represents the probability of agent i achieving goalSpecifically,p = 1, • • • , P is set of goals for the agent i and 0 < p't.p < 1 represents the probability that agent i will achieve specific goal p within time instant t. The capability vector i 'jtis encoded using an encoder (fully connected network) to project into low-dimensional space.

[0082] As noted above, the agent’s current performance is input for better sub-goal assignment. For this, a state-action-goal tuples'. ^') °ft^eagent is created, where si is the state of the agent i at time step t, nj. is the action taken by the agent i at time step t andis the goal assigned for agent i at time step t. This tuple is merged with the encoder output obtained earlier to create an embedding ???;,.

[0083] For every agent, a separate encoder and merger are used to create N embeddings i = 1 , • • • , N. Finally, the embeddings are computed for all the agents and fused using the fusion layer. Also, the global goals (derived from intents) are input to the fusion layer since it is desirable for the sub-goals generated to be conditional on the global goals.

[0084] In addition, the utility function is input to the agents to create efficient sub-goals. To make it generic instead of passing the absolute values of the utility function, the relative values are passed to the network. For this one more fully connected network is used which consumes the ratio of the current utility function value to maximum possible utility function valueNormally a domain expert can have the idea on maximum utility function value which is worst possible.

[0085] The ratio for each agent i is computed and the values are passed to a fully connected network to create a combined utility contextfor time instant t+I.

[0086] The output of the fusion layer obtained by fusing the N embeddingsi = 1, • •• , N at time instant t, global goals and utility contextgives the context representation for the future time step denoted asx, which can be used to train the goal policyto generate the sub-goalsi = 1, • • • , N for next time step. In some embodiments, the RL- based actor-critic approach model is used to train the goal policy.

[0087] The actor may be chosen to be a 2-layer gated recurrent unit (GRU) network and critic a 2-layer fully connected network. Here, fully connected layers are used to represent the encoder (2-layer), merger(l-layer), and fusion layers(3-layers).

[0088] In this way, a model can be generated that is generic with respect to the form of the utility function and / or radio environment changes. Also, since embodiments described herein may use parallel execution, some embodiments may be more efficient and / or converge faster when compared with prior approaches that do not use an ad-hoc teaming approach to generate sub-goals.

[0089] Figure 3 illustrates a supervisor 110 according to some embodiments implemented as a standalone system including a processing circuitry 312, a memory 314 coupled to the processing circuitry or that stores computer-readable instructions executableby the processing circuitry 312 for performing the operations described herein, and a transceiver 316 coupled to the processing circuitry for communicating with an environment, such as a wireless communication system.

[0090] Figure 4 illustrates operations of systems / methods according to some embodiments. As shown therein, a method of collaboration between a plurality of trained agents in a multi-agent reinforcement learning, MARL system is provided, wherein the plurality of agents execute respective local policies on an environment. The method includes receiving (block 402), at a fusion layer, an embedding (m\) for each of the respective agents, the embedding representing a performance of the respective agent, one or more global goals for the MARL system (block 404), and a penalty context pt+i) that is based on one or more utility function values associated with the environment (block 406). The supervisor agent generates (block 408), at the fusion layer, a context representation (ct+1) based on the embeddings, the one or more global goals and the penalty context, applies (block 410) a goal policy to the context representation to generate a plurality of sub-goals (gNi+ / ), each of the sub-goals being associated with a respective one of the agents, and provides (block 412) the sub-goals to respective ones of the agents to be used as goals for training the respective local policies of the agents.

[0091] The embedding for each agent may be generated based on an encoded capability vector for the agent and based on a state-action-goal tuple associated with the agent.

[0092] The encoded capability vector of the agent may include a set of goals for the agent and a set of probabilities associated with respective ones of the goals, of the agent achieving the goal within a specified time period.

[0093] The embedding for each agent may be created by merging the encoded capability vector and the state-action-goal tuple of the agent.

[0094] The penalty context may be generated based on one or more utility function values. The utility function values may be normalized by maximum values of the utility function values. The penalty context may be generated as the output of a neural network that takes the one or more utility function values as inputs.

[0095] The fusion layer may be implemented using a fully connected neural network.

[0096] The context representation may be used to train the goal policy.

[0097] To evaluate the method described herein, a network emulator was used to create three services (i) Conversational Video (CV), (ii) URLLC and (iii) mloT. The emulator has capability to route traffic from / to application layer to / from UE’s through UPF and gNodeB. For each of the service, a KPI to control is selected, e.g., QoE for CV, Packet Error Rate(PER) for URLLC and PER for mloT service. To control the above KPI’s, two different MARL agents were trained independently which can modify priority and MBR respectively. The supervisor agent 110 can change sub-goals for each individual agent within each MARL agent.

[0098] For the sake of comparison, three different scenarios were created including one in which sub-goals were created at the service level, one in which sub-goals were generated at the agent level, and oe in which sub-goals were generated by a supervisor using a rules- based approach.

[0099] For the approach using sub-goals generated at the service level, one goal was generated for each KPI and halved it before being sent to each individual agent within a MARL agent. For example:QoE_Priority = QoE_Goal / 2QoE_MBR = QoE_MBR / 2

[0100] For the approach with sub-goals generated at the agent level, a supervisor agent was used to generate sub-goals for each individual agent, namely, QoE_Priority, QoE_MBR, etc.

[0101] For the rule-based approach, a rules-based supervisor agent was used to perform switching between the agents.

[0102] To quantify the performance a metric known as the integral absolute relative error (IAE), which is widely used in closed-loop control, was used. The IAE is measured as:

[0103] Some embodiments may generalize with respect to the form of utility function used. To quantify the generalization with respect to the form of the utility function, two forms of utility function are considered:1. Form rgff2. FormTarget j

[0104] Form 1 ensures the effect of the deviation is more on the value of the utility function which ensures the KPI with highest function value converges faster than other KPI’s. However, form 2 ensures the effect of the deviation is comparatively less and it makes the KPI’s converges slower when compared to earlier case of form 1.

[0105] The system was trained with form 1 with the value of p. = itVi =f2t3-The system was tested with different values of s;varying between 1 to 10 and also when the form changed to form 2.

[0106] Table 3 shows the IAE values obtained for each KPI when the same form of utility function but different values of p;are used.Table 3 — IAE Values With Different Values of pi

[0107] From Table 3, it can be seen that as the value of p, increases the deviation of KPI ‘i’ decreases. This indicates that the value of p, is directly proportional to the importance of the KPI and it should converge faster to the target. Also, as the form of utility function is linear, it shows that the effect of deviation is more on the value of i.

[0108] Also, experiments were performed when the form of the utility function changes. Table 4 illustrates the results when agents trained on form 1 of the utility function were applied using form 2 of the utility function.Table 4 — IAE Values With Different Utility Function

[0109] From the values shown in Table 4, it can be seen that the IAE values obtained for form 2 are not higher when compared with form 1 , which agrees with intuitive understanding.

[0110] Systems / methods according to some embodiments may also generalize with respect to the utility function used. To demonstrate this, three different scenarios were created namely:1. UE’ s distributed uniformly across 4 gNodeB ’ s.2. UE’s distributed in normal distribution across 4 gNodeB ’s.3. UE’s distributed in Gamma distribution across 4 gNodeB’s.

[0111] A plot of the distributions of UE’s is shown Figure 5.

[0112] For case 1 in which the UE’s are distributed uniformly across 4 gNodeB’s, the IAE values obtained for all the different scenarios are shown in the table below. From Table 5 it can be seen that an AT-MARL approach according to some embodiments results in lower IAE values when compared with other two cases, especially the previous rule-based approach.

[0113] Table 5 — IAE calculated on KPIs compared Across Approaches — Uniform Distribution

[0114] For case 2 in which the UE’s are distributed in a Gaussian way across 4 gNodeB’s, the IAE values obtained for all different scenarios are shown in Table 6.

[0115] Table 6 — IAE calculated on KPIs compared Across Approaches — Gaussian Distribution

[0116] From Table 6, it can be seen that the approach described herein (AT- MARL) gives lower degradation when compared to existing approaches.

[0117] Also, the degradation obtained using the AT-MARL is almost similar in the previous two cases suggesting the good generalization capabilities. Also, the rule-based approach results in poor generalization capabilities which can be seen huge difference in degradation values.

[0118] Finally, in real systems, the UE distribution may change with time. Thus, the AT-MARL system needs to compensate for dynamic changes in UE distribution. For this a scenario was created in which the distribution changes in a continuous fashion.

[0119] The simulated performance of the agents in maintaining the KPIs is shown in Figure 6, which is a graph of QoE for the CV service 602 and packet loss (PL) percentage for mloT 604 and URLLC 606 services compared to target values (dashed lines). Initially, the UEs are uniformly distributed. After 20 time steps, the distribution was changed to Gaussian, and after 30 time steps, the distribution was changed to Gamma. The results shown in Figure 6 indicate that the AT-MARL approach results in better generalization relative to UE distribution.

[0120] Also, from both results, it can be seen that the AT-MARL approach results in better conflict handling and better generalization capabilities.

[0121] Figure 7 shows an example of a communication system 700 in which some embodiments may be advantageously employed.

[0122] In the example, the communication system 700 includes a telecommunication network 702 that includes an access network 704, such as a radio access network (RAN), and a core network 706, which includes one or more core network nodes 708. The access network 704 includes one or more access network nodes, such as networknodes 710a and 710b (one or more of which may be generally referred to as network nodes 710), or any other similar 3rdGeneration Partnership Project (3GPP) access nodes or non- 3GPP access points. Moreover, as will be appreciated by those of skill in the art, a network node is not necessarily limited to an implementation in which a radio portion and a baseband portion are supplied and integrated by a single vendor. Thus, it will be understood that network nodes include disaggregated implementations or portions thereof. For example, in some embodiments, the telecommunication network 702 includes one or more Open-RAN (ORAN) network nodes. An ORAN network node is a node in the telecommunication network 702 that supports an ORAN specification (e.g., a specification published by the O- RAN Alliance, or any similar organization) and may operate alone or together with other nodes to implement one or more functionalities of any node in the telecommunication network 702, including one or more network nodes 710 and / or core network nodes 708.

[0123] Examples of an ORAN network node include an open radio unit (O-RU), an open distributed unit (O-DU), an open central unit (O-CU), including an O-CU control plane (O-CU-CP) or an O-CU user plane (O-CU-UP), a RAN intelligent controller (near-real time or non-real time) hosting software or software plug-ins, such as a near-real time control application (e.g., xApp) or a non-real time control application (e.g., rApp), or any combination thereof (the adjective “open” designating support of an ORAN specification). The network node may support a specification by, for example, supporting an interface defined by the ORAN specification, such as an Al, Fl, Wl, El, E2, X2, Xn interface, an open fronthaul user plane interface, or an open fronthaul management plane interface. Moreover, an ORAN access node may be a logical node in a physical node. Furthermore, an ORAN network node may be implemented in a virtualization environment (described further below) in which one or more network functions are virtualized. For example, the virtualization environment may include an O-Cloud computing platform orchestrated by a Service Management and Orchestration Framework via an O-2 interface defined by the O- RAN Alliance or comparable technologies. The network nodes 710 facilitate direct or indirect connection of user equipment (UE), such as by connecting UEs 712a, 712b, 712c, and 712d (one or more of which may be generally referred to as UEs 712) to the core network 706 over one or more wireless connections.

[0124] Example wireless communications over a wireless connection include transmitting and / or receiving wireless signals using electromagnetic waves, radio waves, infrared waves, and / or other types of signals suitable for conveying information without the use of wires, cables, or other material conductors. Moreover, in different embodiments, thecommunication system 700 may include any number of wired or wireless networks, network nodes, UEs, and / or any other components or systems that may facilitate or participate in the communication of data and / or signals whether via wired or wireless connections. The communication system 700 may include and / or interface with any type of communication, telecommunication, data, cellular, radio network, and / or other similar type of system.

[0125] The UEs 712 may be any of a wide variety of communication devices, including wireless devices arranged, configured, and / or operable to communicate wirelessly with the network nodes 710 and other communication devices. Similarly, the network nodes 710 are arranged, capable, configured, and / or operable to communicate directly or indirectly with the UEs 712 and / or with other network nodes or equipment in the telecommunication network 702 to enable and / or provide network access, such as wireless network access, and / or to perform other functions, such as administration in the telecommunication network 702.

[0126] In the depicted example, the core network 706 connects the network nodes 710 to one or more hosts, such as host 716. These connections may be direct or indirect via one or more intermediary networks or devices. In other examples, network nodes may be directly coupled to hosts. The core network 706 includes one more core network nodes (e.g., core network node 708) that are structured with hardware and software components. Features of these components may be substantially similar to those described with respect to the UEs, network nodes, and / or hosts, such that the descriptions thereof are generally applicable to the corresponding components of the core network node 708. Example core network nodes include functions of one or more of a Mobile Switching Center (MSC), Mobility Management Entity (MME), Home Subscriber Server (HSS), Access and Mobility Management Function (AMF), Session Management Function (SMF), Authentication Server Function (AUSF), Subscription Identifier De-concealing function (SIDE), Unified Data Management (UDM), Security Edge Protection Proxy (SEPP), Network Exposure Function (NEF), and / or a User Plane Function (UPF).

[0127] The host 716 may be under the ownership or control of a service provider other than an operator or provider of the access network 704 and / or the telecommunication network 702, and may be operated by the service provider or on behalf of the service provider. The host 716 may host a variety of applications to provide one or more service. Examples of such applications include live and pre-recorded audio / video content, data collection services such as retrieving and compiling data on various ambient conditions detected by a plurality of UEs, analytics functionality, social media, functions for controllingor otherwise interacting with remote devices, functions for an alarm and surveillance center, or any other such function performed by a server.

[0128] As a whole, the communication system 700 of Figure 7 enables connectivity between the UEs, network nodes, and hosts. In that sense, the communication system may be configured to operate according to predefined rules or procedures, such as specific standards that include, but are not limited to: Global System for Mobile Communications (GSM); Universal Mobile Telecommunications System (UMTS); Long Term Evolution (LTE), and / or other suitable 2G, 3G, 4G, 5G standards, or any applicable future generation standard (e.g., 6G); wireless local area network (WLAN) standards, such as the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standards (WiFi); and / or any other appropriate wireless communication standard, such as the Worldwide Interoperability for Microwave Access (WiMax), Bluetooth, Z-Wave, Near Field Communication (NFC) ZigBee, LiFi, and / or any low-power wide-area network (LPWAN) standards such as LoRa and Sigfox.

[0129] In some examples, the telecommunication network 702 is a cellular network that implements 3GPP standardized features. Accordingly, the telecommunications network 702 may support network slicing to provide different logical networks to different devices that are connected to the telecommunication network 702. For example, the telecommunications network 702 may provide Ultra Reliable Low Latency Communication (URLLC) services to some UEs, while providing Enhanced Mobile Broadband (eMBB) services to other UEs, and / or Massive Machine Type Communication (mMTC) / Massive loT services to yet further UEs.

[0130] In some examples, the UEs 712 are configured to transmit and / or receive information without direct human interaction. For instance, a UE may be designed to transmit information to the access network 704 on a predetermined schedule, when triggered by an internal or external event, or in response to requests from the access network 704. Additionally, a UE may be configured for operating in single- or multi-RAT or multistandard mode. For example, a UE may operate with any one or combination of Wi-Fi, NR (New Radio) and LTE, i.e. being configured for multi-radio dual connectivity (MR-DC), such as E-UTRAN (Evolved-UMTS Terrestrial Radio Access Network) New Radio - Dual Connectivity (EN-DC).

[0131] In the example, the hub 714 communicates with the access network 704 to facilitate indirect communication between one or more UEs (e.g., UE 712c and / or 712d) and network nodes (e.g., network node 710b). In some examples, the hub 714 may be acontroller, router, content source and analytics, or any of the other communication devices described herein regarding UEs. For example, the hub 714 may be a broadband router enabling access to the core network 706 for the UEs. As another example, the hub 714 may be a controller that sends commands or instructions to one or more actuators in the UEs. Commands or instructions may be received from the UEs, network nodes 710, or by executable code, script, process, or other instructions in the hub 714. As another example, the hub 714 may be a data collector that acts as temporary storage for UE data and, in some embodiments, may perform analysis or other processing of the data. As another example, the hub 714 may be a content source. For example, for a UE that is a VR headset, display, loudspeaker or other media delivery device, the hub 714 may retrieve VR assets, video, audio, or other media or data related to sensory information via a network node, which the hub 714 then provides to the UE either directly, after performing local processing, and / or after adding additional local content. In still another example, the hub 714 acts as a proxy server or orchestrator for the UEs, in particular if one or more of the UEs are low energy loT devices.

[0132] The hub 714 may have a constant / persistent or intermittent connection to the network node 710b. The hub 714 may also allow for a different communication scheme and / or schedule between the hub 714 and UEs (e.g., UE 712c and / or 712d), and between the hub 714 and the core network 706. In other examples, the hub 714 is connected to the core network 706 and / or one or more UEs via a wired connection. Moreover, the hub 714 may be configured to connect to an M2M service provider over the access network 704 and / or to another UE over a direct connection. In some scenarios, UEs may establish a wireless connection with the network nodes 710 while still connected via the hub 714 via a wired or wireless connection. In some embodiments, the hub 714 may be a dedicated hub - that is, a hub whose primary function is to route communications to / from the UEs from / to the network node 710b. In other embodiments, the hub 714 may be a non-dedicated hub - that is, a device which is capable of operating to route communications between the UEs and network node 710b, but which is additionally capable of operating as a communication start and / or end point for certain data channels.

[0133] Figure 8 shows a UE 800 in accordance with some embodiments. As used herein, a UE refers to a device capable, configured, arranged and / or operable to communicate wirelessly with network nodes and / or other UEs. Examples of a UE include, but are not limited to, a smart phone, mobile phone, cell phone, voice over IP (VoIP) phone, wireless local loop phone, desktop computer, personal digital assistant (PDA), wireless cameras,gaming console or device, music storage device, playback appliance, wearable terminal device, wireless endpoint, mobile station, tablet, laptop, laptop-embedded equipment (LEE), laptop-mounted equipment (LME), smart device, wireless customer-premise equipment (CPE), vehicle, vehicle-mounted or vehicle embedded / integrated wireless device, etc. Other examples include any UE identified by the3rd Generation Partnership Project (3GPP), including a narrow band internet of things (NB-IoT) UE, a machine type communication (MTC) UE, and / or an enhanced MTC (eMTC) UE.

[0134] A UE may support device-to-device (D2D) communication, for example by implementing a 3GPP standard for sidelink communication, Dedicated Short-Range Communication (DSRC), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), or vehicle-to-everything (V2X). In other examples, a UE may not necessarily have a user in the sense of a human user who owns and / or operates the relevant device. Instead, a UE may represent a device that is intended for sale to, or operation by, a human user but which may not, or which may not initially, be associated with a specific human user (e.g., a smart sprinkler controller). Alternatively, a UE may represent a device that is not intended for sale to, or operation by, an end user but which may be associated with or operated for the benefit of a user (e.g., a smart power meter).

[0135] The UE 800 includes processing circuitry 802 that is operatively coupled via a bus 804 to an input / output interface 806, a power source 808, a memory 810, a communication interface 812, and / or any other component, or any combination thereof. Certain UEs may utilize all or a subset of the components shown in Figure 8. The level of integration between the components may vary from one UE to another UE. Further, certain UEs may contain multiple instances of a component, such as multiple processors, memories, transceivers, transmitters, receivers, etc.

[0136] The processing circuitry 802 is configured to process instructions and data and may be configured to implement any sequential state machine operative to execute instructions stored as machine-readable computer programs in the memory 810. The processing circuitry 802 may be implemented as one or more hardware-implemented state machines (e.g., in discrete logic, field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc.); programmable logic together with appropriate firmware; one or more stored computer programs, general-purpose processors, such as a microprocessor or digital signal processor (DSP), together with appropriate software; or any combination of the above. For example, the processing circuitry 802 may include multiple central processing units (CPUs).

[0137] In the example, the input / output interface 806 may be configured to provide an interface or interfaces to an input device, output device, or one or more input and / or output devices. Examples of an output device include a speaker, a sound card, a video card, a display, a monitor, a printer, an actuator, an emitter, a smartcard, another output device, or any combination thereof. An input device may allow a user to capture information into the UE 800. Examples of an input device include a touch-sensitive or presence- sensitive display, a camera (e.g., a digital camera, a digital video camera, a web camera, etc.), a microphone, a sensor, a mouse, a trackball, a directional pad, a trackpad, a scroll wheel, a smartcard, and the like. The presence-sensitive display may include a capacitive or resistive touch sensor to sense input from a user. A sensor may be, for instance, an accelerometer, a gyroscope, a tilt sensor, a force sensor, a magnetometer, an optical sensor, a proximity sensor, a biometric sensor, etc., or any combination thereof. An output device may use the same type of interface port as an input device. For example, a Universal Serial Bus (USB) port may be used to provide an input device and an output device.

[0138] In some embodiments, the power source 808 is structured as a battery or battery pack. Other types of power sources, such as an external power source (e.g., an electricity outlet), photovoltaic device, or power cell, may be used. The power source 808 may further include power circuitry for delivering power from the power source 808 itself, and / or an external power source, to the various parts of the UE 800 via input circuitry or an interface such as an electrical power cable. Delivering power may be, for example, for charging of the power source 808. Power circuitry may perform any formatting, converting, or other modification to the power from the power source 808 to make the power suitable for the respective components of the UE 800 to which power is supplied.

[0139] The memory 810 may be or be configured to include memory such as random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic disks, optical disks, hard disks, removable cartridges, flash drives, and so forth. In one example, the memory 810 includes one or more application programs 814, such as an operating system, web browser application, a widget, gadget engine, or other application, and corresponding data 816. The memory 810 may store, for use by the UE 800, any of a variety of various operating systems or combinations of operating systems.

[0140] The memory 810 may be configured to include a number of physical drive units, such as redundant array of independent disks (RAID), flash memory, USB flash drive,external hard disk drive, thumb drive, pen drive, key drive, high-density digital versatile disc (HD-DVD) optical disc drive, internal hard disk drive, Blu-Ray optical disc drive, holographic digital data storage (HDDS) optical disc drive, external mini-dual in-line memory module (DIMM), synchronous dynamic random access memory (SDRAM), external micro-DIMM SDRAM, smartcard memory such as tamper resistant module in the form of a universal integrated circuit card (UICC) including one or more subscriber identity modules (SIMs), such as a USIM and / or ISIM, other memory, or any combination thereof. The UICC may for example be an embedded UICC (eUICC), integrated UICC (iUICC) or a removable UICC commonly known as ‘SIM card.’ The memory 810 may allow the UE 800 to access instructions, application programs and the like, stored on transitory or non-transitory memory media, to off-load data, or to upload data. An article of manufacture, such as one utilizing a communication system may be tangibly embodied as or in the memory 810, which may be or comprise a device -readable storage medium.

[0141] The processing circuitry 802 may be configured to communicate with an access network or other network using the communication interface 812. The communication interface 812 may comprise one or more communication subsystems and may include or be communicatively coupled to an antenna 822. The communication interface 812 may include one or more transceivers used to communicate, such as by communicating with one or more remote transceivers of another device capable of wireless communication (e.g., another UE or a network node in an access network). Each transceiver may include a transmitter 818 and / or a receiver 820 appropriate to provide network communications (e.g., optical, electrical, frequency allocations, and so forth). Moreover, the transmitter 818 and receiver 820 may be coupled to one or more antennas (e.g., antenna 822) and may share circuit components, software or firmware, or alternatively be implemented separately.

[0142] In the illustrated embodiment, communication functions of the communication interface 812 may include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communications such as Bluetooth, near-field communication, location-based communication such as the use of the global positioning system (GPS) to determine a location, another like communication function, or any combination thereof. Communications may be implemented in according to one or more communication protocols and / or standards, such as IEEE 802.11, Code Division Multiplexing Access (CDMA), Wideband Code Division Multiple Access (WCDMA), GSM, LTE, New Radio (NR), UMTS, WiMax, Ethernet, transmission control protocol / intcrnct protocol (TCP / IP),synchronous optical networking (SONET), Asynchronous Transfer Mode (ATM), QUIC, Hypertext Transfer Protocol (HTTP), and so forth.

[0143] Regardless of the type of sensor, a UE may provide an output of data captured by its sensors, through its communication interface 812, via a wireless connection to a network node. Data captured by sensors of a UE can be communicated through a wireless connection to a network node via another UE. The output may be periodic (e.g., once every 15 minutes if it reports the sensed temperature), random (e.g., to even out the load from reporting from several sensors), in response to a triggering event (e.g., when moisture is detected an alert is sent), in response to a request (e.g., a user initiated request), or a continuous stream (e.g., a live video feed of a patient).

[0144] As another example, a UE comprises an actuator, a motor, or a switch, related to a communication interface configured to receive wireless input from a network node via a wireless connection. In response to the received wireless input the states of the actuator, the motor, or the switch may change. For example, the UE may comprise a motor that adjusts the control surfaces or rotors of a drone in flight according to the received input or to a robotic arm performing a medical procedure according to the received input.

[0145] A UE, when in the form of an Internet of Things (loT) device, may be a device for use in one or more application domains, these domains comprising, but not limited to, city wearable technology, extended industrial application and healthcare. Non-limiting examples of such an loT device are a device which is or which is embedded in: a connected refrigerator or freezer, a TV, a connected lighting device, an electricity meter, a robot vacuum cleaner, a voice controlled smart speaker, a home security camera, a motion detector, a thermostat, a smoke detector, a door / window sensor, a flood / moisture sensor, an electrical door lock, a connected doorbell, an air conditioning system like a heat pump, an autonomous vehicle, a surveillance system, a weather monitoring device, a vehicle parking monitoring device, an electric vehicle charging station, a smart watch, a fitness tracker, a head-mounted display for Augmented Reality (AR) or Virtual Reality (VR), a wearable for tactile augmentation or sensory enhancement, a water sprinkler, an animal- or item-tracking device, a sensor for monitoring a plant or animal, an industrial robot, an Unmanned Aerial Vehicle (UAV), and any kind of medical device, like a heart rate monitor or a remote controlled surgical robot. A UE in the form of an loT device comprises circuitry and / or software in dependence of the intended application of the loT device in addition to other components as described in relation to the UE 800 shown in Figure 8.

[0146] As yet another specific example, in an loT scenario, a UE may represent a machine or other device that performs monitoring and / or measurements, and transmits the results of such monitoring and / or measurements to another UE and / or a network node. The UE may in this case be an M2M device, which may in a 3GPP context be referred to as an MTC device. As one particular example, the UE may implement the 3GPP NB-IoT standard. In other scenarios, a UE may represent a vehicle, such as a car, a bus, a truck, a ship and an airplane, or other equipment that is capable of monitoring and / or reporting on its operational status or other functions associated with its operation.

[0147] In practice, any number of UEs may be used together with respect to a single use case. For example, a first UE might be or be integrated in a drone and provide the drone’s speed information (obtained through a speed sensor) to a second UE that is a remote controller operating the drone. When the user makes changes from the remote controller, the first UE may adjust the throttle on the drone (e.g. by controlling an actuator) to increase or decrease the drone’s speed. The first and / or the second UE can also include more than one of the functionalities described above. For example, a UE might comprise the sensor and the actuator, and handle communication of data for both the speed sensor and the actuators.

[0148] Figure 9 shows a network node 900 in accordance with some embodiments. As used herein, network node refers to equipment capable, configured, arranged and / or operable to communicate directly or indirectly with a UE and / or with other network nodes or equipment, in a telecommunication network. Examples of network nodes include, but are not limited to, access points (aPs) (e.g., radio access points), base stations (BSs) (e.g., radio base stations, Node Bs, evolved Node Bs (eNBs) and NR NodeBs (gNBs)), O-RAN nodes or components of an O-RAN node (e.g., O-RU, O-DU, O-CU).

[0149] Base stations may be categorized based on the amount of coverage they provide (or, stated differently, their transmit power level) and so, depending on the provided amount of coverage, may be referred to as femto base stations, pico base stations, micro base stations, or macro base stations. A base station may be a relay node or a relay donor node controlling a relay. A network node may also include one or more (or all) parts of a distributed radio base station such as centralized digital units, distributed units (e.g., in an O- RAN access node) and / or remote radio units (RRUs), sometimes referred to as Remote Radio Heads (RRHs). Such remote radio units may or may not be integrated with an antenna as an antenna integrated radio. Parts of a distributed radio base station may also be referred to as nodes in a distributed antenna system (DAS).

[0150] Other examples of network nodes include multiple transmission point (multi-TRP) 5G access nodes, multi-standard radio (MSR) equipment such as MSR BSs, network controllers such as radio network controllers (RNCs) or base station controllers (BSCs), base transceiver stations (BTSs), transmission points, transmission nodes, multi- cell / multicast coordination entities (MCEs), Operation and Maintenance (O&M) nodes, Operations Support System (OSS) nodes, Self-Organizing Network (SON) nodes, positioning nodes (e.g., Evolved Serving Mobile Location Centers (E-SMLCs)), and / or Minimization of Drive Tests (MDTs).

[0151] The network node 900 includes a processing circuitry 902, a memory 904, a communication interface 906, and a power source 908. The network node 900 may be composed of multiple physically separate components (e.g., a NodeB component and a RNC component, or a BTS component and a BSC component, etc.), which may each have their own respective components. In certain scenarios in which the network node 900 comprises multiple separate components (e.g., BTS and BSC components), one or more of the separate components may be shared among several network nodes. For example, a single RNC may control multiple NodeBs. In such a scenario, each unique NodeB and RNC pair, may in some instances be considered a single separate network node. In some embodiments, the network node 900 may be configured to support multiple radio access technologies (RATs). In such embodiments, some components may be duplicated (e.g., separate memory 904 for different RATs) and some components may be reused (e.g., a same antenna 910 may be shared by different RATs). The network node 900 may also include multiple sets of the various illustrated components for different wireless technologies integrated into network node 900, for example GSM, WCDMA, LTE, NR, WiFi, Zigbee, Z-wave, LoRaWAN, Radio Frequency Identification (RFID) or Bluetooth wireless technologies. These wireless technologies may be integrated into the same or different chip or set of chips and other components within network node 900.

[0152] The processing circuitry 902 may comprise a combination of one or more of a microprocessor, controller, microcontroller, central processing unit, digital signal processor, application-specific integrated circuit, field programmable gate array, or any other suitable computing device, resource, or combination of hardware, software and / or encoded logic operable to provide, either alone or in conjunction with other network node 900 components, such as the memory 904, to provide network node 900 functionality.

[0153] In some embodiments, the processing circuitry 902 includes a system on a chip (SOC). In some embodiments, the processing circuitry 902 includes one or more of radiofrequency (RF) transceiver circuitry 912 and baseband processing circuitry 914. In some embodiments, the radio frequency (RF) transceiver circuitry 912 and the baseband processing circuitry 914 may be on separate chips (or sets of chips), boards, or units, such as radio units and digital units. In alternative embodiments, part or all of RF transceiver circuitry 912 and baseband processing circuitry 914 may be on the same chip or set of chips, boards, or units.

[0154] The memory 904 may comprise any form of volatile or non-volatile computer-readable memory including, without limitation, persistent storage, solid-state memory, remotely mounted memory, magnetic media, optical media, random access memory (RAM), read-only memory (ROM), mass storage media (for example, a hard disk), removable storage media (for example, a flash drive, a Compact Disk (CD) or a Digital Video Disk (DVD)), and / or any other volatile or non-volatile, non-transitory device-readable and / or computer-executable memory devices that store information, data, and / or instructions that may be used by the processing circuitry 902. The memory 904 may store any suitable instructions, data, or information, including a computer program, software, an application including one or more of logic, rules, code, tables, and / or other instructions capable of being executed by the processing circuitry 902 and utilized by the network node 900. The memory 904 may be used to store any calculations made by the processing circuitry 902 and / or any data received via the communication interface 906. In some embodiments, the processing circuitry 902 and memory 904 is integrated.

[0155] The communication interface 906 is used in wired or wireless communication of signaling and / or data between a network node, access network, and / or UE. As illustrated, the communication interface 906 comprises port(s) / terminal(s) 916 to send and receive data, for example to and from a network over a wired connection. The communication interface 906 also includes radio front-end circuitry 918 that may be coupled to, or in certain embodiments a part of, the antenna 910. Radio front-end circuitry 918 comprises filters 920 and amplifiers 922. The radio front-end circuitry 918 may be connected to an antenna 910 and processing circuitry 902. The radio front-end circuitry may be configured to condition signals communicated between antenna 910 and processing circuitry 902. The radio front-end circuitry 918 may receive digital data that is to be sent out to other network nodes or UEs via a wireless connection. The radio front-end circuitry 918 may convert the digital data into a radio signal having the appropriate channel and bandwidth parameters using a combination of filters 920 and / or amplifiers 922. The radio signal may then be transmitted via the antenna 910. Similarly, when receiving data, the antenna 910 may collect radio signals which are then converted into digital data by the radio front-end circuitry918. The digital data may be passed to the processing circuitry 902. In other embodiments, the communication interface may comprise different components and / or different combinations of components.

[0156] In certain alternative embodiments, the network node 900 does not include separate radio front-end circuitry 918, instead, the processing circuitry 902 includes radio front-end circuitry and is connected to the antenna 910. Similarly, in some embodiments, all or some of the RF transceiver circuitry 912 is part of the communication interface 906. In still other embodiments, the communication interface 906 includes one or more ports or terminals 916, the radio front-end circuitry 918, and the RF transceiver circuitry 912, as part of a radio unit (not shown), and the communication interface 906 communicates with the baseband processing circuitry 914, which is part of a digital unit (not shown).

[0157] The antenna 910 may include one or more antennas, or antenna arrays, configured to send and / or receive wireless signals. The antenna 910 may be coupled to the radio front-end circuitry 918 and may be any type of antenna capable of transmitting and receiving data and / or signals wirelessly. In certain embodiments, the antenna 910 is separate from the network node 900 and connectable to the network node 900 through an interface or port.

[0158] The antenna 910, communication interface 906, and / or the processing circuitry 902 may be configured to perform any receiving operations and / or certain obtaining operations described herein as being performed by the network node. Any information, data and / or signals may be received from a UE, another network node and / or any other network equipment. Similarly, the antenna 910, the communication interface 906, and / or the processing circuitry 902 may be configured to perform any transmitting operations described herein as being performed by the network node. Any information, data and / or signals may be transmitted to a UE, another network node and / or any other network equipment.

[0159] The power source 908 provides power to the various components of network node 900 in a form suitable for the respective components (e.g., at a voltage and current level needed for each respective component). The power source 908 may further comprise, or be coupled to, power management circuitry to supply the components of the network node 900 with power for performing the functionality described herein. For example, the network node 900 may be connectable to an external power source (e.g., the power grid, an electricity outlet) via an input circuitry or interface such as an electrical cable, whereby the external power source supplies power to power circuitry of the power source 908. As a further example, the power source 908 may comprise a source of power in the form of abattery or battery pack which is connected to, or integrated in, power circuitry. The battery may provide backup power should the external power source fail.

[0160] Embodiments of the network node 900 may include additional components beyond those shown in Figure 9 for providing certain aspects of the network node’s functionality, including any of the functionality described herein and / or any functionality necessary to support the subject matter described herein. For example, the network node 900 may include user interface equipment to allow input of information into the network node 900 and to allow output of information from the network node 900. This may allow a user to perform diagnostic, maintenance, repair, and other administrative functions for the network node 900.

[0161] Figure 10 is a block diagram of a host 1000, which may be an embodiment of the host 716 of Figure 7, in accordance with various aspects described herein. As used herein, the host 1000 may be or comprise various combinations hardware and / or software, including a standalone server, a blade server, a cloud-implemented server, a distributed server, a virtual machine, container, or processing resources in a server farm. The host 1000 may provide one or more services to one or more UEs.

[0162] The host 1000 includes processing circuitry 1002 that is operatively coupled via a bus 1004 to an input / output interface 1006, a network interface 1008, a power source 1010, and a memory 1012. Other components may be included in other embodiments. Features of these components may be substantially similar to those described with respect to the devices of previous figures, such as Figures 8 and 9, such that the descriptions thereof are generally applicable to the corresponding components of host 1000.

[0163] The memory 1012 may include one or more computer programs including one or more host application programs 1014 and data 1016, which may include user data, e.g., data generated by a UE for the host 1000 or data generated by the host 1000 for a UE. Embodiments of the host 1000 may utilize only a subset or all of the components shown. The host application programs 1014 may be implemented in a container-based architecture and may provide support for video codecs (e.g., Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), MPEG, VP9) and audio codecs (e.g., FLAC, Advanced Audio Coding (AAC), MPEG, G.711), including transcoding for multiple different classes, types, or implementations of UEs (e.g., handsets, desktop computers, wearable display systems, heads-up display systems). The host application programs 1014 may also provide for user authentication and licensing checks and may periodically report health, routes, and content availability to a central node, such as a devicein or on the edge of a core network. Accordingly, the host 1000 may select and / or indicate a different host for over-the-top services for a UE. The host application programs 1014 may support various protocols, such as the HTTP Live Streaming (HLS) protocol, Real-Time Messaging Protocol (RTMP), Real-Time Streaming Protocol (RTSP), Dynamic Adaptive Streaming over HTTP (MPEG-DASH), etc.

[0164] Figure 11 is a block diagram illustrating a virtualization environment 1100 in which functions implemented by some embodiments may be virtualized. In the present context, virtualizing means creating virtual versions of apparatuses or devices which may include virtualizing hardware platforms, storage devices and networking resources. As used herein, virtualization can be applied to any device described herein, or components thereof, and relates to an implementation in which at least a portion of the functionality is implemented as one or more virtual components. Some or all of the functions described herein may be implemented as virtual components executed by one or more virtual machines (VMs) implemented in one or more virtual environments 1100 hosted by one or more of hardware nodes, such as a hardware computing device that operates as a network node, UE, core network node, or host. Further, in embodiments in which the virtual node does not require radio connectivity (e.g., a core network node or host), then the node may be entirely virtualized. In some embodiments, the virtualization environment 1100 includes components defined by the O-RAN Alliance, such as an O-Cloud environment orchestrated by a Service Management and Orchestration Framework via an O-2 interface.

[0165] Applications 1102 (which may alternatively be called software instances, virtual appliances, network functions, virtual nodes, virtual network functions, etc.) are run in the virtualization environment Q400 to implement some of the features, functions, and / or benefits of some of the embodiments disclosed herein.

[0166] Hardware 1104 includes processing circuitry, memory that stores software and / or instructions executable by hardware processing circuitry, and / or other hardware devices as described herein, such as a network interface, input / output interface, and so forth. Software may be executed by the processing circuitry to instantiate one or more virtualization layers 1106 (also referred to as hypervisors or virtual machine monitors (VMMs)), provide VMs 1108a and 1108b (one or more of which may be generally referred to as VMs 1108), and / or perform any of the functions, features and / or benefits described in relation with some embodiments described herein. The virtualization layer 1106 may present a virtual operating platform that appears like networking hardware to the VMs 1108.

[0167] The VMs 1108 comprise virtual processing, virtual memory, virtual networking or interface and virtual storage, and may be run by a corresponding virtualization layer 1106. Different embodiments of the instance of a virtual appliance 1102 may be implemented on one or more of VMs 1108, and the implementations may be made in different ways. Virtualization of the hardware is in some contexts referred to as network function virtualization (NFV). NFV may be used to consolidate many network equipment types onto industry standard high volume server hardware, physical switches, and physical storage, which can be located in data centers, and customer premise equipment.

[0168] In the context of NFV, a VM 1108 may be a software implementation of a physical machine that runs programs as if they were executing on a physical, non-virtualized machine. Each of the VMs 1108, and that part of hardware 1104 that executes that VM, be it hardware dedicated to that VM and / or hardware shared by that VM with others of the VMs, forms separate virtual network elements. Still in the context of NFV, a virtual network function is responsible for handling specific network functions that run in one or more VMs 1108 on top of the hardware 1104 and corresponds to the application 1102.

[0169] Hardware 1104 may be implemented in a standalone network node with generic or specific components. Hardware 1104 may implement some functions via virtualization. Alternatively, hardware 1104 may be part of a larger cluster of hardware (e.g. such as in a data center or CPE) where many hardware nodes work together and are managed via management and orchestration 1110, which, among others, oversees lifecycle management of applications 1102. In some embodiments, hardware 1104 is coupled to one or more radio units that each include one or more transmitters and one or more receivers that may be coupled to one or more antennas. Radio units may communicate directly with other hardware nodes via one or more appropriate network interfaces and may be used in combination with the virtual components to provide a virtual node with radio capabilities, such as a radio access node or a base station. In some embodiments, some signaling can be provided with the use of a control system 1112 which may alternatively be used for communication between hardware nodes and radio units.

[0170] Figure 12 shows a communication diagram of a host 1202 communicating via a network node 1204 with a UE 1206 over a partially wireless connection in accordance with some embodiments. Example implementations, in accordance with various embodiments, of the UE (such as a UE 712a of Figure 7 and / or UE 800 of Figure 8), network node (such as network node 710a of Figure 7 and / or network node 900 of Figure 9), and host(such as host 716 of Figure 7 and / or host 1000 of Figure 10) discussed in the preceding paragraphs will now be described with reference to Figure 12.

[0171] Like host 1000, embodiments of host 1202 include hardware, such as a communication interface, processing circuitry, and memory. The host 1202 also includes software, which is stored in or accessible by the host 1202 and executable by the processing circuitry. The software includes a host application that may be operable to provide a service to a remote user, such as the UE 1206 connecting via an over-the-top (OTT) connection 1250 extending between the UE 1206 and host 1202. In providing the service to the remote user, a host application may provide user data which is transmitted using the OTT connection 1250.

[0172] The network node 1204 includes hardware enabling it to communicate with the host 1202 and UE 1206. The connection 1260 may be direct or pass through a core network (like core network 706 of Figure 7) and / or one or more other intermediate networks, such as one or more public, private, or hosted networks. For example, an intermediate network may be a backbone network or the Internet.

[0173] The UE 1206 includes hardware and software, which is stored in or accessible by UE 1206 and executable by the UE’s processing circuitry. The software includes a client application, such as a web browser or operator-specific “app” that may be operable to provide a service to a human or non-human user via UE 1206 with the support of the host 1202. In the host 1202, an executing host application may communicate with the executing client application via the OTT connection 1250 terminating at the UE 1206 and host 1202. In providing the service to the user, the U”s client application may receive request data from the hos”s host application and provide user data in response to the request data. The OTT connection 1250 may transfer both the request data and the user data. The U”s client application may interact with the user to generate the user data that it provides to the host application through the OTT connection 1250.

[0174] The OTT connection 1250 may extend via a connection 1260 between the host 1202 and the network node 1204 and via a wireless connection 1270 between the network node 1204 and the UE 1206 to provide the connection between the host 1202 and the UE 1206. The connection 1260 and wireless connection 1270, over which the OTT connection 1250 may be provided, have been drawn abstractly to illustrate the communication between the host 1202 and the UE 1206 via the network node 1204, without explicit reference to any intermediary devices and the precise routing of messages via these devices.

[0175] As an example of transmitting data via the OTT connection 1250, in step 1208, the host 1202 provides user data, which may be performed by executing a host application. In some embodiments, the user data is associated with a particular human user interacting with the UE 1206. In other embodiments, the user data is associated with a UE 1206 that shares data with the host 1202 without explicit human interaction. In step 1210, the host 1202 initiates a transmission carrying the user data towards the UE 1206. The host 1202 may initiate the transmission responsive to a request transmitted by the UE 1206. The request may be caused by human interaction with the UE 1206 or by operation of the client application executing on the UE 1206. The transmission may pass via the network node 1204, in accordance with the teachings of the embodiments described throughout this disclosure. Accordingly, in step 1212, the network node 1204 transmits to the UE 1206 the user data that was carried in the transmission that the host 1202 initiated, in accordance with the teachings of the embodiments described throughout this disclosure. In step 1214, the UE 1206 receives the user data carried in the transmission, which may be performed by a client application executed on the UE 1206 associated with the host application executed by the host 1202.

[0176] In some examples, the UE 1206 executes a client application which provides user data to the host 1202. The user data may be provided in reaction or response to the data received from the host 1202. Accordingly, in step 1216, the UE 1206 may provide user data, which may be performed by executing the client application. In providing the user data, the client application may further consider user input received from the user via an input / output interface of the UE 1206. Regardless of the specific manner in which the user data was provided, the UE 1206 initiates, in step 1218, transmission of the user data towards the host 1202 via the network node 1204. In step 1220, in accordance with the teachings of the embodiments described throughout this disclosure, the network node 1204 receives user data from the UE 1206 and initiates transmission of the received user data towards the host 1202. In step 1222, the host 1202 receives the user data carried in the transmission initiated by the UE 1206.

[0177] One or more of the various embodiments improve the performance of OTT services provided to the UE 1206 using the OTT connection 1250, in which the wireless connection 1270 forms the last segment.

[0178] In an example scenario, factory status information may be collected and analyzed by the host 1202. As another example, the host 1202 may process audio and video data which may have been retrieved from a UE for use in creating maps. As another example, the host 1202 may collect and analyze real-time data to assist in controlling vehiclecongestion (e.g., controlling traffic lights). As another example, the host 1202 may store surveillance video uploaded by a UE. As another example, the host 1202 may store or control access to media content such as video, audio, VR or AR which it can broadcast, multicast or unicast to UEs. As other examples, the host 1202 may be used for energy pricing, remote control of non-time critical electrical load to balance power generation needs, location services, presentation services (such as compiling diagrams etc. from data collected from remote devices), or any other function of collecting, retrieving, storing, analyzing and / or transmitting data.

[0179] In some examples, a measurement procedure may be provided for the purpose of monitoring data rate, latency and other factors on which the one or more embodiments improve. There may further be an optional network functionality for reconfiguring the OTT connection 1250 between the host 1202 and UE 1206, in response to variations in the measurement results. The measurement procedure and / or the network functionality for reconfiguring the OTT connection may be implemented in software and hardware of the host 1202 and / or UE 1206. In some embodiments, sensors (not shown) may be deployed in or in association with other devices through which the OTT connection 1250 passes; the sensors may participate in the measurement procedure by supplying values of the monitored quantities exemplified above, or supplying values of other physical quantities from which software may compute or estimate the monitored quantities. The reconfiguring of the OTT connection 1250 may include message format, retransmission settings, preferred routing etc.; the reconfiguring need not directly alter the operation of the network node 1204. Such procedures and functionalities may be known and practiced in the art. In certain embodiments, measurements may involve proprietary UE signaling that facilitates measurements of throughput, propagation times, latency and the like, by the host 1202. The measurements may be implemented in that software causes messages to be transmitted, in particular empty or ‘dummy’ messages, using the OTT connection 1250 while monitoring propagation times, errors, etc.

[0180] Although the computing devices described herein (e.g., UEs, network nodes, hosts) may include the illustrated combination of hardware components, other embodiments may comprise computing devices with different combinations of components. It is to be understood that these computing devices may comprise any suitable combination of hardware and / or software needed to perform the tasks, features, functions and methods disclosed herein. Determining, calculating, obtaining or similar operations described herein may be performed by processing circuitry, which may process information by, for example,converting the obtained information into other information, comparing the obtained information or converted information to information stored in the network node, and / or performing one or more operations based on the obtained information or converted information, and as a result of said processing making a determination. Moreover, while components are depicted as single boxes located within a larger box, or nested within multiple boxes, in practice, computing devices may comprise multiple different physical components that make up a single illustrated component, and functionality may be partitioned between separate components. For example, a communication interface may be configured to include any of the components described herein, and / or the functionality of the components may be partitioned between the processing circuitry and the communication interface. In another example, non-computationally intensive functions of any of such components may be implemented in software or firmware and computationally intensive functions may be implemented in hardware.

[0181] In certain embodiments, some or all of the functionality described herein may be provided by processing circuitry executing instructions stored on in memory, which in certain embodiments may be a computer program product in the form of a non-transitory computer-readable storage medium. In alternative embodiments, some or all of the functionality may be provided by the processing circuitry without executing instructions stored on a separate or discrete device -readable storage medium, such as in a hard-wired manner. In any of those particular embodiments, whether executing instructions stored on a non-transitory computer-readable storage medium or not, the processing circuitry can be configured to perform the described functionality. The benefits provided by such functionality are not limited to the processing circuitry alone or to other components of the computing device, but are enjoyed by the computing device as a whole, and / or by end users and a wireless network generally.Embodiments:Embodiment 1. A method of operating a supervisor agent for collaboration with a plurality of trained agents in a multi-agent reinforcement learning, MARL system, wherein the plurality of agents execute respective local policies on an environment, the method comprising: receiving (402), at a fusion layer, an embedding (mi) for each of the respective agents, the embedding representing a performance of the respective agent; receiving (404), at the fusion layer, one or more global goals for the MARL system;receiving (406), at the fusion layer, a penalty context (pt+ ) that is based on one or more utility function values associated with the environment; generating (408), at the fusion layer, a context representation (ct+i) based on the embeddings, the one or more global goals and the penalty context; applying (410) a goal policy to the context representation to generate a plurality of sub-goals (g'v / + / ), each of the sub-goals being associated with a respective one of the agents; and providing (412) the sub-goals to respective ones of the agents to be used as goals for training the respective local policies of the agents.Embodiment 2. The method of Embodiment 1, wherein the embedding for each agent is generated based on an encoded capability vector for the agent and based on a state-action- goal tuple associated with the agent.Embodiment 3. The method of Embodiment 2, wherein the encoded capability vector of the agent comprises a set of goals for the agent and a set of probabilities associated with respective ones of the goals, of the agent achieving the goal within a specified time period.Embodiment 4. The method of Embodiment 3, wherein the embedding for each agent is created by merging the encoded capability vector and the state- action-goal tuple of the agent.Embodiment 5. The method of any previous Embodiment, wherein the penalty context is generated based on one or more utility function values.Embodiment 6. The method of Embodiment 5, wherein the utility function values are normalized by maximum values of the utility function values.Embodiment 7. The method of Embodiments 5 or 6, wherein the penalty context is generated as the output of a neural network that takes the one or more utility function values as inputs.Embodiment 8. The method of any of Embodiments 5 to 7, wherein the penalty context is generated by a utility function network, wherein an output of the neural network isgeneric to a form of utility function values.Embodiment 9. The method of Embodiment 8, wherein the output of the neural network is generic as to priority of utility functions used to generate the utility function values.Embodiment 10. The method of any previous Embodiment, wherein the fusion layer comprises a fully connected neural network.Embodiment 11. The method of any previous Embodiment, wherein the context representation is used to train the goal policy.Embodiment 12. A supervisor agent adapted to perform operations according to any of Embodiments 1 to 11.Embodiment 13. A supervisor agent comprising: a processing circuitry; and a memory coupled to the processing circuitry, wherein the memory comprises computer program instructions that, when executed by the processing circuitry, cause the supervisor agent to perform operations according to any of Embodiments 1 to 11.

Claims

Claims:

1. A method of operating a supervisor agent (110) for collaboration with a plurality of trained agents (112, 114) in a multi-agent reinforcement learning, MARL system, wherein the plurality of agents execute respective local policies on an environment, the method comprising: receiving (402), at a fusion layer (130), an embedding (mi) for each of the respective agents, the embedding representing a performance of the respective agent; receiving (404), at the fusion layer, one or more global goals for the MARL system; receiving (406), at the fusion layer, a penalty context (pt+i) that is based on one or more utility function values associated with the environment; generating (408), at the fusion layer, a context representation (ct+i) based on the embeddings, the one or more global goals and the penalty context; applying (410) a goal policy to the context representation to generate a plurality of sub-goals (gNt+i), each of the sub-goals being associated with a respective one of the agents; and providing (412) the sub-goals to respective ones of the agents to be used as goals for training the respective local policies of the agents.

2. The method of Claim 1, wherein the embedding for each agent is generated based on an encoded capability vector for the agent and based on a state-action-goal tuple associated with the agent.

3. The method of Claim 2, wherein the encoded capability vector of the agent comprises a set of goals for the agent and a set of probabilities, associated with respective ones of the goals, of the agent achieving the goal within a specified time period.

4. The method of Claim 3, wherein the embedding for each agent is created by merging the encoded capability vector and the state-action-goal tuple of the agent.

5. The method of any previous Claim, wherein the penalty context is generated based on one or more utility function values.

6. The method of Claim 5, wherein the utility function values are normalized by maximum values of the utility function values.

7. The method of Claims 5 or 6, wherein the penalty context is generated as the output of a neural network that takes the one or more utility function values as inputs.

8. The method of any of Claims 5 to 7, wherein the penalty context is generated by a utility function network, wherein an output of the neural network is generic to a form of utility function values.

9. The method of Claim 8, wherein the output of the neural network is generic as to priority of utility functions used to generate the utility function values.

10. The method of any previous Claim, wherein the fusion layer comprises a fully connected neural network.

11. The method of any previous Claim, wherein the context representation is used to train the goal policy.

12. The method of any previous Claim, wherein the MARL system is generic with respect to a form of the utility function.

13. The method of any previous Claim, wherein the MARL system is generic with respect to changes in the environment.

14. A supervisor agent (110) for collaboration with a plurality of trained agents (112, 114) in a multi-agent reinforcement learning, MARL system, wherein the plurality of agents execute respective local policies on an environment, wherein the supervisor agent is adapted to perform operations comprising: receiving (402), at a fusion layer (130), an embedding (mi) for each of the respective agents, the embedding representing a performance of the respective agent; receiving (404), at the fusion layer, one or more global goals for the MARL system; receiving (406), at the fusion layer, a penalty context (pt+i) that is based on one or more utility function values associated with the environment;generating (408), at the fusion layer, a context representation (ct+i) based on the embeddings, the one or more global goals and the penalty context; applying (410) a goal policy to the context representation to generate a plurality of sub-goals (gNt+i), each of the sub-goals being associated with a respective one of the agents; and providing (412) the sub-goals to respective ones of the agents to be used as goals for training the respective local policies of the agents.

15. The supervisor agent of Claim 14, wherein the supervisor agent is adapted to perform operations according to any of Claims 2 to 13.

16. A supervisor agent (110) for collaboration with a plurality of trained agents (112, 114) in a multi-agent reinforcement learning, MARL system, wherein the plurality of agents execute respective local policies on an environment, the supervisor agent comprising: a processing circuitry (312); and a memory (314) coupled to the processing circuitry, wherein the memory comprises computer program instructions that, when executed by the processing circuitry, cause the supervisor agent to perform operations comprising: receiving (402), at a fusion layer (130), an embedding (mi) for each of the respective agents, the embedding representing a performance of the respective agent; receiving (404), at the fusion layer, one or more global goals for the MARL system; receiving (406), at the fusion layer, a penalty context (pt+i) that is based on one or more utility function values associated with the environment; generating (408), at the fusion layer, a context representation (ct+i) based on the embeddings, the one or more global goals and the penalty context; applying (410) a goal policy to the context representation to generate a plurality of sub-goals (gNt+i), each of the sub-goals being associated with a respective one of the agents; and providing (412) the sub-goals to respective ones of the agents to be used as goals for training the respective local policies of the agents.

17. The supervisor agent of Claim 16, wherein the computer program instructions further cause the supervisor agent to perform operations according to any of Claims 2 to 13.

Citation Information

Patent Citations

  • Method and apparatus for multiple reinforcement learning agents in a shared environment

    US20230214725A1