Model exploration coordination
The coordination of model exploration across network nodes through a cooperation negotiation process addresses the challenges of resource allocation and decision-making in wireless communication systems, resulting in improved performance and efficiency.
Patent Information
- Application Number
- PCT/EP2024/073873
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-08-27
- Publication Date
- 2025-05-08
AI Technical Summary
Current wireless communication systems face challenges in efficiently coordinating model exploration across multiple network nodes for optimal resource allocation and decision-making, particularly in the context of reinforcement learning and AI/ML integration.
The proposed solution involves an apparatus and method for coordinating model exploration through a cooperation negotiation process, where a coordinating entity identifies network nodes for time domain resource allocation, determines exploration order policies, and synchronizes operations to optimize sequential model exploration based on outcomes from parallel model exploration.
This approach enables efficient resource allocation and improved decision-making by optimizing time domain patterns for sequential model exploration, thereby enhancing the overall performance and effectiveness of wireless communication systems.
Smart Images

Figure EP2024073873_08052025_PF_FP_ABST
Abstract
Description
[0001] Title: Model Exploration Coordination
[0002] Field of the Disclosure
[0003] Various example embodiments relate to communications.
[0004] Background
[0005] 3GPP is the world's leading standardization specification organization for mobile networks. Its ongoing work on the 5G- Advanced study of AI / ML in air interface offers the first glimpse of AI / ML that is embedded across device, radio and RAN. It will create the foundation for AI / ML features for all future releases to come, including 6G. The study focuses on a general framework, enabling a family of use cases utilizing AI / ML techniques. These include enhancements in channel information feedback, beam management and accuracy of device positioning. In wireless communication, 5G-Advanced paves the way, and 6G is seen to be the first truly Al-native system.
[0006] Reinforcement learning is a feedback-based machine learning approach in which agent learns by performing actions and by outcomes of the actions. In reinforcement learning, the agent is not aware of the different states, the actions available in all the states, the associated rewards and transition to resulting states. The agent learns more and more about it by interacting with the environment.
[0007] In exploration techniques, target is to gather more information about each action instead of getting more rewards for the actions for making the best overall decision. Exploration is a concept of the agent improving its knowledge about each action with at a target of long term benefit.
[0008] Brief Description Various example embodiments of the di sclosure are set out by the independent claims .
[0009] Some examples relate to an apparatus compri sing at least one proces sor , and at least one memory storing instructions that , when executed by the at least one proces sor, cause a coordinating entity to : carry out a cooperation negotiation proces s , wherein the cooperation negotiation proces s compri ses identifying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes , and determining an exploration order policy for the at least one sequential model exploration ; receive a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes ; determine a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes , wherein the determination i s based on the received report , the time domain resource allocation and the exploration order policy; transmit , to the plurality of network nodes , information on the time domain pattern, and receive , from the plurality of network nodes , a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern .
[0010] In some examples , the cooperation negotiation proces s may further compri se at least one of : a ) exchanging capability information, or b ) determining an input measure for the at least one parallel model exploration, or c ) determining an input measure for the at least one sequential model exploration .
[0011] In some examples , the time domain resource allocation may further compri se determining the time domain resource allocation as one or more time units . In some examples, at least one of a) the report of the at least one outcome of the at least one parallel model exploration, or b) the report of the at least one outcome of the at least one sequential model exploration may further comprise results of a training of the model.
[0012] In some examples, the report of the at least one outcome of the at least one parallel model exploration may further comprise at least one of: a) an indication of one or more reward values, or b) one or more performance indicators.
[0013] In some examples, the instructions, when executed by the at least one processor, may cause the coordinating entity to evaluate the received at least one outcome for determining at least one of: a) whether the at least one sequential model exploration is to be repeated, or b) whether the cooperation negotiation process is to be repeated, or c) whether one or more additional sequential model explorations are to be performed.
[0014] In some examples, the cooperation negotiation process may further comprise at least one of: negotiating on transferring the time domain resource allocation, or determining the time domain pattern to one of the plurality of network nodes.
[0015] In some examples, the coordinating entity may be a coordinating entity for a wireless, e.g. , cellular, communication system.
[0016] In some examples, the communication system may adhere to and / or may be based on some accepted (and / or planned) specification, e.g. , standard, such as, e.g. , 3G, 4G, 5G, 6G, or some other wireless communication standard.
[0017] In some examples, the coordinating entity may be located in at least one of: a core network, or an edge cloud, or a radio access network . Some examples relate to an apparatus compri sing means for causing a coordinating entity to : carry out a cooperation negotiation proces s , wherein the cooperation negotiation process comprises identi fying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes , and determining an exploration order policy for the at least one sequential model exploration ; receive a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes ; determine a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes , wherein the determination i s based on the received report , the time domain resource allocation and the exploration order policy; transmit , to the plurality of network nodes , information on the time domain pattern, and receive , from the plurality of network nodes , a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern .
[0018] In some examples , the means for causing the coordinating entity to perform at least one of the above-mentioned aspects may compri se at least one processor , and at least one memory storing instructions that , when executed by the at least one proces sor , cause the coordinating entity to perform at least one of the above-mentioned aspects . In some examples , the means for causing the coordinating entity to perform at least one of the above- mentioned aspects may, e . g . , compri se circuitry configured to perform at least one of the aforementioned aspects .
[0019] Some examples relate to a coordinating entity compri sing the apparatus according to the embodiments . In some examples , the coordinating entity may be a machine learning orchestrator , MLO, e . g . , for a wireles s communication system. Some examples relate to a method, comprising : carrying out , by a coordinating entity, a cooperation negotiation proces s , wherein the cooperation negotiation proces s compri ses identifying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes , and determining an exploration order policy for the at least one sequential model exploration; receiving a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes ; determining a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes , wherein the determination i s based on the received report , the time domain resource allocation and the exploration order policy; transmitting, to the plurality of network nodes , information on the time domain pattern, and receiving, from the plurality of network nodes , a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern .
[0020] Some examples relate to an apparatus comprising at least one proces sor , and at least one memory storing instructions that , when executed by the at least one proces sor, cause a coordinating network node to carry out a cooperation negotiation proces s , wherein the cooperation negotiation proces s compri ses identi fying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes , and determining an exploration order policy for the at least one sequential model exploration ; carry out a synchronization to a common reference time with the plurality of network nodes ; receive a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes ; determine a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes, wherein the determination is based on the received report, the time domain resource allocation and the exploration order policy; transmit to a coordinating entity, a request for grant for the determined time domain pattern; in response to receiving the grant, transmit, to the plurality of network nodes, information on the time domain pattern, and receive, from the plurality of network nodes, a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern.
[0021] In some examples, the cooperation negotiation process may further comprise at least one of: a) exchanging capability information or b) determining an input measure for the at least one parallel model exploration, or c) determining an input measure for the at least one sequential model exploration.
[0022] In some examples, the time domain resource allocation may further comprise determining the time domain resource allocation as one or more time units.
[0023] In some examples, at least one of a) the report of the at least one outcome of the at least one parallel model exploration, or b) the report of the at least one outcome of the at least one sequential model exploration may further comprise results of a training of the model.
[0024] In some examples, the report of the at least one outcome of the at least one parallel model exploration may further comprise at least one of: a) an indication of one or more reward values, or b) one or more performance indicators.
[0025] In some examples, the instructions, when executed by the at least one processor, may cause the coordinating network node to evaluate the received at least one outcome for determining at least one of: a) whether the at least one sequential model exploration is to be repeated, or b) whether the cooperation negotiation process is to be repeated, or c) whether one or more additional sequential model explorations are to be performed.
[0026] In some examples, the cooperation negotiation process may further comprise at least one of: negotiating on transferring the time domain resource allocation, or determining the time domain pattern to the coordinating entity.
[0027] Some examples relate to a coordinating network node comprising an apparatus according to the embodiments.
[0028] In some examples, the coordinating network node may be a network device for a wireless, e.g. , cellular, communication system, such as, for example, a base station, e.g. , gNB. In some examples, the apparatus according to the embodiments or its functionality, respectively, may, e.g. , be integrated into at least one component of the network device.
[0029] In some examples, the network device may adhere to and / or may be based on some accepted (and / or planned) specification, e.g. , standard, such as, e.g. , 3G, 4G, 5G, 6G, or some other wireless communication standard.
[0030] In some examples, the coordinating network node may be, in a distributed case, a central unit, CU, of a network node, e.g. , gNB, and the plurality of network nodes may comprise or represent, respectively, distributed units, DU, of the network node, e.g. , gNB, with radio unit accessibility.
[0031] Some examples relate to an apparatus comprising means for causing a coordinating network node to carry out a cooperation negotiation process, wherein the cooperation negotiation process comprises identifying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes , and determining an exploration order policy for the at least one sequential model exploration; carry out a synchroni zation to a common reference time with the plurality of network nodes ; receive a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes ; determine a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes , wherein the determination is based on the received report , the time domain resource allocation and the exploration order policy; transmit to a coordinating entity, a request for grant for the determined time domain pattern ; in response to receiving the grant , transmit , to the plurality of network nodes , information on the time domain pattern, and receive , from the plurality of network nodes , a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern .
[0032] In some examples , the means for causing the coordinating network node to perform at least one of the above-mentioned aspects may compri se at least one processor , and at least one memory storing instructions that , when executed by the at least one proces sor , cause the coordinating network node to perform at least one of the above-mentioned aspects . In some examples , the means for causing the coordinating network node to perform at least one of the above-mentioned aspects may, e . g . , comprise circuitry configured to perform at least one of the aforementioned aspects .
[0033] Some examples relate to a method, compri sing : carrying out , by a coordinating network node , a cooperation negotiation process , wherein the cooperation negotiation proces s compri ses identi fying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes, and determining an exploration order policy for the at least one sequential model exploration; carrying out a synchronization to a common reference time with the plurality of network nodes; receiving a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes; determining a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes, wherein the determination is based on the received report, the time domain resource allocation and the exploration order policy; transmitting to a coordinating entity, a request for grant for the determined time domain pattern; in response to receiving the grant, transmitting, to the plurality of network nodes, information on the time domain pattern, and receive, from the plurality of network nodes, a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern .
[0034] Some examples relate to a communication system, e.g. , a wireless, e.g. , cellular, communication system, comprising at least one of: a) at least one apparatus according to the embodiments, or b) a coordinating entity according to the embodiments, or c) a coordinating network node according to the embodiments.
[0035] Some examples relate to a computer program comprising instructions which, when executed by an apparatus, cause the apparatus to perform at least some aspects of the method according to the embodiments.
[0036] Some examples relate to a data carrier signal carrying and / or characterizing the computer program according to the embodiments.
[0037] Brief Description of the Figures
[0038] Fig. 1A schematically depicts a simplified block diagram according to some examples, Fig. IB schematically depicts a simplified block diagram according to some examples,
[0039] Fig. 2 schematically depicts a simplified block diagram according to some examples,
[0040] Fig. 3 schematically depicts a simplified flow chart according to some examples,
[0041] Fig. 4 schematically depicts a simplified flow chart according to some examples,
[0042] Fig. 5A schematically depicts a simplified diagram according to some examples,
[0043] Fig. 5B schematically depicts a simplified diagram according to some examples,
[0044] Fig. 6 schematically depicts a simplified flow chart according to some examples,
[0045] Fig. 7 schematically depicts a simplified flow chart according to some examples,
[0046] Fig. 8A schematically depicts a simplified block diagram according to some examples,
[0047] Fig. 8B schematically depicts a simplified block diagram according to some examples,
[0048] Fig. 9 schematically depicts a simplified flow chart according to some examples,
[0049] Fig. 10 schematically depicts a simplified flow chart according to some examples,
[0050] Fig. 11 schematically depicts a simplified flow chart according to some examples, Fig. 12 schematically depicts a simplified signaling diagram according to some examples,
[0051] Fig. 13A schematically depicts a simplified block diagram according to some examples,
[0052] Fig. 13B schematically depicts a simplified block diagram according to some examples,
[0053] Fig. 14 schematically depicts a simplified signaling diagram according to some examples,
[0054] Fig. 15 schematically depicts a simplified diagram according to some examples,
[0055] Fig. 16 schematically depicts a simplified block diagram according to some examples,
[0056] Fig. 17 schematically depicts a simplified block diagram according to some examples,
[0057] Fig. 18A schematically depicts aspects of reinforcement learning according to some examples,
[0058] Fig. 18B schematically depicts aspects of reinforcement learning according to some examples.
[0059] Description of some Example Embodiments
[0060] Some examples, see Fig. 1A, 2, 3, relate to an apparatus 100 comprising at least one processor 102, and at least one memory 104 storing instructions 106 that, when executed by the at least one processor 102, cause a coordinating entity 10 to: carry out 300 a cooperation negotiation process COOP-NEGOT, wherein the cooperation negotiation process comprises identifying 300a (Fig. 4) a plurality 20 (Fig. 2) of network nodes for time domain resource allocation TD-RA for at least one parallel model exploration PAR-EXPLOR and for at least one sequential model exploration SEQ-EXPLOR to be carried out by the plurality 20 of network nodes, and determining 300b an exploration order policy ORD-POL for the at least one sequential model exploration SEQ- EXPLOR; receive 302 (Fig. 3) a report REP-PAR of at least one outcome OUT-PAR-EXPLORE (Fig. 5A) of the at least one parallel model exploration carried out by the plurality of network nodes; determine 304 (Fig. 3) a time domain pattern TDP-SEQ for the at least one sequential model exploration SEQ-EXPLOR to be carried out by the plurality of network nodes, wherein the determination 304 is based on the received report REP-PAR, the time domain resource allocation TD-RA and the exploration order policy ORD- POL; transmit 306, to the plurality 20 of network nodes, information I-TDP-SEQ on the time domain pattern TDP-SEQ, and receive 308, from the plurality 20 of network nodes, a report REP-SEQ of at least one outcome OUT-SEQ-EXPLORE (Fig. 5B) of the at least one sequential model exploration SEQ-EXPLOR carried out according to the time domain pattern TDP-SEQ.
[0061] In some examples, Fig. 2, the parallel model exploration PAR- EXPLOR and the sequential model exploration SEQ-EXPLOR may be associated with at least one machine learning, ML, model, which may, e.g. , be based on reinforcement learning, RL, e.g. , involving a plurality of RL agents RLA-1, RLA-2, ....
[0062] In some examples, Fig. 2, at least one of the RL agents RLA-1, RLA-2, ... may be associated with at least one network node 20a, 20b, 20c of the plurality 20 of network nodes. In some examples, at least one of the RL agents RLA-1, RLA-2, ... may also be associated with at least one terminal device, e.g. , user equipment, 30, that may, e.g. , be associated with at least one network node 20a, 20b, 20c of the plurality 20 of network nodes.
[0063] In some examples, Fig. 2, the principle according to the embodiments may enable to coordinate the operation of the RL agents RLA-1, RLA-2, .... In some examples, performing the at least one sequential model exploration SEQ-EXPLOR based on the time domain pattern TDP-SEQ may, e.g. , comprise the first RL agent RLA-1 performing exploration in a first time resources as, e.g. , characterized by the time domain pattern TDP-SEQ and / or, for example, the information I-TDP-SEQ, wherein, for example, other RL agent (s) RLA-2, ... do not perform exploration in the first time resources, e.g. , to ensure a sequential model exploration. In some examples, the further RL agent RLA-2 may, e.g. , perform exploration in second time resources which are different from the first time resources.
[0064] In some examples, reinforcement learning (RL) may be considered as an area of machine learning, where an intelligent agent, e.g. , RL agent, e.g. , the first RL agent RLA-1 (and / or the at least one further RL agent RLA-2) of Fig. 2, may take one or more actions in an environment (e.g. , associated with a communication system 1000 comprising the entities 10, 20, 30) in order to maximize its objective. Some elements associated with a Markov Decision Process (MDP) framework in reinforcement learning may be S, A, p, R, y) , where S is a finite set of environment states, A is a finite set of the RL agent's action, p S X A X S [0,1] is a policy or state transition function, R: S X A X S J? is a reward function, and y is a discount factor.
[0065] Fig. 18A depicts aspects of a Markov Decision Process according to some examples. Block RLA symbolizes an RL agent, e.g. , at least similar to the RL agents RLA-1, RLA-2 of Fig. 2, block ENV symbolizes an environment. Arrow S_i symbolizes an i-th state, arrow A_i symbolizes an i-th action, and arrow R_i symbolizes an i-th reward associated with the i-th action. Correspondingly, arrow S_i+1 symbolizes an (i+l) -th state, and arrow R_i+1 symbolizes an (i+l) -th reward.
[0066] In some examples, a setup of a multi-agent configuration involving a plurality of RL agents, e.g. , at least similar to the RL agents RLA-1, RLA-2, ... of Fig. 2, may, e.g. , be defined as S, A, p, R, T) , where S = S±X S2X ... X Sj is a set of, e.g. all, states available to the RL agents, A = A±X A2X ... X Aj is a set of, e.g. all, actions available to the RL agents, p S X A X S [0,1], is a transition function, R = R X R2X ... X Rj is a set of, e.g. all, rewards by the RL agents, Rj'. S X A X S 5?, is a reward of RL agent j, and r = Yi Xy2X ...XY7is a set of, e.g. all, discount factors of the RL agents.
[0067] Fig. 18B depicts aspects of an example multi-agent MDP configuration involving k many RL agents RLA-1' , ... , RLA-k associated with an environment ENV' wherein the arrows la, lb indicate that the RL agents RL-1A' , . . . , RLA-k may mutually affect their respective policies, wherein the arrows 2a, 2b symbolize respective states and rewards, and wherein the arrows 3a, 3b symbolize respective actions of the RL agents RLA-1', ... ,
[0068] RLA-k.
[0069] In some examples, Fig. 2, at least some RL agents RLA-1, RLA-2 associated with the communication system 1000 may, at least temporarily, form a multi-agent configuration as depicted by Fig. 18B, and / or may, at least temporarily, perform sequential learning or model exploration.
[0070] In some examples, a priority or sequence indicating which RL agent RLA-1, RLA-2 should perform exploration at which time, e.g. , for the sequential model exploration SEQ-EXPLOR, may be indicated by the exploration order policy ORD-POL, as, e.g. , obtained by block 300b of Fig. 4.
[0071] In some examples, Fig. 2, at least some RL agents RLA-1, RLA-2 associated with the communication system 1000 can, e.g. , be used for at least one of the following aspects: a) channel state information, CSI, feedback enhancement (e.g. , reducing an overhead, or improving an accuracy / prediction) , or b) beam management (e.g. , beam prediction in time domain and / or spatial domain, e.g. , for overhead and / or latency reduction, or beam selection accuracy improvement) , or c) enhancing positioning accuracy, e.g. , for, for example heavy, non-line of sight, NLOS, conditions .
[0072] In some examples, Fig. 4, the cooperation negotiation process COOP-NEGOT may further comprise at least one of: a) exchanging 300c capability information I-CAP, or b) determining 300d an input measure IM-PAR for the at least one parallel model exploration PAR-EXPLOR, or c) determining 300e an input measure IM-SEQ for the at least one sequential model exploration SEQ- EXPLOR.
[0073] In some examples, Fig. 4, exchanging 300c the capability information I-CAP, e.g. , with at least one further entity 20a, 20b, 20c, may be used, e.g. , to determine at least one of: a) first time resources for performing the at least one parallel model exploration PAR-EXPLOR involving the RL agents RLA-1, RLA- 2, ... (e.g. , for preparing the sequential model exploration) , b) parameters, e.g. , hyper parameters, for at least one of the bl) sequential model exploration SEQ-EXPLOR or b2) parallel model exploration PAR-EXPLOR. In some examples, the parameters may comprise at least one parameter, e.g. , hyper parameter, related to the RL agents RLA-1, RLA-2, such as, e.g. , at least one of: a) a learning rate, or b) an exploration probability, or c) a policy update periodicity, etc.
[0074] In some examples, Fig. 4, exchanging 300c the capability information I-CAP may comprise a modeling of MDP related configurations, e.g. , comprising at least one of: a) state, or b) action, or c) reward, or d) policy, etc. In some examples, at least some aspects of the MDP related configurations may be use case specific. In some examples, Fig. 4, exchanging 300c the capability information I-CAP may comprise determining a, for example total, number of time resources, e.g. , expressed in milliseconds, ms, or characterized by a System Frame Number (SFN) , e.g. , for executing at least one of: a) a, for example initial, parallel model exploration PAR-EXPLOR, or b) a sequential model exploration SEQ- EXPLOR.
[0075] In some examples, Fig. 4, the determining 300d of the input measure IM-PAR for the at least one parallel model exploration PAR-EXPLOR and / or the determining 300e of the input measure IM- SEQ for the at least one sequential model exploration SEQ-EXPLOR may comprise at least one of the following aspects.
[0076] In some examples, an average reward value of the RL agents RLA-1, RLA-2, ... (Fig. 2) may be used as an input measure, wherein, for example, the time domain pattern TD-SEQ for the sequential model exploration SEQ-EXPLOR may be based on: t0=r° X Ttotal, r total rr where r0, rlfr2are, for example buffered, totalrtotal reward values from each agent RLA-1, RLA-2, .... Note that the present example, without loss of generality, relates to three RL agents, while, for example, only two RL agents RLA-1, RLA-2 are explicitly depicted by Fig. 2. In some other examples related to Fig. 12 et seq. , as explained further below, example configurations are disclosed related to three or more RL agents.
[0077] In some examples, a time portion t0for performing sequential model exploration by a specific RL agent may be determined based on a configured total time resource Ttotalfor the sequential model exploration and based on a ratio rr° of a respective individual total reward value r0of the specific RL agent and an aggregated reward value rtotai• In some examples, this may enable to prioritize such RL agents, e.g. , providing a comparatively large time resource for sequential model exploration for such RL agent, which is associated with a comparatively high share of the aggregated reward value rtotal•
[0078] In some examples, the individual reward value r0, rlrr2of the specific RL agents may be used to determine or indicate a scheduling order for the sequential model exploration by the RL agents. Thus, in some examples, a scheduling order may, e.g. , be defined such that a specific RL agent with the largest individual reward value is scheduled as the first RL agent to perform sequential model exploration, e.g. , followed by another RL agent with the second largest individual reward value, and so on. In some examples, this scheduling order may be characterized by the exploration order policy ORD-POL, see also block 300b of Fig. 4.
[0079] In some examples, a total reward rtotalmay be defined in many ways, for example by rtotal= ro +ri +r2 • Insome other examples, the total reward rtotalmay be formulated by using a set of linear weighting coefficients, i.e. , rtotal= <z0X r0+ og X Tq + <z2xr2, where <z0+ a1+ a2= l.
[0080] In some examples, at least one of the following input measures may be used, e.g. , for determining the exploration order policy ORD-POL, e.g. , scheduling sequence: a) cell throughput, or b) cell average load, or c) traffic conditions, or d) average delay performance, e) user happiness (e.g. , as may be characterized by an XR capacity) . Note that in some examples further aspects may be used alternatively or additionally for determining a scheduling sequence.
[0081] In some examples, the devices 20a, 20b, 20c or their associated RL agents RLA-1, RLA-2, ... respectively, may execute aspects of the sequential model exploration, e.g. , based on the time domain pattern TDP-SEQ or based on the related information I-TDP-SEQ, and based on the exploration order policy ORD-POL, which, in some examples, may, e.g. , characterize the determined time resources and scheduling orders.
[0082] In some examples, Fig. 2, the RL agents RLA-1, RLA-2 may perform their respective aspects of the sequential model exploration in an ordered way as defined by the exploration order policy ORD- POL, wherein, for example, the first RL agent RLA-1 may have the highest priority in a scheduling order, so it may start exploring firstly, while the further RL agent RLA-2 may, e.g. , stay in an idle mode, e.g. , as long as the first RL agent RLA-1 is performing its part of the sequential model exploration. After that, e.g. , when the first RL agent RLA-1 has finished its part of the sequential model exploration, the further RL agent RLA-2 may perform its part of the sequential model exploration.
[0083] In some examples, Fig. 2, the time domain resource allocation TD- RA further comprises determining the time domain resource allocation TD-RA as one or more time units. In some examples, the one or more time units may, e.g. , comprise or may be characterized by at least one of: a) a frame, or b) a frame number, or c) a system frame number, or d) a time value, or e) a transmission time interval, TTI.
[0084] In some examples, Fig. 5A, 5B, at least one of a) the report REPPAR of the at least one outcome of the at least one parallel model exploration, or b) the report REP-SEQ of the at least one outcome of the at least one sequential model exploration may further comprise results RES-TRAIN, RES-TRAIN' of a training of the model.
[0085] In some examples, Fig. 5A, the report REP-PAR of the at least one outcome OUT-PAR-EXPLORE of the at least one parallel model exploration may further comprise at least one of: a) an indication IND-REW of one or more reward values (e.g. , average reward values) , e.g. , as associated with the one or more RL agents, or b) one or more performance indicators IND-PERF, e.g. , key performance indicators.
[0086] In some examples, the one or more performance indicators IND-PERF may comprise at least one of: a) a continuous measure, e.g. , based on at least one of: al) cell-level key performance indicators (KPIs) , or a2) cell throughput, or a3) user happiness (as may, e.g. , be characterized by a quality of service, QoS, measure for extended reality, XR, services) , etc. , or b) respective quantized and / or discretized measure (s) , e.g. organized in the form of at least one bitmap.
[0087] In some examples, Fig. 6, the instructions 106, when executed by the at least one processor 102, may cause the coordinating entity 10 to evaluate 310 the received at least one outcome OUT-PAR- EXPLORE, OUT-SEQ-EXPLORE (and / or the associated report REP-PAR, REP-SEQ) , e.g. , for determining at least one of: a) whether the at least one sequential model exploration SEQ-EXPLOR is to be repeated, see, for example, the optional block 312, or b) whether the cooperation negotiation process COOP-NEGOT is to be repeated (see optional block 312) , or c) whether one or more additional sequential model explorations SEQ-EXPLOR-ADD are to be performed, see the optional block 314.
[0088] In some examples, Fig. 7, the cooperation negotiation process COOP-NEGOT may further comprise at least one of: negotiating 300f on transferring the time domain resource allocation TD-RA, or determining 300g the time domain pattern TDP-SEQ to one of the plurality 20 of network nodes, e.g. , for certain period of time.
[0089] In some examples, Fig. 2, the coordinating entity 10 may be a coordinating entity for a wireless, e.g. , cellular, communication system 1000.
[0090] In some examples, Fig. 2, the communication system 1000 may adhere to and / or may be based on some accepted (and / or planned) specification, e.g. , standard, such as, e.g. , 3G, 4G, 5G, 6G, or some other wireless communication standard.
[0091] In some examples, Fig. 2, the coordinating entity 10 may be located in at least one of: a core network CN, or an edge cloud
[0092] EC, or a radio access network RAN.
[0093] Some examples, Fig. IB, relate to an apparatus 100' comprising means 102' for causing a coordinating entity 10 (Fig. 2) to: carry out 300 (Fig. 3) a cooperation negotiation process, wherein the cooperation negotiation process comprises identifying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes, and determining an exploration order policy for the at least one sequential model exploration; receive 302 a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes; determine 304 a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes, wherein the determination is based on the received report, the time domain resource allocation and the exploration order policy; transmit 306, to the plurality of network nodes, information on the time domain pattern, and receive 308, from the plurality of network nodes, a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern.
[0094] In some examples, Fig. IB, the means 102' for causing the coordinating entity 10 to perform at least one of the above- mentioned aspects 300, 302, 304, 306, 308 may comprise at least one processor 102 (see, for example, Fig. 1A) , and at least one memory 104 storing instructions 106 that, when executed by the at least one processor 102, cause the coordinating entity 10 to perform at least one of the above-mentioned aspects 300, ..., 308. In some examples, Fig. IB, the means 102' for causing the coordinating entity to perform at least one of the above- mentioned aspects 300, ..., 308 may, e.g. , comprise circuitry (not shown) configured to perform at least one of the aforementioned aspects .
[0095] Some examples, Fig. 2, relate to a coordinating entity 10 comprising the apparatus 100, 100' according to the embodiments.
[0096] In some examples, the coordinating entity 10 may be a machine learning orchestrator , MLO, e.g. , for a wireless communication system 1000.
[0097] Some examples, Fig. 3, relate to a method, comprising: carrying out 300, by a coordinating entity 10, a cooperation negotiation process, wherein the cooperation negotiation process comprises identifying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes, and determining an exploration order policy for the at least one sequential model exploration; receiving 302 a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes; determining 304 a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes, wherein the determination is based on the received report, the time domain resource allocation and the exploration order policy; transmitting 306, to the plurality of network nodes, information on the time domain pattern, and receiving 308, from the plurality of network nodes, a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern.
[0098] Some examples, Fig. 8A, 9, relate to an apparatus 200 comprising at least one processor 202, and at least one memory 204 storing instructions 206 that, when executed by the at least one processor 202, cause a coordinating network node 20a (Fig. 2) to carry out 350 a cooperation negotiation process COOP-NEGOT ' , wherein the cooperation negotiation process COOP-NEGOT', see Fig. 4, comprises identifying 300a (e.g. , determining an identity and / or number of) a plurality 20 of network nodes for time domain resource allocation TD-RA for at least one parallel model exploration PAR-EXPLORE and for at least one sequential model exploration SEQ-EXPLORE to be carried out by the plurality 20 of network nodes, and determining 300b an exploration order policy ORD-POL for the at least one sequential model exploration; carry out 352 (Fig. 9) a synchronization SYNCH to a common reference time with the plurality 20 of network nodes; receive 354 a report REP-PAR of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes; determine 356 a time domain pattern TDP-SEQ for the at least one sequential model exploration to be carried out by the plurality 20 of network nodes, wherein the determination 356 is based on the received report REP-PAR, the time domain resource allocation TD-RA and the exploration order policy ORD-POL; transmit 358 to a coordinating entity 10 (Fig. 2) , a request REQ for grant for the determined time domain pattern TDP-SEQ; in response to receiving the grant (e.g. , from the coordinating entity 10) , transmit 360, to the plurality 20 of network nodes, e.g. , at least to network nodes 20b, 20c, information I-TDP-SEQ on the time domain pattern TDP-SEQ, and receive 362, from the plurality of network nodes, a report REP-SEQ (see, for example, Fig. 5B) of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern.
[0099] In some examples, the common reference time may, e.g. , be determined based on at least one of: a) a global navigation satellite system (GNSS) , or b) a system frame number (SFN) , etc. In some examples, the synchronization SYNCH may, e.g. , be used for the at least one parallel model exploration and / or for the at least one sequential model exploration.
[0100] In some examples, an absolute time indicator with, e.g. , a certain precision, e.g. , UTC time with 1 ms precision, may be used, e.g. , for the synchronization SYNCH. In some examples, a relative time indicator may be used, e.g. , relative to an initial synchronization time (e.g. , epoch) . In some examples, the time indicator may, e.g. , be based on a counter, e.g. , with a certain increment step or step size, respectively.
[0101] In some examples, the parallel model exploration PAR-EXPLOR may be performed to provide a prior knowledge, e.g. , related to the RL agents RLA-1, RLA-2, e.g. , related to learning priorities associated with the involved entities 20a, 20b, 20c.
[0102] In some examples, Fig. 9, the cooperation negotiation process COOP-NEGOT' may further comprise at least one of: a) exchanging 300c (Fig. 4) capability information I-CAP or b) determining 300d an input measure IM-PAR for the at least one parallel model exploration, or c) determining 300e an input measure IM-SEQ for the at least one sequential model exploration.
[0103] In some examples, the time domain resource allocation may further comprise determining the time domain resource allocation as one or more time units. In some examples, as mentioned above, the one or more time units may, e.g. , comprise or may be characterized by at least one of: a) a frame, or b) a frame number, or c) a system frame number, or d) a time value, or e) a transmission time interval, TTI.
[0104] In some examples, as mentioned above, at least one of a) the report REP-PAR (Fig. 5A) of the at least one outcome of the at least one parallel model exploration, or b) the report REP-SEQ (Fig. 5B) of the at least one outcome of the at least one sequential model exploration may further comprise results RESTRAIN, RES-TRAIN' of a training of the model.
[0105] In some examples, as mentioned above, the report REP-PAR of the at least one outcome of the at least one parallel model exploration may further comprise at least one of: a) an indication of one or more reward values, or b) one or more performance indicators.
[0106] In some examples, Fig. 10, the instructions 206, when executed by the at least one processor 202, may cause the coordinating network node 20a to evaluate 370 the received at least one outcome OUT-PAR-EXPLORE, OUT-SEQ-EXPLORE (and / or the associated report REP-PAR, REP-SEQ) , e.g. , for determining at least one of: a) whether the at least one sequential model exploration is to be repeated, see the optional block 372 of Fig. 10, or b) whether the cooperation negotiation process COOP-NEGOT ' is to be repeated, see the optional block 372, or c) whether one or more additional sequential model explorations SEQ-EXPLOR-ADD are to be performed, see the optional block 374.
[0107] In some examples, the evaluating 310 (Fig. 6) and / or the evaluating 370 (Fig. 10) may comprise at least one of: a) using a timer for determining a time associated with the sequential model exploration SEQ-EXPLORE, b) determining a stability parameter, e.g. , associated with at least one RL agent RLA-1, RLA-2, ... , c) determining the performance indicator IND-PERF, or d) storing the outcome OUT-PAR-EXPLORE, OUT-SEQ-EXPLORE.
[0108] In some examples, using the timer may, e.g. , be performed to determine whether a total time resource Ttotal, e.g. , for the sequential model exploration, has elapsed. In some examples, using the timer may, e.g. , be performed to determine whether the time portion t0for performing sequential model exploration by a specific RL agent RLA-1 has elapsed.
[0109] In some examples, a resolution of the timer may, e.g. , be in the millisecond range, e.g. , 1 ms.
[0110] In some examples, determining the stability parameter may comprise evaluating at least one stability or convergence condition associated with at least one RL agent RLA-1. In some examples, it may be determined, whether a temporal difference (TD) error of each RL agent RLA-1, RLA-2 reaches a stable status. For example, in some examples, TD2<62, and TD3<63, it may be concluded that the sequential model exploration reaches stability. In other words, in some examples, determining the stability parameter may comprise determining, for at least some of the RL agents RLA-1, RLA-2 involved in the sequential model exploration, whether a respective TD error TDj of an i-th one of the RL agents is equal to or less than a respective predetermined threshold . Note that, in some examples, the threshold may be specific for a respective RL agent.
[0111] In some examples, Fig. 5A, determining the performance indicator IND-PERF may comprise evaluating at least one radio performance KPI stability condition. In some examples, it may be proposed to select one common KPI for the entities 20a, 20b, 20c associated with the RL agents RLA-1, RLA-2, ..., and their respective performance may, e.g. , be subjected to thresholds. In some examples, a cell throughput of a radio cell (e.g. , indicated in Mbps) may be selected as a common KPI, and, in some examples, the sequential model exploration may be terminated if at least one condition related to the cell throughput is satisfied, such as, e.g. , if Tput-L ^ q, Tput2<y2, and Tput3< y3all hold true. In other words, in some examples, the sequential model exploration may be terminated if a respective cell throughput Tput1(Tput2, Tput2associated with the entities 20a, 20b, 20c (Fig. 2) , is equal to or less than a respective predetermined threshold ylry2, y2.
[0112] In some examples, Fig. 7, the cooperation negotiation process COOP-NEGOT' may further comprise at least one of: negotiating 300f on transferring the time domain resource allocation TD-RA, or determining 300g' the time domain pattern TDP-SEQ to the coordinating entity 10.
[0113] Some examples, Fig. 8B, relate to an apparatus 200' comprising means 202' for causing a coordinating network node 20a to carry out 350 a cooperation negotiation process, wherein the cooperation negotiation process comprises identifying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes, and determining an exploration order policy for the at least one sequential model exploration; carry out 352 a synchronization to a common reference time with the plurality of network nodes; receive 354 a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes; determine 356 a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes, wherein the determination is based on the received report, the time domain resource allocation and the exploration order policy; transmit 358 to a coordinating entity, a request for grant for the determined time domain pattern; in response to receiving the grant, transmit 360, to the plurality of network nodes, information on the time domain pattern, and receive 362, from the plurality of network nodes, a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern . In some examples, Fig. 8B, the means 202' for causing the coordinating network node 20a to perform at least one of the above-mentioned aspects 350, ..., 362 may comprise at least one processor 202 (Fig. 8A) , and at least one memory 204 storing instructions 206 that, when executed by the at least one processor 202, cause the coordinating network node 20a to perform at least one of the above-mentioned aspects 350, ..., 362. In some examples, the means 202' for causing the coordinating network node 20a to perform at least one of the above-mentioned aspects 350, ..., 362 may, e.g. , comprise circuitry (not shown) configured to perform at least one of the aforementioned aspects 350, ..., 362.
[0114] Some examples, Fig. 2, relate to a coordinating network node 20a comprising an apparatus 200, 200' according to the embodiments.
[0115] In some examples, Fig. 2, the coordinating network node 20a may be a network device for a wireless, e.g. , cellular, communication system 1000, such as, for example, a base station, e.g. , gNB. In some, but not necessarily all, examples, the apparatus 200, 200' according to the embodiments or its functionality, respectively, may, e.g. , be integrated into at least one component of the network device.
[0116] In some examples, the network device 20a may adhere to and / or may be based on some accepted (and / or planned) specification, e.g. , standard, such as, e.g. , 3G, 4G, 5G, 6G, or some other wireless communication standard.
[0117] In some examples, the coordinating network node 20a may be, in a distributed case, a central unit, CU, of a network node, e.g. , gNB, and the plurality of network nodes may comprise distributed units, DU, of the network node, e.g. , gNB, with radio unit accessibility. In some examples, element 20a of Fig. 2 may, e.g. , symbolize a gNB-CU, and elements 20b, 20c, may, e.g. , symbolize gNB-DUs associated with the gNB-CU 20a.
[0118] Some examples, Fig. 9, relate to a method, comprising: carrying out 350, by a coordinating network node, a cooperation negotiation process, wherein the cooperation negotiation process comprises identifying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes, and determining an exploration order policy for the at least one sequential model exploration; carrying out 352 a synchronization to a common reference time with the plurality of network nodes; receiving 354 a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes; determining 356 a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes, wherein the determination is based on the received report, the time domain resource allocation and the exploration order policy; transmitting 358 to a coordinating entity, a request for grant for the determined time domain pattern; in response to receiving the grant, transmitting 360, to the plurality of network nodes, information on the time domain pattern, and receive 362, from the plurality of network nodes, a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern .
[0119] Some examples, Fig. 2, relate to a communication system, e.g. , a wireless, e.g. , cellular, communication system 1000, comprising at least one of: a) at least one apparatus 100, 100' , 200, 200' according to the embodiments, or b) a coordinating entity 10 according to the embodiments, or c) a coordinating network node
[0120] 20a according to the embodiments. Fig. 11 schematically depicts a simplified flow chart according to some examples, wherein, for example, a plurality 20 (Fig. 2) of network nodes 20a, 20b, 20c or devices, e.g. , gNB may be associated with a respective RL agent RLA-1, RLA-2, ... each. Block 500 of Fig. 11 symbolizes a capability exchange between at least the plurality 20 of gNB, e.g. , an "inter-gNB" capability exchange, e.g. , with the coordinating entity, e.g. , MLO 10, e.g. , to determine gNB which may, at least temporarily, be coordinated, e.g. , with respect to a parallel model exploration or sequential model exploration, e.g. , with respect to their associated RL agents RLA-1, RLA-2, ....
[0121] Block 502 of Fig. 11 comprises determining a time allocation for the sequential model exploration involving the plurality of RL agents of the coordinated gNB 20a, 20b, 20c. In some examples, in accordance with block 502, the MLO 10 (Fig. 2) may determine time resources to be allocated to the plurality 20 of gNB or their respective RL agents. In some examples, in accordance with block 502, at least one of the plurality 20 of gNB, e.g. , a coordinating gNB 20a, may determine time resources to be allocated to the plurality 20 of gNB or their respective RL agents. In some examples, in accordance with block 502, dynamic (e.g. , during operation) switching between a) a determination of time resources by the MLO 10 and b) a determination of time resources by at least one of the gNB may be provided, wherein, for example, the dynamic switching may, e.g. , be controlled by the MLO 10.
[0122] Block 504 of Fig. 11 symbolizes an execution of the sequential model exploration SEQ-EXPLOR, wherein, for example, the gNB 20a, 20b, 20c may use the determined time resources (see block 502) as the exploration time for the sequential model exploration.
[0123] Block 506 symbolizes at least one, for example all, gNB providing an outcome OUT-SEQ-EXPLORE (Fig. 5B) of the sequential model exploration, e.g. , comprising buffering of a priority measure or metrics, e.g. , within an allocated time interval, and transmission of the outcome OUT-SEQ-EXPLORE of the sequential model exploration to a further entity, e.g. , the MLO 10, e.g. , in the form of the report REP-SEQ.
[0124] Block 508 of Fig. 11 symbolizes an evaluation of the outcome of the sequential model exploration by the MLO 10, e.g. , deciding to reschedule time resources for at least one of a future parallel model exploration or sequential model exploration, e.g. , based on the outcome of the sequential model exploration.
[0125] Fig. 12 schematically depicts a simplified signaling diagram according to some examples. Element El symbolizes a coordinating entity, e.g. , MLO, and elements E2a, E2b, E2c symbolize base stations, e.g. , gNB, each of which may be associated with, e.g. , comprises, an RL agent (not shown in Fig. 12, see, for example elements RLA-1, RLA-2, ... of Fig. 2) . I.e. , in some examples, at each gNB E2a, E2b, E2c, a respective RL agent may be deployed, e.g. , for the parallel and sequential model exploration.
[0126] Element E3 of Fig. 12 symbolizes a capability exchange between the entities El, E2a, E2b, E2c. In some examples, the capability exchange E3 of Fig. 12 may comprise at least some of the aspects of block 300 of Fig. 3.
[0127] In some examples, the MLO El may, e.g. , be a 5G network entity that, for example, supports Network Data Analytics, such as 0AM (Operations And Management) or NWDAF (Network Data Analytics Function) or, for example also, RIC (RAN Intelligent Controller) . In some other examples, e.g. , alternatively, the RL agents may be deployed at gNB-DUs, and the MLO can, e.g. , be hosted at a gNB- CU.
[0128] In some examples, Fig. 12, a number of coordinated gNB E2a, E2b
[0129] E2c may be determined, e.g. , in the capability exchange E3. In some examples, the gNB E2a may be denoted as "source gNB" or "gNBO". In some examples, two further gNB E2b, E2c may be identified, e.g. , for coordination with the source gNB E2a, e.g. , regarding the sequential model exploration, e.g. , using a neighboring selection algorithm, example aspects of which are explained further below with reference to Fig. 16.
[0130] In some examples, Fig. 12, the MLO El, e.g. , during the capability exchange E3, e.g. , based on a total number of coordinated gNBs E2a, E2b, E2c, may determine a total number of time resources (e.g. , expressed in ms or System Frame Number (SEN) ) for executing at least one of a) an initial parallel model exploration E4, or b) a sequential model exploration E7. In some examples, an ordering policy ORD-POL and / or relevant input measures IM-SEQ, e.g. , for the sequential model exploration, may be agreed in block E3, e.g. , between the gNBs E2a, E2b, E2c and the MLO El.
[0131] In some examples, parameters, e.g. , hyper parameters, e.g. , related to the RL agents, including at least one of: a) learning rate, or b) exploration probability, or c) policy update periodicity may be determined in the capability exchange E3. In some examples, a modeling or determination of MDP related, e.g. , RL agent related, configurations such as at least one of a) state, or b) action, or c) reward, or d) policy, may be determined in the capability exchange E3.
[0132] Element E4 of Fig. 12 symbolizes a synchronization, wherein, for example, the gNBs E2a, E2b, E2c may synchronize to a common reference timing e.g. , GNSS or SEN, and a parallel model exploration involving the RL agents of the gNBs E2a, E2b, E2c. Arrows al, a2, a3 of Fig. 12 symbolize the gNBs E2a, E2b, E2c reporting, to the MLO El, their respective outcome of the parallel model exploration, e.g. , using at least one continuous measure or quantized / discretized measure, e.g. , in form of at least one bitmap.
[0133] Element E5 of Fig. 12 symbolizes an evaluation of the reports al, a2, a3 by the MLO El, and element E6 symbolizes a determination of time resources for a future sequential model exploration E7, based on the evaluation E5, see, for example, the abovementioned time allocation, e.g. , determining time resources based on: t0=
[0134] The arrows a4, a5, a6 symbolize transmissions of the respective time allocation, as, e.g. , obtained by MLO El in block E6, e.g. , in the form of the information I-TDP-SEQ, to the gNBs E2a, E2b, E2c. On this basis, in some examples, the RL agents of the gNBs E2a, E2b, E2c may execute the sequential model exploration E7.
[0135] Element E8 of Fig. 12 symbolizes an evaluation of one or more termination conditions for the sequential model exploration E7, e.g. , jointly, e.g. , by the gNBs E2a, E2b, E2c, e.g. , based on at least one of the aspects a) time, or b) stability parameter, or c) performance indicator.
[0136] The arrows a7, a8, a9 symbolize the gNBs E2a, E2b, E2c reporting, to the MLO El, their respective outcome (s) of the sequential model exploration E7, e.g. , using at least one continuous measure or quantized / discretized measure, e.g. , in form of at least one bitmap, e.g. , similar to the arrows al, a2, a3.
[0137] Element E9 of Fig. 12 symbolizes an optional process of storing the outcome of the sequential model exploration, e.g. , for use as a reference.
[0138] Element E10 symbolizes an evaluation, e.g. , at least similar to block 310 of Fig. 6, and / or an optional repetition of at least one block of Fig. 12, e.g. , (re- ) starting the procedure, for example at element E3 or element E5. In some examples, when restarting the procedure at element E3, the MLO El may reconfigure and identify the coordinated gNBs E2a, E2b, E2c for synchronization and coordination. In some examples, when restarting the procedure at element E5, the MLO may continue evaluating and updating the time resource allocation, e.g. , for the same gNBs E2a, E2b, E2c, e.g. , for coordination.
[0139] Alternatively, at block E10, no further iterations, e.g. , restarting from block E3 or E5, may be performed, but rather, for example, an exploitation may be executed, e.g. , using the RL agents of the gNBs E2a, E2b, E2c.
[0140] Fig. 13A, 13B depict aspects of an operation of RL agents RLA-1', RLA-2 ' , RLA-3', each of which is associated with, e.g. , implemented by, a respective gNB, according to some examples. In this regard, block 510 of Fig. 13B symbolizes starting a sequential model exploration of the RL agents RLA-1' , RLA-2', RLA-3', and block T of Fig. 13A symbolizes a total amount of time resources that may be provided for the process of the sequential model exploration by all three RL agents RLA-1' , RLA-2' , RLA-3' , e.g. , corresponding to Ttotalas explained above.
[0141] Element tO of Fig. 13A symbolizes a scheduled time resource for the RL agent RLA-1', e.g. , similar to the time portion t0as explained above, element tl symbolizes a scheduled time resource for the RL agent RLA-2', e.g. , similar to the time portion as explained above, and element t2 symbolizes a scheduled time resource for the RL agent RLA-3' , e.g. , similar to the time portion t2as explained above.
[0142] Block 510a of Fig. 13B symbolizes the RL agent RLA-1' executing the model exploration, in the sense of the sequential model exploration according to the examples, for the duration of tO, wherein the further RL agents RLA-2' , RLA-3' do not perform exploration, but, e.g. , may stay in an idle mode. Block 510b of Fig. 13B symbolizes the RL agent RLA-2 ' executing the model exploration, in the sense of the sequential model exploration, for the duration of tl, wherein the further RL agents RLA-1', RLA-3' do not perform exploration, but may, e.g. , stay in an idle mode. Block 510c of Fig. 13B symbolizes the RL agent RLA-3' executing the model exploration, in the sense of the sequential model exploration, for the duration of t2, wherein the further RL agents RLA-1' , RLA-2' do not perform exploration, but, may e.g. , stay in an idle mode.
[0143] In some examples, the scheduled time resources may be changed, e.g. , re-ordered, see block arrow BAI of Fig. 13A, e.g. , based on respective priorities associated with the RL agents, wherein changed scheduled time resources t0', tl' , t2 ' can be obtained, which may, e.g. , be used for a further sequential model exploration .
[0144] Fig. 14 schematically depicts a simplified signaling diagram according to some examples. Element E20 symbolizes a coordinating device, e.g. , MLO, e.g. , similar or identical to MLO El of Fig. 12, and elements E21a, E21b, E21c symbolize base stations, e.g. , gNB, each of which is associated with, e.g. , comprises, an RL agent (not shown in Fig. 14) . In some examples, element E21a may be configured or operate as a gNB-CU (e.g. , "source gNB") , and elements E21b, E21c may be configured or operate as a gNB-DUs.
[0145] Element E22 of Fig. 14 symbolizes a capability exchange between the entities, e.g. , similar or identical to element E3 of Fig. 12. In some examples of Fig. 14, in difference to Fig. 12, the source gNB E21a may operate as a master node, e.g. , to control a collection of information from the further gNBs (e.g. , "target gNBs") E21b, E21c and to report to the MLO E20. Element E23 symbolizes a synchronization of the entities E21a, E21b, E21c, e.g. , and a parallel model exploration, similar or identical to element E4 of Fig. 12.
[0146] The arrows a20, a21 symbolize the target gNBs E21b, E21c reporting, to the source gNB E21a, an outcome of the parallel model exploration, e.g. , comprising buffered exploration and training results as obtained during the parallel model exploration, e.g. , using at least one continuous measure or quantized / discretized measure, e.g. , in form of at least one bitmap .
[0147] Element E24 symbolizes generating, based on the outcome of the parallel model exploration E23, a determination of time pattern information, e.g. , for an execution of a subsequent sequential model exploration E25, e.g. , based on the time resources as explained above with reference to: t0X Ttotal, t2= - ^2 r X Ttotai • Insome examples, the time pattern information may total be used to control which RL agent of which entity E21a, E21b, E21c performs exploration during which period of time, i.e. , controlling a duration for the model exploration as well as a priority order indicating which RL agent may start with the sequential model exploration.
[0148] Fig. 15 schematically depicts an illustration of the time pattern information according to some examples for three RL agents RLA- 1' , RLA-2 ' , RLA-3', as may, e.g. , be obtained at element E24 of Fig. 14. Bracket TP-SL of Fig. 15 symbolizes a time pattern which may be repeated, e.g. , in a slot level, whereas bracket T-SEQ- EXPLOR symbolizes a total time which may be used for a sequential model exploration. In some examples, on a slot level, the first four slots Al may be assigned to the RL agent RLA-1', the next three slots A2 may be assigned to the RL agent RLA-2 ' , and the last two slots A3 may be assigned to the RL agent RLA-3' . From Fig. 15, it can be seen that, in some examples, the time pattern information may be interpreted as a time pattern framework having characteristics of a time division duplexing, TDD, technique or a TDD frame configuration.
[0149] In some examples, Fig. 14, the source gNB E21a may transmit a request associated with, e.g. , characterizing aspects of, the time pattern information for the sequential model exploration to the MLO E20, see arrow a22, and the MLO E20 may respond with an acknowledgement, see arrow a23.
[0150] In some examples, the source gNB E21a may transmit the granted (e.g. , based on the acknowledgement a23) time pattern information (see, for example, also the information I-TDP-SEQ of block 360 of Fig. 9) to the target gNBs E21b, E21c, see the arrows a24, a25, e.g. , before the execution of sequential model exploration E25. In some examples, this way, each entity E21b, E21c may receive respective parts of the information I-TDP-SEQ relevant to an operation of its associated RL agent. Thus, in the examples of Fig. 14, the gNBs E21a, E21b, E21c may perform the sequential model exploration E25 based on the time pattern information as obtained by the source gNB E21a, see element E24.
[0151] In some examples, Fig. 14, the gNBs E21a, E21b, E21c may, e.g. , jointly, evaluate one or more termination conditions for the sequential model exploration, see element E26, e.g. , at least similar to element E8 of Fig. 12.
[0152] After that, the target gNBs, E21b, E21c may report an outcome of the sequential model exploration to the source gNB E21a, see the arrows a26, a27, the outcome, e.g. , comprising buffered exploration and training results of the sequential model exploration, e.g. , similar to arrows a20, a21 for the outcome of the parallel model exploration. Element E27 of Fig. 14 symbolizes an optional process of storing the outcome of the sequential model exploration, e.g. for use as a reference, e.g. , similar to element E9 of Fig. 12.
[0153] Element E28 of Fig. 14 symbolizes an optional repetition of at least one block of Fig. 14, e.g. , (re- ) starting the procedure, for example at element E22 or element E24. In some examples, when restarting the procedure at element E22, the MLO E20 may reconfigure and identify the coordinated gNBs E21a, E21b, E21c for synchronization and coordination. In some examples, when restarting the procedure at element E24, the source gNB E21a may continue evaluating and updating the time resource allocation, e.g. , for the same gNBs E21a, E21b, E21c, e.g. , for coordination.
[0154] Alternatively, at block E28, no further iterations, e.g. , restarting from block E22 or E24, may be performed, but rather, for example, an exploitation may be executed, e.g. , using the RL agents of the gNBs E21a, E21b, E21c.
[0155] In some examples, a, for example dynamic, switching between an MLO controlled scheduling mechanism of some examples as, e.g. , illustrated by Fig. 12 and a time pattern framework of some examples as, e.g. , determined by gNB E21a, see Fig. 14, may be performed, e.g. , using the capability exchange of element E3 (Fig. 12) or element E22 (Fig. 14) for the switching. In other words, some examples relate to determining, e.g. , during the capability exchange, whether a switching between the MLO controlled scheduling mechanism of, e.g. , Fig. 12 and the gNB controlled mechanism of, e.g. , Fig. 14 should be performed, and, based on this determination, performing the switching. In some examples, the determination and the switching may be performed by at least one of the MLO El, E20 and a gNB, e.g. , source gNB E2a, E21a. While some of the above-explained examples, e.g. , of Fig. 12 and 14, are related to examples of coordinating entities E2a, E2b, ... of the gNB type (e.g. , "inter-gNB level coordination") , in some examples, the principle according to the embodiments may also be applied to other entities, such as, terminal devices, e.g. , UE 30 (Fig. 2) or combinations of network devices such as, e.g. , gNB 20 or MLO 10, and terminal devices 30 (e.g. , "gNB-UE level coordination") .
[0156] In some examples, a gNB-CU or gNB-DU (not shown) may act as an MLO 10, and one or more RL agents may, e.g. , be deployed at a terminal device, e.g. , UE, side, e.g. , for sequential model exploration .
[0157] In some examples, for example for inter-gNB level coordination, information may be exchanged between the involved entities, e.g. , over at least one of the following interfaces: a) Xn, or b) Fl (e.g. , using XnAP or F1AP signalling) . In some examples, the MLO may be a 5G network entity that may support Network Data Analytics, such as QAM or NWDAF or also RIC (Radio Intelligent Controller) .
[0158] In some examples, for example for gNB-UE level coordination, information may be exchanged over a Uu interface (e.g. , using RRC (Radio Resource Control) or MAC (Medium Access Control) signalling) . In some examples, the MLO 30 (Fig. 2) may be hosted at a gNB-CU or a gNB-DU.
[0159] In some examples, one or more information elements may be provided, e.g. , for transmitting and / or receiving at least one aspect of: a) the cooperation negotiation process, or b) the reports REP-PAR, REP-SEQ, or c) the information I-TDP-SEQ, or e) the request REQ, or f) an acknowledgement of the request REQ, or g) the time pattern information I-TIM- PATTERN, or h) aspects of the capability exchange 300c. In some examples, e.g. , depending on timing constraints, the one or more information elements may, e.g. , be transmitted, e.g. , conveyed, e.g. , by RRC signaling and / or LI (Layer 1) / L2 (Layer 2) signaling, e.g. , MAC CE (control element) or DCI (Downlink Control Information) .
[0160] In some examples, at least one accepted or planned specification, e.g. , standard, may be extended, e.g. , for providing the abovementioned one or more information elements, e.g. , for signaling, e.g. , between gNBs or gNB-CU and gNB-DU or between gNB and UE or between MLO and gNBs. In some examples, one or more information elements for intra-RAN signaling, e.g. , elements between gNBs or gNB-CU and gNB-DU may be provided, e.g. , by adding to or extending at least one of the following specifications: TS38.423 (XnAP) , TS 38.473 (F1AP) . In some examples, one or more information elements for signaling between gNB and UE may be provided, e.g. , by adding to or extending at least one of the following specifications: TS 38.331 (RRC) , TS 38.321 (MAC) .
[0161] Fig. 16 schematically depicts aspects of a neighbouring selection algorithm based on a UE clustering method, which may be used in some examples.
[0162] In some examples, at least some aspects of the following example of a neighbouring selection algorithm may be used, e.g. , by a source gNB, see, for example, element gNBO of Fig. 16 (also see gNB E2a (Fig. 12) , E21a (Fig. 14) ) , e.g. , to verify a list of potential neighbouring gNBs that, in some examples, may create an impact on different groups / clusters of UEs served by the source gNB.
[0163] *** start of algorithm example ***
[0164] Step 0: Start the algorithm, source gNB gNBO may collect relevant input data to drive a clustering method. In some examples, a slow faded RSRP may be collected from UE measurement report (s) to establish clusters in a cell Cl.
[0165] Step 1: UE clusters CLUST-1, CLUST-2 may be created in the source gNB gNBO based on RSRP measurement, where each cluster CLUST-1, CLUST-2 may correspond to a critical neighbouring gNB gNBl, gNB2 for coordination. In some examples, at this step, multiple UE clusters CLUST-1, CLUST-2 (e.g. , each with at least one UE) may be created based on the following example criterion:
[0166] Reference option (Strongest RSRP rule) : UEs may belong to a certain cluster if they share a critical neighbouring cell, which may be determined by the strongest neighbouring cell RSRP measurement. In some examples, by doing so, the creation of each UE cluster CLUST-1, CLUST-2 served by gNBO may result in each distinctive collaborative gNB with the highest interference correlation. In this regard, Fig. 16 shows an example case with two UE clusters CLUST-1, CLUST-2 generated, where each of the clusters CLUST-1, CLUST-2 may correspond to a critical neighbouring gNB, i.e. , gNBl and gNB2.
[0167] Step 2: the UE clusters CLUST-1, CLUST-2 may be ordered in the source gNB gNBO according to the predefined criterion, which may rank the critical neighbouring gNBs for joint exploration. The ranking methodology may generate the sorted list of critical neighbouring gNBs. In some examples, possible criteria for ordering UE clusters CLUST-1, CLUST-2 may be: i) Number of UEs in each cluster, ii) Average serving cell RSRP value for each cluster, iii) Average strongest neighbouring cell RSRP value for each cluster, iv) Average RSRPdiff value for each cluster, etc.
[0168] In some examples, UE clusters that contain fewer UEs may be discarded. Or given the minimum number of UEs in each cluster is nminr for clusters where the number of UEs UEs (of these clusters) may be re-assigned to other clusters. This step may aim at reducing the potential list of critical neighboring gNBs to reach the acceptable coordination complexity.
[0169] Step 3: i) the UE cluster IDs served by gNBO and ii) the selected list of critical neighbouring gNB IDs for the following coordinated exploration phase may be output / generated.
[0170] *** end of algorithm example ***
[0171] Some examples, Fig. 17, relate to a computer program PRG comprising instructions INSTR, which, when executed by an apparatus 100, 100' , 200, 200', cause the apparatus to perform at least some aspects of the method according to the examples.
[0172] In some examples, the computer program PRG may be provided on a computer readable storage medium SM, e.g. , a non-transitory computer readable medium.
[0173] Some examples, Fig. 17, relate to a data carrier signal DCS carrying and / or characterizing the computer program PRG according to the examples.
[0174] In some examples, the principle according to the embodiments may, e.g. , be used to provide a, for example dynamic, multi-agent RL coordination mechanism in time domain that may account for the sequential scheduling of a plurality of RL agents, e.g. , for a coordinated model exploration, e.g. , within an NG-RAN framework.
[0175] In some examples, different levels of coordination may be provided, such as, e.g. , inter-gNB and / or gNB to UE, etc. , optionally using a, for example central, entity such as an MLO.
[0176] In some examples, a, for example intra-RAN, coordination signaling configuration exchange may be provided, e.g. , between NG-RAN nodes and MLO such that: a) MLO may determine an exploration time resource to be allocated for a plurality of agents, e.g. , for each agent, e.g. , RL agent (e.g. , at gNB or UE) , e.g. , for sequential model exploration, or b) agents, e.g. , RL agents, (e.g. , at gNBs or UEs) may determine an exploration time resource, e.g. , in a compound pattern, e.g. , for sequential model exploration, or c) a dynamic switching between the aforementioned aspects a) , b) , e.g. , based on an outcome of the sequential model exploration.
[0177] In some examples, configuration information (and / or corresponding information elements, e.g. , for exchanging such configuration information) for at least one of the following aspects may be provided: a) total time resource for initial parallel model exploration among the RL agents, b) a maximum number of, e.g. coordinated, RL agents, e.g. , gNBs or UEs, c) a total time resource for sequential model exploration among the RL agents, d) ordering metrics for RL agents, e) sequential model exploration time pattern frame configuration, f) sequential model exploration exit conditions and related thresholds.
Claims
Claims1 . An apparatus comprising at least one processor, and at least one memory storing instructions that , when executed by the at least one processor, cause a coordinating entity to : carry out a cooperation negotiation process , wherein the cooperation negotiation process comprises identi fying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes , and determining an exploration order policy for the at least one sequential model exploration; receive a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes ; determine a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes , wherein the determination is based on the received report , the time domain resource allocation and the exploration order policy; transmit , to the plurality of network nodes , information on the time domain pattern, and receive , from the plurality of network nodes , a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern .2 . The apparatus according to claim 1 , wherein the cooperation negotiation process further comprises at least one of : a ) exchanging capability information, or b ) determining an input measure for the at least one parallel model exploration, or c ) determining an input measure for the at least one sequential model exploration .3 . The apparatus according to any of the preceding claims , wherein the time domain resource allocation further comprises determining the time domain resource allocation as one or more time units .4 . The apparatus according to any of the preceding claims wherein at least one of a ) the report of the at least one outcome of the at least one parallel model exploration, or b ) the report of the at least one outcome of the at least one sequential model exploration further comprises results of a training of the model .5 . The apparatus according to any of the preceding claims , wherein the report of the at least one outcome of the at least one parallel model exploration further comprises at least one of : a ) an indication of one or more reward values , or b ) one or more performance indicators .6 . The apparatus according to any of the preceding claims , wherein the instructions , when executed by the at least one processor, cause the coordinating entity to evaluate the received at least one outcome for determining at least one of : a ) whether the at least one sequential model exploration is to be repeated, or b ) whether the cooperation negotiation process is to be repeated, or c ) whether one or more additional sequential model explorations are to be performed .7 . The apparatus according to any of the preceding claims , wherein the cooperation negotiation process further comprises at least one of : negotiating on trans ferring the time domain resource allocation, or determining the time domain pattern to one of the plurality of network nodes .8 . The apparatus according to any of the preceding claims , wherein the coordinating entity is located in at least one of : a core network, or an edge cloud, or a radio access network .9 . An apparatus comprising at least one processor, and at least one memory storing instructions that , when executed by the at least one processor, cause a coordinating network node to carry out a cooperation negotiation process , wherein the cooperation negotiation process comprises identi fying a plurality of network nodes for time domain resource allocation for at least one parallel model exploration and for at least one sequential model exploration to be carried out by the plurality of network nodes , and determining an exploration order policy for the at least one sequential model exploration; carry out a synchroni zation to a common reference time with the plurality of network nodes ; receive a report of at least one outcome of the at least one parallel model exploration carried out by the plurality of network nodes ; determine a time domain pattern for the at least one sequential model exploration to be carried out by the plurality of network nodes , wherein the determination is based on the received report , the time domain resource allocation and the exploration order policy; transmit to a coordinating entity, a request for grant for the determined time domain pattern; in response to receiving the grant , transmit , to the plurality of network nodes , information on the time domain pattern, and receive , from the plurality of network nodes , a report of at least one outcome of the at least one sequential model exploration carried out according to the time domain pattern .10 . The apparatus according to claim 9 , wherein the cooperation negotiation process further comprises at least one of : a ) exchanging capability information or b ) determining an input measure for the at least one parallel model exploration, or c ) determining an input measure for the at least one sequential model exploration .11 . The apparatus according to any of the claims 9 to 10 , wherein the time domain resource allocation further comprises determining the time domain resource allocation as one or more time units .12 . The apparatus according to any of the claims 9 to 11 , wherein at least one of a ) the report of the at least one outcome of the at least one parallel model exploration, or b ) the report of the at least one outcome of the at least one sequential model exploration further comprises results of a training of the model .13 . The apparatus according to any of the claims 9 to 12 , wherein the report of the at least one outcome of the at least one parallel model exploration further comprises at least one of : a ) an indication of one or more reward values , or b ) one or more performance indicators .14 . The apparatus according to any of the claims 9 to 13 , wherein the instructions , when executed by the at least one processor, cause the coordinating network node to evaluate the received at least one outcome for determining at least one of : a ) whether the at least one sequential model exploration is to be repeated, or b ) whether the cooperation negotiation process is to be repeated, or c )whether one or more additional sequential model explorations are to be performed .15 . The apparatus according to any of the claims 9 to 14 , wherein the cooperation negotiation process further comprises at least one of : negotiating on trans ferring the time domain resource allocation, or determining the time domain pattern to the coordinating entity .16 . The apparatus according to any of the claims 9 to 15 , wherein the coordinating network node is , in a distributed case , a central unit of a network node and the plurality of network nodes comprise distributed units of the network node with radio unit accessibility .
Citation Information
Patent Citations
Balanced resource allocator for heterogeneous multi-objective systems
US11403142B1
Method and system for implementing reinforcement learning agent using reinforcement learning processor
US20180260700A1
Artificial Intelligence Engine Having Various Algorithms to Build Different Concepts Contained Within a Same AI Model
US20180357552A1