Synchronizing machine learning model training

A synchronization mechanism for ML model training in communication networks addresses inefficiencies by managing resource utilization and convergence times through context-based process control, enhancing energy efficiency and resource management.

WO2026135536A1PCT designated stage Publication Date: 2026-06-25TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
Filing Date
2025-05-13
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Existing machine learning model training in communication networks, particularly in distributed environments, is resource-intensive and inefficient due to prolonged convergence times and resource wastage when processes are resumed after timer-based interruptions.

Method used

Implementing a synchronization mechanism through network and training nodes to manage ML model training by sending synchronization requests with context, evaluating these requests, and configuring the training process accordingly to pause, stop, or resume based on detected events, thereby optimizing resource utilization.

Benefits of technology

Enhances energy efficiency and resource management by synchronizing training processes, reducing wastage and minimizing prolonged convergence times through proactive resource release and adaptive training management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2025050453_25062026_PF_FP_ABST
    Figure SE2025050453_25062026_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure provides methods and apparatus for synchronizing machine learning, ML, model training in a communication network (101) The method (200, 400) performed by a network node (103), the method (200, 400) comprises obtaining (201, 402a, 402b), from a training party involved in the ML model training in the communication network (101), a synchronization request comprising a first context indicating a need for synchronizing the ML model training. The method comprises evaluating (202, 403) the first context comprised in the synchronization request. The method comprises notifying (203, 404) the training party a result of the evaluation; and configuring (204, 407) the ML model training based on the result of the evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYNCHRONIZING MACHINE LEARNING MODEL TRAINING

[0002] TECHNICAL FIELD

[0003] The disclosure relates to a system comprising a network node and a training node, methods performed by the network node and the training node, computer programs and computer program products comprising the computer programs, to synchronize machine learning model training in a communication network.

[0004] BACKGROUND

[0005] 3rd generation partnership project system architecture and services, 3GPP SA2, has specified a network function, NF, called network data analytics function, NWDAF, in 5thgeneration mobile network core, 5GC, responsible for optimized data collection, training and deploying machine learning, ML, models. During the training, the model training logical function, MTLF, comprised in the NWDAF can get the data required by the training process from various data sources as specified in 3GPP TS23.288 version 18.5.0 or version 18.7.0. NWDAF supports training architectures such as centralized training or distributed training via horizontal federated learning (support of vertical federated learning can be added in Rel19).

[0006] In all these training architectures, data collection and training process can be very time and energy consuming, especially when large ML models are trained, or many participants are required in distributed training cases.

[0007] WO 2023 / 227937 A1 discloses an analytics service training pipeline, ASTP, defining a pipeline structure of an analytics service implementation, and providing information about the training and re-training of ML models associated with the analytics service. The ASTP describes the data stages within the analytics (sub)service implementation, and how they apply to ML model (re-)training operations.

[0008] Training machine learning models, particularly in a distributed environment such as communication network, involves potentially numerous parties and substantial consumption of hardware resources such as memory, compute power, specialized hardware, and communication resources. This resource consumption is even larger when multiple parties are engaged not only in the actual training process but also in gathering training data and executing necessary data transformations. Consequently, the number of locked resources becomes substantial. To mitigate resource wastage and enhance energy efficiency, these resources should be promptly released upon the completion, termination, or interruption of the training process. Although resource release is currently managed by means of timers, this approach adds complexity to the model training systems and may lead to further inefficiencies. When timers expire, participants might inadvertently clear the current state of the optimization algorithm (such as some learning rate or batch size that depends on the iteration number and how close we are to the optimal solution), even if the training operations could be resumed later, this leads to prolonged convergence time upon resuming the process.

[0009] SUMMARY

[0010] The object of the invention is to enable an efficient handling of synchronization on machine leaning model trainings and improve resource utilization associated with training of a machine learning model in a communication network.

[0011] A first aspect of the invention is a method for synchronizing ML model training in the communication network. The method performed by a network node. The method comprises obtaining, from a training party involved in the training of the ML model in the communication network, a synchronization request comprising a first context indicating a need for synchronizing the ML model training. The method comprises evaluating the first context comprised in the synchronization request. The method comprises notifying the training party a result of the evaluation. The method comprises configuring the ML model training based on the result of the evaluation. Hereby, the invention enables synchronizing ML model training in the communication network by means of the synchronization request and the first context indicating the need for synchronizing the ML model training.

[0012] A second aspect of the invention is a method for synchronizing ML model training the communication network. The method performed by a training node in the communication network. The method comprises transmitting, to the network node in the communication network, the synchronization request comprising the first context indicating the need for synchronizing the ML model training. The method comprises receiving, from the network node, the notification indicating the result of the evaluation of the ML model training based on the first context indicating the need for synchronizing the ML model training. The method comprises configuring the ML model training based on the result of the evaluation. A third aspect of the invention is the network node configured to synchronize machine learning, ML, model training in the communication network, the network node comprising processing circuitry and a memory , the memory containing instructions executable by the processing circuitry such that the network node is operable to obtain, from the training party involved in the ML model training in the communication network, the synchronization request comprising the first context indicating the need for synchronizing the ML model training. The network node is operable to evaluate the first context comprised in the synchronization request. The network node is operable to notify the training party a result of the evaluation. The network node is operable to configure the ML model training based on the result of the evaluation.

[0013] A fourth aspect of the invention is the training node configured to synchronize ML model training in the communication network. The training node comprising processing circuitry and a memory, the memory containing instructions executable by the processing circuitry such that the network node is operable to transmit, to the network node in the communication network, a synchronization request comprising a first context indicating a need for synchronizing the ML model training. The network node is operable to receive, from the network node, the notification indicating the result of the evaluation of the ML model training based on the first context indicating the need for synchronizing the ML model training. The network node is operable to configure the ML model training based on the result of the evaluation.

[0014] According to a fifth aspect of the invention, there is presented a computer program comprising instructions which when executed on a processor of the network node, causes the network node to perform a method according to the first aspect or any of the embodiments of the first aspect.

[0015] According to a sixth aspect of the invention, there is presented a computer program product which comprises a computer readable storage medium on which a computer program according to the fifth aspect is stored.

[0016] According to a seventh aspect of the invention, there is presented a computer program comprising instructions which when executed on a processor of the training node, causes the training node to perform a method according to the second aspect or any of the embodiments of the second aspect. According to an eight aspect of the invention, there is presented a computer program product which comprises a computer readable storage medium on which a computer program according to the seventh aspect is stored.

[0017] Other objectives, features and advantages of the enclosed embodiments will be apparent from the following detailed disclosure, from the attached dependent claims as well as from the drawings.

[0018] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise herein. All references to "a / an / the element, apparatus, component, means, module, action, etc." are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, module, action, etc., unless explicitly stated otherwise. The actions of any method disclosed herein.

[0019] BRIEF DESCRIPTION OF THE DRAWINGS

[0020] A more complete understanding of the present embodiments, and the attendant advantages and features thereof, will be more readily understood by reference to the following detailed description when considered in conjunction with the accompanying drawings wherein:

[0021] Figure 1 depicts a system according to one or more aspects of the invention.

[0022] Figure 2 illustrates a method according to a first aspect of the invention.

[0023] Figure 3 illustrates a method according to a second aspect of the invention.

[0024] Figure 4 illustrates a signaling diagram according to one or more aspects and one or more embodiments of the invention.

[0025] Figure 5 illustrates an example of a network node.

[0026] Figure 6 illustrates an example of a training node.

[0027] DETAILED DESCRIPTION

[0028] The invention discloses system and methods for adding new signals to a procedure of machine learning, ML, model training in a communication network. Note that all references herein to an ML model refer broadly to any kind of artificial intelligence AI / ML model, regardless of whether traditional ML techniques or more advanced Al approaches are used in its creation, training or execution. The ML model training can be distributed or centralized, in either training architectures, data collection and training process can be time and energy consuming, especially when large ML models are trained, or many participants are required for the training (e.g., in distributed training cases). In such scenarios, it is desirable to be able to pause or stop the training process if some of the participants or data sources have higher priority tasks to handle and resume the process later. It is also desirable to synchronize the training process amongst the participants of the training process. When the participants are communicated in advance regarding the pausing or stopping of the training process, that is, when the participants are in sync regarding the training process, the participants can temporarily release the memory and processing resources allocated for the training and use them for more urgent services. The invention discloses adding new signals to a procedure of ML model training in the communication network in order to synchronize the training process and to provide advance indication regarding the interruptions to the training procedure to the involved participants of the training. Synchronization of the training process by providing advance indication regarding the interruption of the training process enables the mitigation of resource wastage and enhance energy efficiency, by prompt release of resources used for the training process.

[0029] Figure 1 depicts a system according to one or more aspects of the invention. System 100 comprises a communication network 101 , a service consumer node 102, a network node 103 and a training node 104. The service consumer node 102, the network node 103 and the training node 104 are configured to operate in the communication network 101. The service consumer node 102, the network node 103 and the training node 104 communicate with each other by means of the communication network 101. The communication network 101 as illustrated in Figure 1 enables connectivity between user equipment, UE, and communication network nodes, such as service consumer node 102, network node 103 and the training node 104. In that sense, the communication network 100 may be a 3rd Generation Partnership Project, 3GPP, network and may be configured to operate according to predefined rules or procedures, such as specific standards that comprise, but are not limited to: Global System for Mobile Communications, GSM; Universal Mobile Telecommunications System, UMTS; Evolved Packet System, EPS, 5th Generation System, 5GS, or any applicable future generation standard (e.g.,3GPP 6G). The communication network 101 here comprises one or more radio access network, RAN, nodes, and one or more core network nodes of an Evolved Packet Core, EPC, or an 5th Generation Core, 5GC or of any applicable future generation standard (e.g.,3GPP 6G). A RAN node may comprise a base station, eNodeB of an LTE network, gNodeB of a New Radio, NR, network, or any other current or future implementation of functionality facilitating the exchange of radio network signals between nodes of the communication network 100 and / or UE. Moreover, as will be appreciated by those of skill in the art, a network node is not necessarily limited to an implementation in which a radio portion and a baseband portion are supplied and integrated by a single vendor. Thus, it will be understood that network nodes include disaggregated implementations or portions thereof. For example, in some embodiments, the communication network 100 comprises one or more Open-RAN, ORAN, network nodes. An ORAN network node is a node in the communication network 100 that supports an ORAN specification (e.g., a specification published by the O-RAN Alliance, or any similar organization) and may operate alone or together with other nodes to implement one or more functionalities of any node in the communication network, including one or more access network nodes and / or core network nodes.

[0030] The service consumer 102 is a network node in the communication network 101 subscribing to a service offered by the network node 103. The service consumer 102 subscribes to a service associated with the ML model. The service consumer 102 can be a core network node in the communication network 101 . The service consumer 102 can be a RAN network node in the communication network 101 . The service consumer 102 can be an ORAN node in the communication network 101 . The service consumer 102 can be a UE in the communication network 101.

[0031] The network node 103 is responsible for training the ML model to offer the service associated with the ML model to the service consumer 102. The network node 103 monitors the one or more training nodes performing the ML model training. The network node 103 may also perform ML model training. The network node 103 can be a core network node in the communication network 101. The network node 103 comprises a Network Function, NF, responsible for collecting data required for training the ML model, training the ML model and deploying ML model in the communication network 101 . The network node 103 may comprise a NF called network data analytics function, NWDAF. The network node 103 can be a core network node in 5thgeneration core, 5GC, network and comprises NWDAF in 5GC. NWDAF is responsible for data collection, training and inference [3GPP TS23.288 version 18.5.0 or 18.7.0], 3GPP SA2 has specified the NWDAF in 5GC, responsible for optimized data collection, training and deploying ML models. During the training, the Model Training Logical Function, MTLF, comprised in NWDAF obtains the data required by the training process from various data sources specified in [3GPP TS23.288 version 18.5.0 or 18.7.0], NWDAF supports training on a central location or distributed training via horizontal federated learning, HFL, (support of vertical federated learning, VFL, will be added in Rel19). The network node 103 can be a Service Management and Orchestration, SMO, node in ORAN. The network node 103 can be a non-real-time RAN Intelligent controller, Non-RT RIC, in ORAN.

[0032] The training node 104 is a network node in the communication network 101 performing data collection for training the ML model and / or performing the ML model training. The training node 104 can also be a source for data collection in the communication network 101 . The training node 104 can be a core network node in the communication network 101 . The training node 104 may comprise Access and Mobility Function, AMF. The training node 104 may comprise Session Management Function, SMF. The training node 104 may comprise Analytics Data Repository Function, ADRF. The training node 104 may comprise MLTF capabilities. The training node 104 may comprise intent management function, IMF. The training node 104 can be a RAN node in the communication network 101. The training node 104 can be a UE in the communication network 101. There can be one or more training nodes 104 in the communication network 101 performing the ML model training. ML model training also comprises collecting data for training the ML model in the communication network 101 .The communication interfaces between the network node 103 and the one or more training nodes 104 can be implemented using RESTful APIs over HTTP / 2, utilizing the specified prefixes and service names as mentioned in [3GPP TS 23.288 v18.5.0 or v18.7.0], Detailed services and service names are specified for NWDAF in [3GPP TS 23.288 v18.5.0, Section 7] and for ADRF in [3GPP TS 23.288 v18.5.0 Section 10] of the same specification. Communication with the ADRF may be necessary if the training node 104 is a data provider or holds datasets within the ADRF. Similarly, if communication with other network functions such as UDM is required for data collection, the communication can be handled the same way using proper service prefix.

[0033] For example, the service consumer 102 subscribes to a service such as Quality of Experience, QoE, estimation in the communication network 101. The service consumer 102 subscribes to QoE estimation service offered by the network node 103. The network node 103 is responsible for training the ML model to perform QoE estimation. The service associated with the ML model is QoE estimation. Training the ML model to estimate QoE is performed by one or more training nodes 104. The network node 103 is responsible for coordinating the one or more training nodes 104 to perform the training.

[0034] Figure 2 illustrates a method 200 according to a first aspect of the invention. The method 200 performed by the network node 103 for synchronizing ML model training in the communication network 101 . The ML model training can be a centralized training or a distributed training in the communication network 101. The ML model training is performed by one or more training nodes in the communication network 101 . The ML model training is monitored by the network node 103. The method 200 comprises obtaining 201 , from a training party involved in the ML model training in the communication network 101 , a synchronization request comprising a first context indicating a need for synchronizing the ML model training. The training party comprises the network node 103 and / or the training node 104.

[0035] The first context may comprise one or more events experienced by the network node

[0036] 103 and / or, by one or more training nodes 104 that may affect the training of the ML model. The first context may comprise one or more events related to the training process. The event related to the training process may indicate an issue affecting the training process. The event related to the training process may indicate resolution of an issue affecting the training process. The synchronization request comprising the first context is an advance indication to the involved participants of the ML model training regarding the interruptions to the ML model training due to the events experienced by other participants of the ML model training.

[0037] Synchronizing ML model training in the communication network 101 comprises adapting a way in which the network node 103 and / or the one or more training nodes

[0038] 104 train the ML model based on the first context. The method 200 comprises evaluating 202 the first context comprised in the synchronization request. Evaluation 202 comprises the network node 103 determining whether the need indicated by the first context is an acceptable need for one or more training nodes 104 and / or network node 103 to synchronize the ML model training. The network node 103, upon obtaining the first context, performs an evaluation of the ML model training corresponding to the first context. The method comprises notifying 203 the training party a result of the evaluation. The method comprises configuring 204 the ML model training based on the result of the evaluation.

[0039] In a first embodiment of the invention, the training party comprises the network node 103. The network node 103 monitors the one or more training nodes during the ML model training and detects one or more of the following example events:

[0040] • persistent problem in convergence of the ML model training,

[0041] • an unusual / abnormal model updates during training performed by one or more training nodes,

[0042] • internal exceptions in the network node 103 or in the one or more training nodes,

[0043] • malicious activities by one or more training nodes,

[0044] • overload experienced by one or more training nodes or the network node,

[0045] • low efficiency,

[0046] • waiting for straggler training nodes to catch up in case of synchronous model update during training,

[0047] • needing more time to analyze logs and decide on which training nodes to be involved in the next ML model training cycle,

[0048] • unavailability of one or more training nodes,

[0049] • resolution of persistent problem in convergence of the ML model training,

[0050] • resolved unusual / abnormal model updates during training performed by one or more training nodes

[0051] • resolved internal exceptions in the network node 103 or in the one or more training nodes 104,

[0052] • resolved malicious activities by one or more training nodes,

[0053] • resolved overload experienced by one or more training nodes or the network node, • resolved low efficiency,

[0054] • resolved waiting for straggler training nodes to catch up in case of synchronous model update during training,

[0055] • resolved needing more time to analyze logs and decide on which training nodes to be involved in the next ML model training cycle,

[0056] • resolved unavailability of one or more training nodes.

[0057] The first context comprises one or more of the above detected events, thus, indicating the need for synchronizing the ML model training. The first context is not restricted to the above-mentioned events but can comprise any event affecting the training of the ML model in the communication network or security events related to the ML model training.

[0058] In the first embodiment, the first context may be specific to the one or more training nodes or specific to the network node 103. The events comprised in the first context is detected by the network node 103.

[0059] In a second embodiment of the invention, the training party comprises the training node 104. The training node 104 detects one or more of the following example events:

[0060] • some exceptions or internal errors,

[0061] • getting temporarily overloaded with analytic requests (a phenomenon also known as “signaling storm”),

[0062] • hyper-parameter inconsistency and problem in local model update,

[0063] • data unavailability (in terms of volume of data to train, or in terms of data quality),

[0064] • higher priority task to be attended,

[0065] • waiting for human feedback to fine-tune the ML model,

[0066] • exceeding staleness threshold in distributed training,

[0067] • resolved exceptions or internal errors,

[0068] • resolved signaling storm,

[0069] • resolved hyper-parameter inconsistency and problem in local model update,

[0070] • resolved data unavailability (in terms of volume of data to train, or in terms of data quality),

[0071] • resolved task priority issues, resolved waiting for human feedback to fine-tune the ML model, resolved exceeding staleness threshold in distributed training.

[0072] The first context, according to the second embodiment, comprises one or more of the above events detected by the training node 104, thus, indicating the need for synchronizing the ML model training. The first context is not restricted to the above- mentioned events but can comprise any event affecting the training of the ML model in the communication network or security events related to the ML model training. In the second embodiment, the first context may be specific to the training node 104. The events comprised in the first context is detected by the training node 104.

[0073] Figure 3 illustrates a method 300 according to a second aspect of the invention. The method 300 performed by the training node 104 for synchronizing ML model training in the communication network 101. The method 300 comprising transmitting, to the network node 103, the synchronization request comprising the first context indicating the need for synchronizing the ML model training. The method 300 comprising receiving 302, from the network node 103, the notification indicating the result of the evaluation of the ML model training based on the first context indicating the need for synchronizing the ML model training. The method 300 comprising configuring the ML model training based on the result of the evaluation.

[0074] The first context according to the second aspect may comprise more of the following example events detected by the training node 104:

[0075] • some exceptions or internal errors,

[0076] • getting temporarily overloaded with analytic requests (a phenomenon also known as “signaling storm”),

[0077] • hyper-parameter inconsistency and problem in local model update,

[0078] • data unavailability (in terms of volume of data to train, or in terms of data quality),

[0079] • higher priority task to be attended,

[0080] • waiting for human feedback to fine-tune the ML model,

[0081] • exceeding staleness threshold in distributed training,

[0082] • resolved exceptions or internal errors,

[0083] • resolved signaling storm

[0084] • resolved hyper-parameter inconsistency and problem in local model update, • resolved data unavailability (in terms of volume of data to train, or in terms of data quality),

[0085] • resolved task priority issues,

[0086] • resolved waiting for human feedback to fine-tune the ML model,

[0087] • resolved exceeding staleness threshold in distributed training.

[0088] The first context, according to the second aspect, comprises one or more of the above events detected by the training node 104, thus, indicating the need for synchronizing the ML model training. The first context is not restricted to the above-mentioned events but can comprise any event affecting the training of the ML model in the communication network or security events related to the ML model training. In the second aspect, the first context may be specific to the training node 104. The events comprised in the first context is detected by the training node 104.

[0089] In a third embodiment of the invention, the configuring 204 comprises, the network node 103 notifying a second context to each training node in the one or more training nodes performing the ML model training. The second context comprises a modification to the ML model training based on the first context and the result of the evaluation. The second context can be specific to the training node. The second context can be specific to the network node 103. The second context may indicate an action to be performed by the training node 104. The second context may indicate an action to be performed by the network node 103.

[0090] According to the third embodiment of the invention, the configuring 303 comprises the training node 104 receiving the second context from the network node 103.

[0091] For example, the second context comprises one or more of:

[0092] • removing training node from the ML model training,

[0093] • adding training node to the ML model training,

[0094] • removing some features of the ML model training,

[0095] • updating parameters of the ML model training,

[0096] • performing a recovery process from a certain epoch of the ML model training and restart from the certain epoch,

[0097] • changing training node selection policy. The second context is not restricted to the above-mentioned examples but can comprise any modification to the training of the ML model in the communication network based on the result of the evaluation and the first context.

[0098] According to the first, second aspects of the invention and the first, second and third embodiments of the invention the synchronization request comprises one or more of a pause request, a stop request, or a resume request. The pause request can be implemented using RESTful APIs over HTTP / 2, utilizing the specified prefixes and service names as mentioned in [3GPP TS 23.288 v18.5.0], The stop request can be implemented using RESTful APIs over HTTP / 2, utilizing the specified prefixes and service names as mentioned in [3GPP TS 23.288 v18.5.0], The resume request can be implemented using RESTful APIs over HTTP / 2, utilizing the specified prefixes and service names as mentioned in [3GPP TS 23.288 v18.5.0],

[0099] According to the first, second aspects of the invention and the first, second and third embodiments of the invention the result of the evaluation comprises one or more of to pause the ML model training, to stop the ML model training, to resume the ML model training, or to continue the ML model training.

[0100] According to the first, second aspects of the invention and the first, second and third embodiments of the invention the configuring 204, 303 the ML model training comprises performing one of pausing the ML model training, stopping the ML model training, resuming the ML model training, or continuing the ML model training. Configuring 204, 303 the ML model training comprises modifying the ML model training based on the first context and / or the second context when the when the result of the evaluation is to pause or stop the ML model training.. Modifying the ML model training comprises removing or adding one or more training parties. Modifying the ML model training comprises removing or adding one or more features of the ML model training. Modifying the ML model training comprises changing one or more parameters associated with the ML model training. Modifying the ML model training comprises initiating a recovery process of one or more training parties. Modifying the ML model training comprises changing a training party selection policy. Modifying the ML model training comprises a combination of thereof.

[0101] According to the first, second aspects of the invention and the first, second and third embodiments of the invention the second context comprises one or more of a context associated with pausing the ML model training, a context associated with stopping the ML model training, or a context associated with resuming the ML model training.

[0102] Figure 4 illustrates a signaling diagram according to one or more aspects and one or more embodiments of the invention. The signaling between the service consumer 102 and the network node 103 can be implemented using RESTful APIs over HTTP / 2, utilizing the specified prefixes and service names as mentioned in [3GPP TS 23.288 v18.5.0], The signaling between the training node 104 and the network node 103 can be implemented using RESTful APIs over HTTP / 2, utilizing the specified prefixes and service names as mentioned in [3GPP TS 23.288 v18.5.0], The service consumer may transmit 401 a subscription request to the network node 103, wherein the subscription request indicating that the service consumer is subscribing to a service, such as QoE estimation, associated with the ML model, offered by the network node 103. In this case, the subscription request comprises Nwdaf_MLModelTraining_Subscribe and the subscription procedure follows the procedure outlined in [3GPP TS 23.288 v18.5.0, Clause 6.2F], with at least the contents specified in [Clause 6.2F.1 ] of the same document.The network node 103 monitors the one or more training nodes 104 performing the ML model training to offer the service associated with the ML model to the service consumer 102. The network node 103 obtains 402a the synchronization request comprising the first context indicating a need for synchronizing the ML model training from the network node 103. The network node 103 obtains 402b, from the training node 104, the synchronization request comprising the first context indicating a need for synchronizing the ML model training. The network node may obtain the synchronization request from the network node 103 and / or the training node 104. The first context may comprise one or more events related to the training process. The event related to the training process may indicate an issue affecting the training process. The event related to the training process may indicate resolution of an issue affecting the training process.

[0103] The synchronization request may comprise the pause request. The pause request indicates pausing the training process carried out by the network node 103 and one or more training nodes 104. The pause request can be invoked by any of the training participants- the network node 103 and one or more training nodes 104. The first context comprises one or more events according to the first and second embodiments, indicating a need for invoking the pause request. The pause request may be invoked by the network node 103 due to one or more events, according to the first embodiment, such as, internal exceptions in the network node 103 or malicious activities by one or more training nodes, comprised in the first context, in the training process detected by the network node 103. The pause request may be invoked by the training node 104 due to one or more events according to the second embodiment, such as, hyperparameter inconsistency and problem in local model update or data unavailability, comprised in the first context, in the training process detected by the training node 104. The network node 103 may obtain 402a or 402b the pause request from the network node 103 and / or the training node 104, respectively. The pause request may also comprise a duration of pause required by the training participant invoking the pause request. The pause request may also comprise the duration of pause required by the network node 103 based on the one or more events according to the first embodiment, such as, internal exceptions in the network node 103 or malicious activities by one or more training nodes, comprised in the first context, in the training process detected by the network node 103. In the context of the network node 103 or the training node 104, internal exceptions refer to unexpected events or errors that occur within the node's software or hardware, disrupting its normal operation. Malicious attacks may comprise training nodes providing incorrect or misleading data to the training process. This can degrade the model's performance or cause it to behave unpredictable Malicious attacks may comprise attackers directly modifying the model updates they contribute. This can subtly alter the model's decision boundaries to produce biased or incorrect outputs. The pause request may also comprise the duration of pause required by the training node 104 based on the one or more events, according to the second embodiment, such as, hyper-parameter inconsistency and problem in local model update or data unavailability, comprised in the first context, in the training process detected by the training node 104. The pause request can be periodical. The training node 104 may invoke the pause request periodically. The network node 103 may invoke the pause request periodically. The pause request can be periodical, and the pause request may comprise a pause frequency indicating a frequency with which the training process has to be paused. The pause request can be time bounded, that is, the pause request will expire after a pause request expiry time. If the network node 103, which obtains the pause request, does not attend to the pause request, the pause request will expire after the pause expiry time. The pause expiry time can be indicated by the pause request. The pause expiry time can be a value of time agreed upon between the network node 103 and the one or more training nodes 104.

[0104] For example, the network node 103 obtains the pause request comprising the first context as waiting for human feedback to fine-tune the ML model from the training node 104. The pause request may additionally comprise the pause duration, such as, five hours, indicating the network node 103 to pause the training process for five hours while the training node 104 waits for the human feedback to fine-tune the ML model. The pause request may additionally comprise the pause frequency, such as, five hours, indicating the network node 103 to pause the training process for every five hours while the training node 104 waits for human feedback to fine-tune the ML model. The pause request may additionally comprise the pause request expiry time, such as, five hours, indicating that the pause request will expire after five hours if the network node 103 does not attend to the pause request.

[0105] The synchronization request may comprise the stop request. The stop request indicates stopping the training process carried out by the network node 103 and one or more training nodes 104. The stop request can be invoked by any of the training participants- the network node 103 and one or more training nodes 104. The first context comprises one or more events according to the first and second embodiments, indicating a need for invoking the stop request. The stop request may be invoked by the network node 103 due to one or more events according to the first embodiment, such as, malicious activities by one or more training nodes or persistent problem in convergence of the ML model training, comprised in the first context, in the training process detected by the network node 103. The stop request may be invoked by the training node 104 due to one or more events according to the second embodiment, such as, experiencing signaling storm or having a higher priority task to attend to, comprised in the first context, in the training process detected by the training node 104. The network node may obtain 402a or 402b the stop request from the network node 103 and / or the training node 104, respectively. The stop request may also comprise a duration of stop or stop duration required by the training participant invoking the stop request. The stop request may also comprise a duration of stop or stop duration required by the network node 103 based on the one or more events, according to the first embodiment, comprised in the first context, in the training process detected by the network node 103. The stop request may also comprise a duration of stop or stop duration required by the training node 104 based on the one or more events, according to the second embodiment, comprised in the first context, in the training process detected by the training node 104.

[0106] The stop request can be periodical. The training node 104 may invoke the stop request periodically. The network node 103 may invoke the stop request periodically. The stop request can be periodical, and the stop request may comprise a stop frequency indicating a frequency with which the training process has to be stopped. The stop request can be time bounded, that is, the stop request will expire after a stop request expiry time. If the network node 103, which obtains the stop request, does not attend to the stop request, the stop request will expire after the stop expiry time. The stop expiry time can be indicated by the stop request. The stop expiry time can be a value of time agreed upon between the network node 103 and the one or more training nodes 104.

[0107] For example, the network node 103 obtains the stop request comprising the first context as waiting for human feedback to fine-tune the ML model from the training node 104. The stop request may additionally comprise the stop duration, such as, five hours, indicating the network node 103 to stop the training process for five hours while the training node 104 waits for the human feedback to fine-tune the ML model. The stop request may additionally comprise the stop frequency, such as, five hours, indicating the network node 103 to stop the training process for every five hours while the training node 104 waits for human feedback to fine-tune the ML model. The stop request may additionally comprise the stop request expiry time, such as, five hours, indicating that the stop request will expire after five hours if the network node 103 does not attend to the stop request.

[0108] The synchronization request may comprise the resume request. The resume request indicates resuming the training process carried out by the network node 103 and one or more training nodes 104. The resume request can be invoked by any of the training participants- the network node 103 and one or more training nodes 104. The first context comprises one or more events according to the first and second embodiments, indicating a need for invoking the resume request. The resume request may be invoked by the network node 103 due to one or more events according to the first embodiment, such as, resolved exceptions or internal errors, or resolved signaling storm comprised in the first context, detected by the network node 103 in the training process. The resume request may be invoked by the training node 104 due to one or more events according to the second embodiment, such as, resolved exceptions or internal errors or resolved signaling storm, comprised in the first context, detected by the training node 104 in the training process. The network node may obtain 402a or 402b the resume request from the network node 103 and / or the training node 104. The resume request can be invoked by the network node 103 or the training node 104 if the network node 103 or the training node 104 determines that the training process has been paused or stopped. The resume request can be invoked by the network node 103 or the training node 104 if the network node 103 or the training node 104 determines that the training process has been paused for a certain time period. The resume request can be invoked by the network node 103 or the training node 104 upon resolution of one or more events that caused the pausing or stopping of the training process.

[0109] Upon obtaining the synchronization request comprising the first context, the network node 103 evaluates 403 the first context comprised in the synchronization request. Evaluation 403 comprises the network node 103 determining whether the need indicated by the first context is an acceptable need for one or more training nodes 104 and / or network node 103 to synchronize the ML model training. The network node 103 may perform the evaluation by individually communicating with each training node of the one or more training nodes, to determine if the need indicated by the first context is acceptable for the training node 104 to synchronize the ML model training. The network node 103 may maintain a list of acceptable needs for synchronizing the ML model training and during evaluation 403, the network node 103 may determine if the need indicated by the first context is comprised in the list of acceptable needs maintained by the network node 103.

[0110] For example, the first context comprises one or more training nodes 104 or the network node 103 experiencing signaling storm according to the first embodiment or the first context comprises the one or more training nodes 104 experiencing signaling storm according to the second embodiment. The network node 103 or the one or more training nodes 104 experiencing signaling storm indicates the need for synchronizing, that is, stopping or pausing the ML model training. The network node 103 evaluates whether this need, experiencing signaling storm, indicated by the first context is an acceptable need for one or more training nodes 104 and / or network node 103 to pause or stop the ML model training. The network node 103 may perform the evaluation by individually communicating with each training node of the one or more training nodes, to determine if the network node 103 or the one or more training nodes experiencing signaling storm is acceptable for the training node 104 to stop or pause the ML model training. The network node 103 may maintain a list of acceptable needs for synchronizing the ML model training and during evaluation 403, the network node 103 may determine if the network node 103 or the one or more training nodes experiencing signaling storm is comprised in the list of acceptable needs maintained by the network node 103. If the evaluation indicates that the network node 103 or the one or more training nodes experiencing signaling storm is an acceptable need for stopping or pausing the ML model training, then the result of the evaluation is to stop or pause the ML model training. If the evaluation indicates that the network node 103 or the one or more training nodes experiencing signaling storm is not an acceptable need for stopping or pausing the ML model training, then the result of the evaluation is to continue the ML model training.

[0111] For example, the first context comprises an exception experienced by the network node 103 according to the first embodiment. The network node 103 experiencing exceptions indicates the need for synchronizing the ML model training. The network node 103 evaluates whether this need, experiencing exception, indicated by the first context is an acceptable need for one or more training nodes 104 to pause or stop the ML model training. The network node 103 may perform the evaluation by individually communicating with each training node of the one or more training nodes, to determine if the network node 103 experiencing exception is acceptable for the training node 104 to stop or pause the ML model training. The network node 103 may maintain a list of acceptable needs for synchronizing the ML model training and during evaluation 403, the network node 103 may determine if the network node 103 experiencing exception is comprised in the list of acceptable needs maintained by the network node 103. If the evaluation indicates that the network node 103 experiencing exception is an acceptable need for stopping or pausing the ML model training, then the result of the evaluation is to stop or pause the ML model training. If the evaluation indicates that the network node 103 experiencing exception is not an acceptable need for stopping or pausing the ML model training, then the result of the evaluation is to continue the ML model training without stopping or pausing. In this example, the evaluation may indicate that the network node 103 experiencing exception is an acceptable need for stopping or passing the ML model training and the result of the evaluation is to the stop or pause the ML model training. If the evaluation indicates that the network node 103 or the one or more training nodes experiencing signaling storm is not an acceptable need for stopping or pausing the ML model training, then the result of the evaluation is to continue the ML model training. To pause or stop the ML model training depends on a severity of the event detected by the network node 103 (according to the first embodiment) or the training node (according to the second embodiment). The network node 103 and the one or more training nodes 104 may collectively agree upon the severity of the events. Upon agreeing the network node 103 and the one or more training nodes 104 may determine to stop the ML model training for some events that have high severity than for other events that have low severity. Upon agreeing, the network node 103 and the one or more training nodes 104 may determine to pause the ML model training for events that have low severity.

[0112] In the case of synchronization request comprising the resume request, the first context, for example, comprises resolved signaling storm according to the first or the second embodiment. The network node 103 evaluates whether this need, resolved signaling storm, indicated by the first context is an acceptable need for one or more training nodes 104 to resume the ML model training. The network node 103 may perform the evaluation by individually communicating with each training node of the one or more training nodes, to determine if resolved signaling storm according to the first or the second embodiment is an acceptable need for the training node 104 to resume the ML model training. The network node 103 may maintain a list of acceptable needs for synchronizing the ML model training and during evaluation 403, the network node 103 may determine if resolved signaling storm according to the first or the second embodiment is comprised in the list of acceptable needs maintained by the network node 103. If the evaluation indicates that the network node 103 resolved signaling storm is an acceptable need for resuming the ML model training, then the result of the evaluation is to resume the ML model training. If the evaluation indicates that the network node 103 resolved signaling storm is not an acceptable need for resuming the ML model training, then the result of the evaluation is to maintain the ML model training in either stopped or paused state. In step 404, the result of the evaluation is notified by the network node 103 to the training parties, that is, the one or more training nodes.

[0113] In some embodiments, the result of the evaluation is notified 405 by the network node 103 to the service consumer 102 when the result of the evaluation comprises one of: to pause the ML model training, to stop the ML model training, or to resume the ML model training.

[0114] Service consumers 102 are typically unaware of model training operations, as they interact with an Analytics Logical Function, AnLF, of NWDAF comprised in the network node 103. AnLF is a component of the NWDAF that performs data analytics and exposes analytics service(s) to other NFs through the Nnwdaf_AnalyticsSubscription and Nnwdaf_Analyticslnfo services [3GPP TS23.288],

[0115] Upon subscribing to receive analytics reports through Nnwdaf_AnalyticsSubscription_Subscribe request, service consumers 102 are expected to deploy appropriate resources to handle received analytics reports, such as deploying several replicas of the microservices responsible for processing analytics reports, providing storage capabilities for storing the received analytics reports and computing ML model metrics. If the AnLF is unable to provide analytics reports because the ML model is not yet available (e.g., training has not started or has been paused / stopped), service consumers 102 simply wait for the analytics report to be received asynchronously whenever the ML model is ready. It is not sensible to deploy computing and storage resources if they are not necessary, as resource deployment comes with a cost. To prevent this the result of the evaluation is notified in step 405 as a flag in a response to the Nnwdaf_AnalyticsSubscription_Subscribe request to inform the service consumer 102 about the unavailability of the ML model, either because model training has not started or is paused. This approach allows the service consumer 102 to be informed about the availability of the ML model and to adjust to the ML model's status without wasting resources.

[0116] In step 407, the network node 103 configures the ML model training. Configuring 407 comprises performing one of: pausing the ML model training, stopping the ML model training, resuming the ML model training, or continuing the ML model training. Configuring the ML model training further comprises notifying the second context to each training node in one or more training nodes 104 performing the ML model training, wherein the second context comprises the modification to the ML model training based on the first context and the result of the evaluation. For example, the second context comprises the modifications that the network node 103 wants to make to the ML model training based on the first context and the result of the evaluation.

[0117] When the result of the evaluation is to pause the ML model training, in step 407, the network node 103 pauses the ML model training. In step 407, the training node 104 pauses the ML model training.

[0118] When the result of the evaluation is to stop the ML model training, in step 407, the network node 103 stops the ML model training. In step 407, the training node 104 stops the ML model training.

[0119] When the result of the evaluation is to resume the ML model training, in step 407, the network node 103 resumes the ML model training. In step 407, the training node 104 resumes the ML model training.

[0120] When the result of the evaluation is to continue the ML model training, in step 407, the network node 103 continues the ML model training. In step 407, the training node 104 continues the ML model training.

[0121] Configuring 407 the ML model training may comprise modifying the ML model training based on the first context and / or the second context when the synchronization request comprises one of the pause or the stop requests.

[0122] For example, the first context comprises unavailability of a training node according to the first embodiment. The result of the evaluation is to pause or stop the ML model training. Configuring 407 comprises either pausing the ML model training or stopping the ML model training. The second context comprises removing the unavailable training node. Configuring 407 the ML model training comprises modifying the ML model training based on the first context and / or the second context. In this case, configuring 407 the ML model training comprises pausing or stopping the ML model training and removing the unavailable training node from the ML model training as per the second context, as the first context indicated the unavailability of the training node. In step 407, the network node 103 stops or pauses the ML model training and removes the unavailable training node from the ML model training as per the second context, as the first context indicated the unavailability of the training node. In step 407, the training node 104 stops or pauses the ML model training and is notified of the removal of the unavailable training node.

[0123] For example, the first context comprises hyper-parameter inconsistency according to the second embodiment. The result of the evaluation is to pause or stop the ML model training. Configuring 407 comprises either pausing the ML model training or stopping the ML model training. The second context comprises updating parameters of the ML model training. Configuring 407 the ML model training comprises modifying the ML model training based on the first context and / or the second context. In this case, configuring 407 the ML model training comprises pausing or stopping the ML model training and updating parameters of the ML model training as per the second context, as the first context indicated the inconsistency in the parameters. In step 407, the network node 103 stops or pauses the ML model training and updates parameters of the ML model training as per the second context, as the first context indicated the inconsistency in the parameters. In step 407, the training node 104 stops or pauses the ML model training and updates parameters of the ML model training as per the second context, as the first context indicated the inconsistency in the parameters.

[0124] In some embodiments, when the synchronization request comprises pause or stop request, the training process is resumed, after the ML model training is modified based on the first and / or the second context, by the network node 103 and / or by the one or more training nodes 104.

[0125] In a fourth embodiment of the invention, the second context can be provided by a recommendation system by analyzing logs of at least one of: previous synchronization requests, first context, second context and corresponding modifications to the ML model training. Such recommendation system can be realized via, for example, a multi-armed bandit that takes as input one or a combination of: a) the synchronization request, b) first context comprised in the synchronization request, c) a vector or matrix comprising similarity scores indicating similarity two or more training nodes 104, that is, between one training node and another training node. The vector or matrix comprising similarity scores is inputted for a possibility of a training node being replaced by another training node. Creating such vector / matrix is outside the scope of this invention. However, various approaches such as similarity in inputs to the training nodes can be used.

[0126] In a fifth embodiment of the first and second aspects, the synchronization request may comprise one or more conditions for synchronizing the ML model training. The ML model training can be configured 407 when the condition is satisfied. The synchronization request can be obtained by the network node 103 when such conditions for synchronizing the ML model training are met. For example, the condition can be a Key Performance Indicator, KPI, of the communication network 101 with some threshold. The condition can be the KPI of the communication network 101 exceeding a threshold. The network node 101 may perform pausing or stopping or resuming the ML model training when the KPI exceeds the threshold.

[0127] In a sixth embodiment of the first and second aspects, the network node 103 may monitor the duration of pausing or stopping of the ML model training using a watch dog timer comprised in the network node 103. The watchdog timer may detect training pauses or stops that are taking too long and have not yet resumed. If the training pause or stop goes beyond a certain threshold, training may stale, and the output may no longer be valid for some use cases. As such, it would be better to clear all and restart the process using fresh data. This watchdog timer can also monitor the pause or stop requests coming from various training parties to detect a potentially compromised training party sending unnecessary pause or stop requests to jeopardize the training.

[0128] In a seventh embodiment of the first and second aspects, in case the training node 104 pauses or stops the training process, the network node 103 may transfers results of the ML model training obtained by one training node, due to the training node pausing or stopping the ML model training, to another training node. For example, results of the ML model training obtained by the training node may comprise partially trained ML model and / or collected training data.

[0129] In some embodiments, the network node 103 receives a request for canceling the subscription to the service associated with the ML model from the service consumer 102. The service consumer 102 may transmit the request for cancellation upon receiving 405 the result of the evaluation comprising pausing or stopping the ML model training. Service consumers 102 may cancel their subscription until they receive a notification about the end of the training procedure and / or the availability of the ML model for the AnLF to use the ML model for inference.

[0130] In some embodiments, the network node 103 comprising NWDAF with MTLF capability is deployed in a public land mobile network, PLMN, and the training nodes 104 are other network node 103 comprising NWDAF instances with MTLF capabilities deployed in the same or different PLMN. The signaling among NWDAFs is based on the concept of the Service-Based Interface (SBI), with the specified prefix “Nwdaf”.

[0131] This invention is described in a way highly relevant to 3GPP SA2 Kl#2, however, is applicable to almost every ML model training process. The invention seeks to improve the ML model training process involving many training parties. This invention proposes new signals to enable pause / stop-resume during a training procedure. The pause process can be requested by the training parties, where the intent would be to pause / stop or resume a training depending on the available resources and the criticality of the task of training the model). This invention is applicable for all training instances (centralized, HFL, VFL, and future training architectures). This invention enables much better handling of issues that may happen during the training process while allowing involved parties to better use their internal resources when the training is halted. This invention improves explainability and training logs, which can in turn be useful later for credit assignment to various participants (lower credit if a similar party adds big delays in many training instances), This invention improves the robustness against various attacks during training, as the process can be paused temporarily without a need to restart the process from scratch. This invention helps fine-tune training schedule and other parameters based on historical pause / stop-resume signals. This invention makes service consumers aware about the status of the training operations and potentials model unavailability periods. This invention has immediate application to NWDAF-based ML model training in both centralized and distributed setups.

[0132] Figure 5 illustrates an example of the network node 103 according to the embodiments of the present invention. The network node 103 illustrated in Figure 5 implements the method 200 and its embodiments as illustrated in Figures 2 and 4, for example on receipt of suitable instructions from a computer program 501 . The network node 103 comprises a processor or processing circuitry 502, and a computer program product 504 in the form of a memory 503. The processing circuitry 502 is operable to perform some or all of the steps of the method 200 and its embodiments as illustrated in Figures 2 and 4, respectively. The memory 503 contains instructions executable by the processing circuitry 502 such that the network node 103 is operable to perform some or all of the steps of the method 200 and its embodiments as illustrated in Figures 2 and 4, respectively. The instructions may also include instructions for executing one or more telecommunications and / or data communications protocols. The instructions may be stored in the form of the computer program 501 . In some examples, the processor or processing circuitry 502 may include one or more microprocessors or microcontrollers, as well as other digital hardware, which may include digital signal processors (DSPs), special-purpose digital logic, etc. The processor or processing circuitry 502 may be implemented by any type of integrated circuit, such as an Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA), etc. The memory 503 may include one or several types of memory suitable for the processor, such as read-only memory (ROM), random-access memory, cache memory, flash memory devices, optical storage devices, solid state disk, hard disk drive, etc.

[0133] Figure 6 illustrates an example of the training node 104 according to the embodiments of the present invention. The training node 104 illustrated in Figure 6 implements the method 300 and its embodiments as illustrated in Figures 3 and 4, for example on receipt of suitable instructions from a computer program 601. The training node 104 comprises a processor or processing circuitry 602, and a computer program product 604 in the form of a memory 503. The processing circuitry 502 is operable to perform some or all of the steps of the method 300 and its embodiments as illustrated in Figures 3 and 4, respectively. The memory 503 contains instructions executable by the processing circuitry 502 such that the training node 104 is operable to perform some or all of the steps of the method 300 and its embodiments as illustrated in Figures 3 and 4, respectively. The instructions may also include instructions for executing one or more telecommunications and / or data communications protocols. The instructions may be stored in the form of the computer program 501 . In some examples, the processor or processing circuitry 502 may include one or more microprocessors or microcontrollers, as well as other digital hardware, which may include digital signal processors (DSPs), special-purpose digital logic, etc. The processor or processing circuitry 502 may be implemented by any type of integrated circuit, such as an Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA), etc. The memory 503 may include one or several types of memory suitable for the processor, such as read-only memory (ROM), random-access memory, cache memory, flash memory devices, optical storage devices, solid state disk, hard disk drive, etc.

Claims

CLAIMS1. A method (200, 400) for synchronizing machine learning, ML, model training in a communication network (101 ), the method (200, 400) performed by a network node (103), the method (200, 400) comprising: obtaining (201 , 402a, 402b), from a training party involved in the ML model training in the communication network (101 ), a synchronization request comprising a first context indicating a need for synchronizing the ML model training; evaluating (202, 403) the first context comprised in the synchronization request; notifying (203, 404) the training party a result of the evaluation; and configuring (204, 407) the ML model training based on the result of the evaluation.

2. The method (200, 400) according to claim 1 , wherein obtaining (201 , 402a, 402b) the synchronization request when one or more conditions for synchronizing the ML model training are met.

3. The method (200, 400) according to any of the preceding claims, wherein the synchronization request comprises one of: one or more pause requests, one or more stop requests, or one or more resume requests.

4. The method (200, 400) according to any of the preceding claims, wherein the result of the evaluation comprises one of: to pause the ML model training, to stop the ML model training, to resume the ML model training, or to continue the ML model training.

5. The method (200, 400) according to any of the preceding claims, wherein the configuring (204, 407) the ML model training comprises performing one of:pausing the ML model training, stopping the ML model training, resuming the ML model training, or continuing the ML model training.

6. The method (200, 400) according to any of the preceding claims, wherein evaluating (202, 403) the first context comprises: determining whether the need indicated by the first context is an acceptable need for one or more training nodes (104) and / or network node (103) to synchronize the ML model training.

7. The method (200, 400) according to any of the preceding claims, wherein configuring (204, 407) the ML model training comprises: notifying a second context, to each training node (104) in one or more training nodes (104) performing the ML model training, wherein the second context comprises a modification to the ML model training based on the first context and the result of the evaluation.

8. The method (200, 400) according to any of claims 2-7, wherein configuring (204, 407) the ML model training comprises: modifying the ML model training based on the first context and / or the second context when the result of the evaluation is to pause or stop the ML model training.

9. The method (200, 400) according to claim 8, wherein the modifying comprises one of or a combination of: removing or adding one or more training nodes (104), removing or adding one or more features of the ML model training, changing one or more parameters associated with the ML model training, initiating a recovery process of one or more training nodes (104), or changing a training node selection policy.

10. The method (200, 400) according to any of claims 8-9, wherein the modifying is performed based on a recommendation from a recommendation system.11 . The method (200, 400) according to claim 10, wherein input to the recommendation system comprises one of or a combination of: the synchronization request, the first context comprised in the synchronization request, a vector or matrix comprising similarity scores indicating similarity between two or more training nodes (104).

12. The method (200, 400) according to any of the preceding claims, comprising: receiving (401 ) a subscription request from a service consumer node (102), wherein the subscription request indicates that the service consumer node (102) is subscribing to a service associated with the ML model, offered by the network node (103).

13. The method (200, 400) according to claim 12, comprising: notifying (405) the result of the evaluation to the service consumer node (102) when the result of the evaluation comprises: to pause the ML model training, to stop the ML model training, or to resume the ML model training.

14. The method (200, 400) according to any of claims 12-13, comprising: receiving (408) a cancellation request from the service consumer node (102) for canceling the subscription to the service when the notified result of evaluation comprises to pause the ML model training or to stop the ML model training.

15. The method (200, 400) according to any of the preceding claims 2-14, wherein the pause request indicates a duration for pausing the ML model training.

16. The method (200, 400) according to any of the preceding claims 2-15, wherein the stop request indicates a duration for stopping the ML model training.

17. The method (200, 400) according to any of the preceding claims 2-16, wherein the pause or stop request is periodical or time bounded.

18. The method (200, 400) according to any of the preceding claims, wherein the synchronization request comprises the one or more conditions for synchronizing the ML model training.

19. The method (200, 400) according to any of the preceding claims 15-18, comprising: monitoring the duration of pausing or stopping of the ML model training using a watch dog timer comprised in the network node (103).

20. The method (200, 400) according to any of the preceding claims 2-19, wherein the network node (103) transfers results of the ML model training obtained by one training node, due to the training node pausing or stopping the ML model training, to another training node.

21. The method (200, 400) according to any of claims 7-20, wherein the second context is provided by the recommendation system.

22. A method (300, 400) for synchronizing machine learning, ML, model training in a communication network (101 ), the method (300, 400) performed by a training node (104) in the communication network (101 ), the method (300, 400) comprising: transmitting (301 , 402b), to a network node (103) in the communication network (101 ), a synchronization request comprising a first context indicating a need for synchronizing the ML model training; receiving (302, 404), from the network node (103), a notification indicating a result of an evaluation of the ML model training based on the first context indicating the need for synchronizing the ML model training; and configuring (303, 407) the ML model training based on the result of the evaluation.

23. The method (300, 400) according to any of the preceding claims, wherein the synchronization request comprises one of: one or more pause requests, one or more stop requests, orone or more resume requests.

24. The method (300, 400) according to any of the preceding claims, wherein the result of the evaluation comprises one of: to pause the ML model training, to stop the ML model training, to resume the ML model training, or to continue the ML model training.

25. The method (300, 400) according to any of the preceding claims, wherein the configuring (303, 407) the ML model training comprises performing one of: pausing the ML model training, stopping the ML model training, resuming the ML model training, or continuing the ML model training.

26. The method (300, 400) according to any of the preceding claims, comprising: receiving a second context from the network node (103) wherein the second context comprises a modification to the ML model training based on the first context and the result of the evaluation.

27. The method (300, 400) according to any of the preceding claims 23-26, wherein configuring (303, 407) the ML model training comprises: modifying the ML model training based on the first context and / or the second context when the result of the evaluation is to pause or stop the ML model training.

28. The method (300, 400) according to any of the preceding claims 23-27, wherein the pause request indicates a duration for pausing the ML model training.

29. The method (300, 400) according to any of the preceding claims 23-28, wherein the stop request indicates a duration for stopping the ML model training.

30. The method (300, 400) according to any of the preceding claims 23-29, wherein the pause or stop request is periodical or time bounded.

31. The method (300, 400) according to any of the preceding claims 1 -30, wherein the training party comprises the network node (103).

32. The method (300, 400) according to claim 31 , the first context comprises one or more events, detected by the network node (103), affecting the ML model training.

33. The method (300, 400) according to any of the preceding claims 1 -30, wherein the training party comprises the training node (104) in the communication network (101 ).

34. The method (300, 400) according to claim 33, wherein the first context comprises one or more events, detected by the training node (104), affecting the ML model training.

35. The method (300, 400) according to any of the claims 1-34, wherein the network node (103) comprises Network Data Analytics Function, NWDAF, or Service Management and Orchestration, SMO, function in open radio access network, ORAN.

36. The method (300, 400) according to any of the claims 1 -35, wherein the training node (104) is a core network node (103) in the communication network (101 ).

37. The method (300, 400) according to claim 36, wherein the training node (104) comprises:Access and Mobility Function, AMF,Session Management Function, SMF,Analytics Data Repository Function, ADRF, orModel Training Logical Function, MTLF38. The method (300, 400) according to any of the claims 1 -35, wherein the training node (104) is a radio access network, RAN, node in the communication network (101 ).

39. The method (300, 400) according to any of the claims 1 -35, wherein the training node (104) is a user equipment, UE, in the communication network (101 ).

40. The method (300, 400) according to any of the claims 1 -39, wherein the ML model training comprises one of: a distributed ML model training performed by a plurality of training nodes (104), a federated ML model training performed by the plurality of training nodes (104),a vertical federated ML model training performed by the plurality of training nodes (104) a horizontal federated ML model training performed by the plurality of training nodes (104). a centralized ML model training.

41. A network node (103) configured to synchronize machine learning, ML, model training in a communication network (101 ), the network node (103) comprising processing circuitry (502) and a memory (503), the memory (503) containing instructions executable by the processing circuitry (502) such that the network node (103) is operable to: obtain (201 , 402a, 402b), from a training party involved in the ML model training in the communication network (101 ), a synchronization request comprising a first context indicating a need for synchronizing the ML model training; evaluate (202, 403) the first context comprised in the synchronization request; notify (203, 404) the training party a result of the evaluation; and configure (204, 407) the ML model training based on the result of the evaluation.

42. The network node (103) of claim 41 , further operable to perform a method according to any of claims 2-21 and claims 31-40.

43. A training node (104) configured to synchronize machine learning, ML, model training in a communication network (101 ), the training node (104) comprising processing circuitry (602) and a memory (603), the memory (603) containing instructions executable by the processing circuitry (602) such that the network node (103) is operable to: transmit (301 , 402b), to a network node (103) in the communication network (101 ), a synchronization request comprising a first context indicating a need for synchronizing the ML model training; receive (302, 404), from the network node (103), a notification indicating a result of an evaluation of the ML model training based on the first context indicating the need for synchronizing the ML model training; andconfigure (303, 407) the ML model training based on the result of the evaluation.

44. The training node (104) of claim 43, further operable to perform a method according to any of claims 23-30 and claims 31 -40.

45. A computer program (501 ), comprising instructions which when run on a processor (502) of a network node (103) causes the network node (103) to perform a method according to any of claims 2-21 and claims 31-40.

46. A computer program product (504) which comprises a computer readable storage medium (503) on which a computer program according to claim 45 is stored.

47. A computer program (601 ), comprising instructions which when run on a processor (602) of a training node (104) causes the training node (104) to perform a method according to any of claims 23-30 or claims 31 -40.

48. A computer program product (604) which comprises a computer readable storage medium (603) on which a computer program according to claim 47 is stored.