Method and apparatus for joint learning in network
By sending request messages in the network and receiving response messages to select appropriate computing nodes to participate in the joint learning process, the problem of inefficient computing node selection in the existing technology is solved, and more efficient resource utilization and joint learning process are achieved.
Patent Information
- Application Number
- CN202480014867.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-05
- Filing Date
- 2024-01-04
- Publication Date
- 2025-10-03
AI Technical Summary
When selecting appropriate computing nodes in the network to participate in the federated learning process, existing technologies have difficulty in effectively identifying and selecting appropriate computing nodes, resulting in inefficiency and waste of resources.
The first computing node sends a request message to multiple client computing nodes, receives a response message, selects a suitable computing node to join the federated learning process based on the response message, and uses the same service extension or reuse to notify the selected computing node to start FL training. The request message contains information such as available data requirements and availability time requirements.
It improves the efficiency of computing node selection and resource utilization, reduces unnecessary computing node participation, and improves the efficiency of the joint learning process.
Smart Images

Figure CN120752648A_ABST
Abstract
Description
[0001] Related applications This patent application claims priority to U.S. Provisional Patent Application No. 63 / 437,283, filed on January 5, 2023, the disclosure of which is incorporated by reference in its entirety. Technical Field
[0002] The present disclosure relates to computing systems, and more particularly to FL (Federated Learning) in networks. Background Art
[0003] Federated learning (FL) is an ML (machine learning) technique that distributes data across multiple compute nodes (e.g., edge devices, servers, etc.) in a network, typically without exchanging those local data samples. Federated learning, also known as collaborative learning, contrasts with traditional centralized ML techniques, where all local datasets reside on a single server, and other decentralized approaches, where local data samples are identically distributed across all compute nodes.
[0004] Federated learning enables multiple vendors to build common machine learning models without sharing data, which addresses issues such as data privacy, data security, data access rights, and access to heterogeneous data. Federated learning applications span many industries, including defense, telecommunications, IoT, and pharmaceuticals.
[0005] A network may include many computing nodes, but not all of these computing nodes may be suitable candidates to participate in the federated learning process. Suitable candidates need to be identified and selected to participate in the federated learning process. Summary of the Invention
[0006] According to one aspect, a method for performing a federated learning process in a network is provided. The method involves a first computing node sending a request message to a plurality of client computing nodes to participate in the federated learning process. The method also involves the first computing node receiving at least one response message in response to the request message. The method also involves the first computing node selecting, based on the at least one response message, which computing nodes from the plurality of client computing nodes to join the federated learning process. The method also involves the first computing node notifying the selected computing nodes to perform the federated learning process.
[0007] The request message has a message type or flag indicating the ML (machine learning) preparation phase. According to embodiments of the present disclosure, the request message also includes information indicating available data requirements and / or availability time requirements. This information may be useful for the client computing node to determine whether to join the federated learning process. This information may also be useful for the client computing node to execute the federated learning process. In some implementations, the information in the request message indicates both the available data requirements and the availability time requirements.
[0008] In some implementations, the request message also includes interoperability information, such that information indicating availability data requirements and / or availability time requirements is additional information that supplements the interoperability information. The combination of information provided to the client computing node may be useful for the client computing node to determine whether to join the federated learning process. This combination of information may also be useful for the client computing node to perform the federated learning process.
[0009] In some implementations, the same service is used to send the request message and notify the selected compute nodes. This service can be expanded or reused to notify the selected compute nodes to begin FL training. Expanding or reusing services can provide various efficiency advantages.
[0010] According to another aspect, a method for performing a federated learning process in a network is provided. The method involves a first computing node sending a request message to a plurality of client computing nodes to participate in the federated learning process. The method also involves the first computing node receiving at least one response message in response to the request message. The method also involves the first computing node selecting, based on the at least one response message, which computing nodes from the plurality of client computing nodes to join the federated learning process. The method also involves the first computing node notifying the selected computing nodes to perform the federated learning process.
[0011] The request message has a message type or flag indicating the ML (machine learning) preparation phase. According to embodiments of the present disclosure, the same service is used to send the request message and notify the selected compute nodes. This allows the service used to send the request message to be expanded or reused to notify the selected compute nodes to begin FL training. Expanding or reusing services can provide various efficiency advantages.
[0012] According to another aspect, there is provided a non-transitory computer-readable medium having recorded thereon statements and instructions that, when executed by a processor of a first computing node, configure the processor to implement the method as described above.
[0013] According to another aspect, a first computing node configured to perform a federated learning process in a network is provided. The first computing node has a network interface configured to communicate with other computing nodes of the network, and a federated learning circuit coupled to the network interface. The federated learning circuit is configured to send a request message for participating in the federated learning process to a plurality of client computing nodes via the network interface. The federated learning circuit is further configured to receive at least one response message in response to the request message via the network interface. The federated learning circuit is further configured to select which computing nodes of the plurality of client computing nodes to join the federated learning process based on the at least one response message. The federated learning circuit is further configured to notify the selected computing nodes via the network interface to perform the federated learning process.
[0014] The request message has a message type or flag indicating the ML (machine learning) preparation phase. According to embodiments of the present disclosure, the request message also includes information indicating available data requirements and / or availability time requirements. This information may be useful for the client computing node to determine whether to join the federated learning process. This information may also be useful for the client computing node to execute the federated learning process. In some implementations, the information in the request message indicates both the available data requirements and the availability time requirements.
[0015] According to another aspect, a first computing node configured to perform a federated learning process in a network is provided. The first computing node has a network interface configured to communicate with other computing nodes of the network, and a federated learning circuit coupled to the network interface. The federated learning circuit is configured to send a request message for participating in the federated learning process to a plurality of client computing nodes via the network interface. The federated learning circuit is further configured to receive at least one response message in response to the request message via the network interface. The federated learning circuit is further configured to select which computing nodes of the plurality of client computing nodes to join the federated learning process based on the at least one response message. The federated learning circuit is further configured to notify the selected computing nodes via the network interface to perform the federated learning process.
[0016] The request message has a message type or flag indicating the ML (machine learning) preparation phase. According to embodiments of the present disclosure, the same service is used to send the request message and notify the selected compute nodes. This allows the service used to send the request message to be expanded or reused to notify the selected compute nodes to begin FL training. Expanding or reusing services can provide various efficiency advantages.
[0017] According to another aspect, a method for performing a federated learning process in a network is provided. The method involves a client computing node receiving a request message to participate in the federated learning process. The method also involves the client computing node determining, based on information provided in the request message and the availability and capabilities of the client computing node, whether to join the federated learning process. The method also involves the client computing node sending a response message indicating whether to join the federated learning process.
[0018] The request message has a message type or flag indicating the ML (machine learning) preparation phase. According to embodiments of the present disclosure, the request message also includes information indicating available data requirements and / or availability time requirements. This information may be useful for the client computing node to determine whether to join the federated learning process. This information may also be useful for the client computing node to execute the federated learning process. In some implementations, the information in the request message indicates both the available data requirements and the availability time requirements.
[0019] In some implementations, the request message also includes interoperability information, such that information indicating availability data requirements and / or availability time requirements is additional information that supplements the interoperability information. The combination of information provided to the client computing node may be useful for the client computing node to determine whether to join the federated learning process. This combination of information may also be useful for the client computing node to perform the federated learning process.
[0020] According to another aspect, there is provided a non-transitory computer-readable medium having recorded thereon statements and instructions that, when executed by a processor of a client computing node, configure the processor to implement the method as described above.
[0021] According to another aspect, a client computing node configured to perform a federated learning process in a network is provided. The client computing node has a network interface configured to communicate with other computing nodes in the network, and federated learning circuitry coupled to the network interface. The federated learning circuitry is configured to receive, via the network interface, a request message for participating in the federated learning process, wherein the request message has a message type or flag indicating an ML (machine learning) preparation phase. The federated learning circuitry is further configured to determine, based on information provided in the request message and the availability and capabilities of the client computing node, whether to join the federated learning process. The federated learning circuitry is further configured to send, via the network interface, a response message indicating whether to join the federated learning process.
[0022] The request message has a message type or flag indicating the ML (machine learning) preparation phase. According to embodiments of the present disclosure, the request message also includes information indicating available data requirements and / or availability time requirements. This information may be useful for the client computing node to determine whether to join the federated learning process. This information may also be useful for the client computing node to execute the federated learning process. In some implementations, the information in the request message indicates both the available data requirements and the availability time requirements.
[0023] Other aspects and features of the present disclosure will become apparent to those of ordinary skill in the art upon reviewing the following description of various embodiments of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Embodiments will now be described with reference to the accompanying drawings, in which: Figure 1 is a block diagram of an example network with computing nodes for federated learning; Figure 2 is a sequence diagram of a method for performing a federated learning process in a network; Figures 3A to 3D is a block diagram of the system, showing signaling for a prepare request, a prepare message in the ML preparation phase, a prepare message in the ML execution phase, and an extended Nnwdaf_MLPreparation service, respectively; Figure 4 It is a sequence diagram of the joint learning and training process between multiple NWDAFs (network data analysis functions); Figure 5 is a sequence diagram of the process for selection of (one or more) client NWDAFs in the federated learning preparation phase; Figure 6 is a sequence diagram of the process for NWDAF monitoring and reselection in the execution phase of joint learning; Figure 7 is a sequence diagram of a process of dynamically discovering a new NWDAF in a federated learning execution phase when information about a server NWDAF is known at (one or more) new client NWDAFs; Figure 8 is a sequence diagram of a process of dynamically discovering a new NWDAF in a federated learning execution phase when information about the server NWDAF is unknown at (one or more) new client NWDAFs; Figure 9 It is a sequence diagram for the process of extending services in the FL preparation phase and the FL execution phase; Figure 10A and Figure 10B is a sequence diagram of another method for performing the federated learning process in a network; Figure 11 is a sequence diagram of methods for subscribing to and unsubscribing from ML models used for analysis; Figure 12 is a schematic diagram of an example cellular communication system in which some embodiments of the present disclosure may be implemented; and Figure 13 and Figure 14 is a block diagram of a wireless communication system represented as a 5G network architecture in which some embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0025] It should be understood at the outset that although illustrative implementations of one or more embodiments of the present disclosure are provided below, the disclosed systems and / or methods may be implemented using any number of techniques. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques shown below (including the exemplary designs and implementations shown and described herein), but may be modified within the scope of the appended claims, along with their full scope of equivalents.
[0026] introduce First reference Figure 1 , a block diagram of an example network 130 with compute nodes 140 and 150a-n for federated learning is shown. Network 130 may also have other components, but these components are not shown for simplicity. In some implementations, network 130 may be a core network of a cellular communication system, such as a 5G core network. However, other implementations are possible and within the scope of the present disclosure. Computing nodes 140 and 150a-n include a first computing node 140 and client computing nodes 150a-n. The number of computing nodes 140 and 150a-n may vary and is implementation-specific.
[0027] The first computing node 140 has a network interface 142 configured to communicate with other computing nodes (e.g., client computing nodes 150a-n) on the network 130. The first computing node 140 also has a federated learning circuit 144 coupled to the network interface 142. In some implementations, the federated learning circuit 144 includes a processor 146 that executes software, which may be sourced from a CRM (computer-readable medium) 148. However, other implementations are possible and within the scope of the present disclosure. The first computing node 140 may have additional components, but these components are not shown for simplicity.
[0028] The client computing nodes 150a-n are shown as having a configuration corresponding to that of the first computing node 140, and thus similarly have a network interface 152 and a federated learning circuit 154. However, it is to be understood that the client computing nodes 150a-n may have different and varying configurations. Note that the computing nodes 140 and 150a-n may originate from different vendors and, therefore, may differ in configuration.
[0029] The federated learning circuitry 144 of the first computing node 140 and the federated learning circuitry 154 of the client computing nodes 150a-n operate to implement a method for performing a federated learning process in the network 130. Figure 2 Describe this operation, Figure 2 is a flow chart of a method for performing a joint learning process in a network. Although the following reference Figure 1 The computing nodes 140 and 150a-n in the network 130 shown in FIG. Figure 2 But to understand, Figure 2 The method is applicable to other communication systems. Figure 2 The method is applicable to computing nodes in any appropriately configured network.
[0030] At step 2-1, first computing node 140 sends a request message to participate in the federated learning process. The request message is sent to client computing nodes 150a-n. This can be implemented in a multicast manner as depicted, or in other manners such as, for example, a daisy chain. At step 2-2, each client computing node 150a-n receives the request message. Not all client computing nodes 150a-n may be suitable candidates for participating in the federated learning process. Therefore, at step 2-3, each client computing node 150a-n determines whether to join the federated learning process based on the availability and capabilities of the client computing nodes, according to the information provided in the request message.
[0031] At step 2-4, each client computing node 150a-n sends a response message. In some implementations, the response message indicates whether to join the federated learning process. More generally, the response message may include any useful information that first computing node 140 can use to determine which client computing nodes 150a-n are suitable candidates for participating in the federated learning process. The response messages may be implemented using separate response messages as depicted, or in other ways, such as, for example, a single, consolidated response message. At step 2-5, first computing node 140 receives the response message(s).
[0032] The request message sent in step 2-1 and received in step 2-2 has a message type or flag indicating the ML preparation phase. For example, the request message may have a message type of "prepare" or a federated learning execution flag of "false." As another example, the request message may have an ML preparation flag that identifies whether the request is for preparing for federated learning or executing federated learning, and such a flag may be set to indicate preparation for federated learning (i.e., the ML preparation phase). According to an embodiment of the present disclosure, the request message also includes information indicating available data requirements and / or availability time requirements. This information may be useful for client computing nodes 150a-n to determine whether to join the federated learning process. This information may also be useful for client computing nodes 150a-n to execute the federated learning process.
[0033] "Available data requirements" refer to requirements for available local data at client compute nodes for training local ML models in the FL process. In some implementations, the available data requirements include a list of event IDs for the local data used for training, and may also include dataset statistical properties, a time window for data samples, and a minimum number of data samples. Client compute nodes can use the available data requirements to determine whether their available data meets the requirements of the FL process. "Availability time requirements" refer to requirements for the availability of client compute nodes participating in the FL process. Client compute nodes can use the availability time requirements to determine their availability for the FL process.
[0034] In some implementations, information from the request message indicates both an available data requirement and an availability time requirement. This combination of information may be particularly useful for client computing nodes 150a-n in determining whether to join the federated learning process. In particular, this combination of information may be more useful than either the available data requirement or the availability time requirement alone, because when considering the combination of the available data requirement and the availability time requirement, it may be possible to select the most suitable candidate to participate in the federated learning process.
[0035] In some implementations, the request message also includes interoperability information, such that information indicating available data requirements and / or availability time requirements is additional information that supplements the interoperability information. The interoperability information provided by the request message can be used to compare with interoperability information stored locally at the client computing node. The interoperability information primarily relates to FL operations, for example, the initial model provided by the first computing node can be trained at the client computing node, or the required operations for training the ML model can be performed at both the server and the client computing node. This is different from information indicating available data requirements and / or availability time requirements, which relate to requirements on the data and time required for the client node to train the ML model.
[0036] The combination of information provided to client computing nodes 150a-n (i.e., the interoperability information and the additional information supplementing the interoperability information) may be useful for client computing nodes 150a-n in determining whether to join the federated learning process. This combination of information may also be useful for client computing nodes 150a-n in performing the federated learning process. Other FL-related parameters provided by the request message may also be used to determine whether the client computing node's capabilities (e.g., computing and communication) meet the requirements for performing the FL process.
[0037] In some specific implementations, the request message includes some or more of the following items: analysis ID, ML model interoperability information, ML model ID identifying the provided ML model, ML model information, ML model file, ML training information such as data availability requirements and time availability requirements, training report information, ML preparation flag, ML model accuracy check flag, ML correlation ID, termination request (when terminating the federated learning identified by the ML correlation ID, and optionally indicating the reason, such as the FL client NWDAF is not selected by the FL server NWDAF for the FL process, or the FL process is paused, etc.), training filter information, the target of the training report, and the use case context.
[0038] In some implementations, the response message(s) received in steps 2-5 include a response message from each client computing node 150a-n indicating whether the client computing node 150a-n can join the federated learning process based on the interoperability information and the additional information. Thus, each client computing node 150a-n sends a response regardless of whether it can join the federated learning process. In other implementations, the response message(s) received in steps 2-5 include a response message only from each client computing node 150a-n that can join the federated learning process based on the interoperability information and the additional information. First computing node 140 may, for example, interpret the absence of a response from a client computing node as meaning that the client computing node cannot join the federated learning process. Other implementations are also possible.
[0039] In some specific implementations, the response message includes some or more of the following items: an indication of the operation execution result when the request is accepted, or an error response with a reason code when the request is not accepted (for example, NWDAF does not meet ML training requirements, ML training is not completed, NWDAF is overloaded, no longer available for the FL process, etc.), ML model ID, analysis ID, ML model information, ML correlation ID for federated learning, corresponding use case context, global ML model accuracy (model accuracy of the global ML model, which is calculated by the FL client NWDAF using local training data as a test dataset), FL training status report: local ML model metrics and training input data information generated by the FL client NWDAF during the FL process (for example, the area covered by the dataset, sampling rate, maximum / minimum value of each dimension of the data, etc.), delay event notification, and global ML model metrics.
[0040] At step 2-6, first computing node 140 selects which client computing nodes 150a-n to join the federated learning process based on the response message(s). In some implementations, first computing node 140 selects all client computing nodes 150a-n that have indicated that they can join the federated learning process. In other implementations, first computing node 140 selects a subset of client computing nodes 150a-n that have indicated that they can join the federated learning process. For example, first computing node 140 may select a subset of client computing nodes 150a-n if more client computing nodes 150a-n than needed or desired have indicated that they can join the federated learning process. First computing node 140 may also select a subset of client computing nodes 150a-n for other reasons, which may depend on, for example, the response message(s).
[0041] In the illustrated example, it is assumed that the first two client computing nodes 150a-b have been selected by first computing node 140 to participate in the federated learning process. In step 2-7, first computing node 140 notifies the selected computing nodes 150a-b to perform the federated learning process. For example, an instruction message may be sent to the selected computing nodes 150a-b. This can be implemented in a multicast manner as depicted or in other manners such as, for example, a daisy chain. In some implementations, as depicted, the instruction message is only sent to the selected computing nodes 150a-b. In other implementations, the instruction message is sent to all client computing nodes 150a-n, regardless of whether they have been selected, so that the instruction message indicates which of the client computing nodes 150a-n have been selected (e.g., the identities of the first two client computing nodes 150a-b), so that each client computing node 150a-n can determine whether they have been selected based on such an indication. Other implementations are also possible.
[0042] Finally, at step 2-8, the federated learning process is performed by the first computing node 140 and the selected computing nodes 150a-b. Note that the other client computing nodes 150n not selected by the first computing node 140 do not participate in the federated learning process. In some implementations, the selected computing nodes 150a-b utilize additional information from the request message (i.e., available data requirements and / or availability time requirements) during the federated learning process. For example, based on the available data requirements, the client computing node uses the requested local data to train the local ML model of the FL process.
[0043] In some implementations, first computing node 140 is a server NWDAF (Network Data Analysis Function), and each client computing node 150a-b is a client NWDAF. Other implementations are also possible. The details of computing nodes 140 and 150a-b may depend on network 130 and may be application-specific.
[0044] In some implementations, first compute node 140 uses the same service used to send the request message in step 2-1 to notify selected compute nodes 150a-b in step 2-7. In this way, the service used to send the request message in step 2-1 can be expanded or reused to notify selected compute nodes 150a-b in step 2-7, thereby starting FL training in step 2-8. Expanding or reusing services can provide various advantages in terms of efficiency and avoiding the need for different services to send the preparation message and notify the client to start FL training. In other implementations, separate services are used for steps 2-1 and 2-7.
[0045] Note that the parameters provided in the prepare message in step 2-7 are typically different from the parameters in the request message in step 2-1. For example, in step 2-7, the parameters in the prepare message may include initial ML model information, message type (= "execute") or FL execution flag (= "true") to indicate that the message is used to notify the client NWDAF to start FL training, FL correlation ID, and guidance information (e.g., the maximum response time for the FL client to provide temporary local ML model information).
[0046] There are many possibilities for the same service mentioned above. In some implementations, the same service is a Nnwdaf_MLPreparation_Request service, such that sending the request message involves invoking the Nnwdaf_MLPreparation_Request service operation with a message type or flag indicating the ML preparation phase, and notifying the selected compute nodes to perform the federated learning process involves invoking the Nnwdaf_MLPreparation_Request service operation with a message type or flag indicating the ML execution phase.
[0047] In other implementations, the same service is a Nnwdaf_MLModelTraining_Subscribe service, such that sending the request message involves calling the Nnwdaf_MLModelTraining_Subscribe service operation with a message type or flag indicating the ML preparation phase, and notifying the selected compute nodes to perform the federated learning process involves calling the Nnwdaf_MLModelTraining_Subscribe service operation with a message type or flag indicating the ML execution phase.
[0048] In other implementations, the same service is a Nnwdaf_MLModelTrainingInfo_Request service, such that sending the request message involves calling the Nnwdaf_MLModelTrainingInfo_Request service operation with a message type or flag indicating the ML preparation phase, and notifying the selected compute nodes to perform the federated learning process involves calling the Nnwdaf_MLModelTrainingInfo_Request service operation with a message type or flag indicating the ML execution phase.
[0049] Other possibilities for the same service not specifically mentioned herein are also possible. In general, any suitable service with an appropriate message type or flag can be used. There are a variety of possibilities for such message types and flags. Specific example details for invoking service operations using various flags are provided in the following sections. Many of these examples focus on a specific service, namely the Nnwdaf_MLPreparation_Request service. However, as mentioned above, other services (e.g., the Nnwdaf_MLModelTraining_Subscribe and Nnwdaf_MLModelTrainingInfo_Request services) are possible and within the scope of this disclosure.
[0050] According to another embodiment of the present disclosure, there is provided a non-transitory computer readable medium having statements and instructions recorded thereon, which, when executed by the processor 146 of the first computing node 140, implement the method described herein. The non-transitory computer readable medium may be Figure 1 The computer-readable medium 148 of the first computing node 140 shown in , or some other non-transitory computer-readable medium.
[0051] According to another embodiment of the present disclosure, there is provided a non-transitory computer readable medium having statements and instructions recorded thereon, which, when executed by the processor 156 of the client computing node 150a, implement the method described herein. The non-transitory computer readable medium may be Figure 1 The computer-readable medium 158 of the client computing node 150a shown in FIG. 1 , or some other non-transitory computer-readable medium.
[0052] Examples of non-transitory computer-readable media include SSD (Solid State Drive), hard drive, CD (Compact Disc), DVD (Digital Video Disc), BD (Blu-ray Disc), memory stick, etc. Other non-transitory computer-readable media are also possible.
[0053] The illustrative examples described herein focus on software implementations. However, other implementations are possible and within the scope of this disclosure. Note that other implementations may include additional or alternative hardware components, such as, for example, any appropriately configured FPGA (field programmable gate array), ASIC (application specific integrated circuit), and / or microcontroller. Thus, the federated learning circuitry 144 of the first computing node 140 and the federated learning circuitry 154 of the client computing nodes 150a-n may alternatively be implemented using any suitable combination of hardware, software, and / or firmware.
[0054] Further example details are provided in the following sections.It is to be understood that the following sections are very specific and are provided for exemplary purposes only, such that other implementations are possible and within the scope of the present disclosure.
[0055] Overview of the example system and signaling Figures 3A to 3D is a block diagram of the system, showing signaling for prepare request, prepare message in ML preparation phase, prepare message in ML execution phase and extended Nnwdaf_MLPreparation service respectively. These block diagrams are briefly described below.
[0056] like Figure 3A As shown in FIG, the server NWDAF sends a prepare request to the client NWDAFs 1...X for selection of (one or more) client NWDAFs. Parameters / information in the prepare request include message type (= "prepare") or FL execution flag (= "false"), interoperability information, available data requirements, availability time requirements, etc.
[0057] The (one or more) client NWDAFs that decide to join the ML process send a response to the server NWDAF, such as Figure 3BThen, the server NWDAF selects (one or more) client NWDAFs from the client NWDAFs that responded to the prepare request.
[0058] After the selection of (one or more) client NWDAFs is completed, the server NWDAF notifies the selected (one or more) client NWDAFs to perform ML training, such as Figure 3C The parameters / information in the prepare message include initial ML model information, message type (= "execution") or ML (e.g., federated learning) execution flag, ML (e.g., FL) correlation ID, guidance information (e.g., maximum response time for (one or more) client NWDAFs to provide temporary local ML model information), etc.
[0059] Figure 3D This is a block diagram of a system illustrating signaling for the extended Nnwdaf_MLPreparation service. During the ML preparation phase, the Nnwdaf_Preparation service is extended by adding new parameters to the preparation request for the selection of client NWDAF(s). The server NWDAF sends the preparation request to the client NWDAF(s) by invoking the Nnwdaf_Preparation_Request service operation. The preparation request parameters / information include the message type ("prepare") or the FL execution flag ("false"), interoperability information, available data requirements, availability time requirements, and more. The client NWDAF(s) that decide to join the ML (e.g., FL) process respond to the server NWDAF by invoking the Nnwdaf_Preparation_Request_Response service operation, indicating that they will join the ML (e.g., FL) process. The server then performs the selection from the client NWDAF(s).
[0060] At the very beginning of the ML execution phase, the server notifies the selected client NWDAF(s) to execute ML (e.g., FL) operations by calling the Nnwdaf_MLPreparation_Request service operation. The preparation message contains parameters / information such as the initial ML model information, the message type ("execution") or the ML (e.g., federated learning) execution flag, the ML (e.g., FL) correlation ID, and guidance information (e.g., the maximum response time for the client NWDAF(s) to provide temporary local ML model information).
[0061] Joint learning among multiple NWDAFs Some solutions for supporting federated learning between multiple NWDAFs in the 5GC (5G Core Network) have been summarized in clause 8.8 of 3rd Generation Partnership Project TR 23.700-81, "Study of Enablers for Network Automation for 5G," Release 18, Version 2.0.0 (2022-11) (hereinafter "TR 23.700-81"). Some descriptions and procedures for federated learning between multiple NWDAFs have been added to clause 6.2C of 3rd Generation Partnership Project TS 23.288, "Architecture enhancements for 5G System (5GS) to support network data analytics services," Release 18, Version 18.0.0 (2022-12) (hereinafter "TS 23.288").
[0062] NWDAF containing MTLF (Model Training Logic Function) can utilize federated learning technology to train ML models, which does not require input data transmission (e.g., centralized to one NWDAF), but instead involves cooperation between multiple NWDAFs (MTLFs) distributed in different areas, i.e., sharing learning results and (one or more) ML models between multiple NWDAFs (MTLFs). Figure 4 This is a sequence diagram of the joint learning and training process between multiple NWDAFs. The following briefly describes each step.
[0063] In step 4-0, the consumer (NWDAF with AnLF) sends a subscription request to the NWDAF with MTLF to retrieve the ML model. This subscription request includes the analysis ID and ML model filter information as described in TS 23.288. The NWDAF with MTLF can be a FL server (server NWDAF) with FL server capabilities or a MTLF without FL server capabilities. Note that there are many possibilities for the MTLF registration and discovery process for FL. Also note that there are many possibilities for defining FL capabilities.
[0064] In step 4-1, the server NWDAF sends a request to the selected NWDAF (client NWDAF) that includes the MTLF and participates in the federated learning to perform local model training for the federated learning.
[0065] In step 4-2, each client NWDAF collects its local data by using the current mechanism in clause 6.2 of TS 23.288.
[0066] In step 4-3, during the federated learning training process, each client NWDAF further trains the ML model retrieved from the server NWDAF based on its own data and reports temporary local ML model information to the server NWDAF. During the FL training process, ML model information is exchanged between the client NWDAF(s) and the server NWDAF. Note that there are many possibilities for the ML model information exchanged between the client NWDAF(s) and the server NWDAF. Also note that there are many possibilities for service-enabled FL-based ML model training in steps 4-1, 4-3, 4-5a, and 4-6.
[0067] In step 4-4, the server NWDAF aggregates all local ML model information retrieved in step 4-3 to update the global ML model.
[0068] In step 4-5a, based on consumer request, the server NWDAF updates the training status (accuracy level) to the consumer periodically (one or more rounds of training or every 10 minutes, etc.) or dynamically when a certain predetermined status (such as a certain accuracy level) is reached.
[0069] Optionally, in step 4-5b, the consumer determines whether the current model can meet various requirements, such as accuracy and time. If the current model can meet the requirements, the consumer modifies the subscription. Note that the various requirements in step 4-5b are specific to the consumer and are different from the available data requirements and availability time requirements described above for the client compute node. Note that there are many possibilities for the accuracy of the FL model process.
[0070] In step 4-5c, according to the request from the consumer, the server NWDAF updates or terminates the current FL training process.
[0071] In steps 4-6, if the FL process continues, the server NWDAF sends the aggregated ML model information to each client NWDAF for the next round of model training.
[0072] In step 4-7, each client NWDAF updates its own ML model based on the aggregated ML model information distributed by the server NWDAF in step 4-6.
[0073] Note that steps 4-2 to 4-7 should be repeated until the training termination condition is reached (e.g., the maximum number of iterations, or the result of the loss function is lower than a threshold). After the training process is completed, the server NWDAF can send the global optimal ML model information to the consumer.
[0074] Maintenance of Federated Learning in 5GC TR 23.700-81 provides a solution for maintaining the FL process between multiple NWDAFs in a 5GC (Solution #51). This solution is proposed to address Key Issue #8: Supporting Federated Learning in 5GC. Key research points for this key issue include: ● Study how to coordinate multiple NWDAFs, including the selection of participant NWDAF instances in a joint learning group, such as the auxiliary information used to perform the selection (if any) and the decision of the roles of the participant NWDAFs.
[0075] ● Study whether and how to perform performance (e.g., network performance and model performance) monitoring of NWDAF federated learning operations.
[0076] To address the challenges outlined above for supporting federated learning in 5GC, this solution focuses on the selection of NWDAF(s) during the federated learning preparation phase and the monitoring and maintenance of NWDAF(s) during the federated learning execution phase. During the federated learning preparation phase, many factors influence the selection of client NWDAF(s), including the capabilities of the NWDAF(s), interoperability, and availability of the client NWDAF(s) for participating in federated learning.
[0077] During the federated learning execution phase, due to the dynamic changes in the federated network, the current client NWDAF(s) may leave or join. This dynamic joining and leaving of client NWDAF(s) in the 5GC should be considered during the federated learning multiple rounds of learning / training. Furthermore, the method can be applied to the server NWDAF to monitor the status changes (e.g., changes in capabilities and availability) of client NWDAF(s).
[0078] In the federated learning preparation phase, the server and (potential) client NWDAFs are discovered via NRF (Network Storage Function), and (one or more) client NWDAFs are selected through a handshake mode method. The selection of (one or more) client NWDAFs is based on availability, capabilities, etc. Figure 5 This is a sequence diagram of the process of selecting (one or more) client NWDAFs in the federated learning preparation phase. The following briefly describes each step.
[0079] In step 5-0, the NWDAF registers the federated learning capability with the NRF. The server NWDAF discovers the client NWDAF based on, for example, the federated learning capability, the analysis ID, and the like.
[0080] In step 5-1, the server NWDAF sends a federated learning preparation request to the client NWDAF(s) by invoking the Nnwdaf_MLPreparation_Request service operation with interoperability information. The preparation request may include an indication of the role of the NWDAF(s), namely, acting as the client NWDAF(s). Note that the interoperability information indicates what capabilities the client NWDAF requires or desires (e.g., the ability to run certain models) to support the FL process. For example, whether and how the server and client NWDAFs can share models. Interoperability information can be determined between different vendors.
[0081] In step 5-2, the client NWDAF(s) decide whether to join the federated learning process based on their availability, capabilities, and interoperability information.
[0082] In step 5-3, the client NWDAF(s) send a response to the server NWDAF indicating whether they want to join the FL process.
[0083] In step 5-4, the server NWDAF can send a test task to the client NWDAF(s) that want to join the FL process. The client NWDAF(s) run the test task and send the results to the server NWDAF. Note that a test task can be a micro-computation or training task, with the requirements for completing the micro-task being the same or similar to the main task. The test task can be a small task that lets the client NWDAF collect local data and send local model weights back to the server; or it can be some kind of test to ensure that the server and client NWDAFs can communicate (if they use the same FL framework or library). There are many ways to retrieve and run a test task.
[0084] In step 5 - 5 , the server NWDAF selects client NWDAF(s). The server NWDAF may consider the results of the test task for selecting the client NWDAF(s).
[0085] During the federated learning execution phase, the server NWDAF monitors the status changes of the client NWDAF(s) and may reselect the client NWDAF(s) based on their updated status, availability, and / or capabilities for the FL task. Figure 6 This is a sequence diagram of the process for NWDAF monitoring and reselection during the joint learning execution phase. The following briefly describes each step.
[0086] In step 6-1, the server NWDAF, which monitors the status of the client NWDAF(s) during the federated learning process, receives the updated status of the client NWDAF(s). The server NWDAF may monitor and obtain the updated status of the client NWDAF(s) directly and / or via the NRF. Note that the status of the client NWDAF may include changes in NF load, NF availability, or its capabilities, such as no longer supporting FL.
[0087] In step 6-2, the server NWDAF checks the status of the client NWDAF(s) based on the received information and determines whether reselection of the client NWDAF(s) for the next round(s) of federated learning is necessary. This determination is based on the updated status of the client NWDAF(s), including availability, capabilities, etc.
[0088] In step 6-3, if reselection is necessary as determined in step 6-2, then Figure 5 In steps 5-1 to 5-5, the server NWDAF reselects (one or more) client NWDAFs. The process of discovering (one or more) new client NWDAFs in the federated learning execution phase is given in clause 2.1.2.2.3 of TS 23.288.
[0089] In step 6 - 4 , if a termination request is received from the server NWDAF, the client NWDAF(s) terminate the operation of the federated learning.
[0090] There are two possible scenarios for the server NWDAF to obtain information about the new client NWDAF(s), namely, obtaining information directly from the new client NWDAF(s) or obtaining information via the NRF. The two possible scenarios include: (1) the new client NWDAF(s) directly notifies the server NWDAF, and (2) the server NWDAF obtains information about the new client NWDAF(s) via the NRF. The two possible scenarios are described below.
[0091] Figure 7This is a sequence diagram illustrating the process of dynamically discovering new NWDAFs during the federated learning execution phase, when information about the server NWDAF is known at one or more new client NWDAFs. Client NWDAFs 1 through N are selected by the server NWDAF to participate in the current round of federated learning. Client NWDAFs N+1 through N+X (which are new) are enabled to join the next several rounds of training. The new client NWDAFs available and / or capable of joining the federated learning process become aware of information about the server NWDAF and directly notify the server NWDAF. The following briefly describes each step.
[0092] In step 7-0, the server NWDAF registers the federated learning process with the NRF using the following parameters: the federated learning (FL) correlation ID and the analysis ID. Note that the FL correlation ID is used to identify a specific FL process. For example, a server NWDAF or a client NWDAF can simultaneously participate in different FL processes. When they receive messages or data from other NWDAFs, they must know which FL process the message or data is intended for.
[0093] When the server NWDAF initiates the FL process, it registers the FL process with the NRF using the FL correlation ID and analysis ID. Later, when the client NWDAF wants to dynamically join the FL process—for example, to update its local model with global information—it queries the NRF for an ongoing FL process for the analysis ID. The NRF then provides the server NWDAF ID and FL correlation ID to the client NWDAF, which can then contact the server NWDAF to join the FL process. Using the FL correlation ID, the server NWDAF knows which FL process the client NWDAF wants to join and which model it should provide to the client.
[0094] In step 7 - 1 , if information about the server NWDAF and the corresponding FL process is known via NRF, the new client NWDAF(s) inform the server NWDAF to the client NWDAF(s) of their interoperability and availability by calling the Nnwdaf_MLPreparation_Request service operation.
[0095] In step 7-2, before starting the next round of training, the server NWDAF selects (one or more) client NWDAFs from NWDAF 1 to N+X based on the updated information of (one or more) client NWDAFs. Figure 4 The same as steps 4-1 to 4-5 in .
[0096] Figure 8This is a sequence diagram of the process of dynamically discovering new NWDAFs in the federated learning execution phase when information about the server NWDAF is unknown at (one or more) new client NWDAFs. Client NWDAFs 1 to N are selected by the server NWDAF to participate in the current round of federated learning. Client NWDAFs N+1 to N+X (which are new) have the ability to join the next few rounds of training. Similar to the above, Figure 7 As described above, in step 8-0, the server NWDAF registers with the NRF regarding the federated learning process. In step 8-1, the server NWDAF dynamically obtains information about (one or more) new client NWDAFs via the NRF by subscribing to the event of new client NWDAF registration or discovering the NRF when it wishes to reselect a client NWDAF in step 8-2.
[0097] The ML (e.g., distributed machine learning / FL) process can be divided into two phases: an ML preparation phase and an ML execution phase. Some implementations described herein provide details regarding how parameters can be sent from the server NWDAF to the client NWDAF(s) for selection during the ML preparation phase. Furthermore, during the ML execution phase, which service can be used to notify the client NWDAF(s) to perform ML training. Therefore, some implementations described herein provide details regarding how parameters can be sent from the server NWDAF to the client NWDAF(s) for selection, and / or which service can be used to notify the client NWDAF(s) to perform ML training.
[0098] Some implementations described in this document extend the Nnwdaf_MLPreparation service given in Solution #51 of TR 23.700-81. This extension may include • In the ML preparation phase, new parameters are added to the prepare request sent from the server NWDAF to the client NWDAF(s) for selection by the client NWDAF(s).
[0099] ● In the ML execution phase, reuse Nnwdaf_MLPreparation with extended parameters to notify (one or more) client NWDAFs to perform ML training.
[0100] The process of using the extended service to send prepare requests in the ML preparation phase and send prepare messages in the ML execution phase is given.
[0101] In some implementations, the Nnwdaf_MLPreparation service described in Solution #51 of TR 23.700-81 is extended by adding new parameters to the prepare request during the ML preparation phase and reused and extended during the ML execution phase. During the ML preparation phase, the prepare request is sent from the server NWDAF to the client NWDAF(s) for selection. During the ML execution phase, the server NWDAF uses the extended Nnwdaf_MLPreparation service to notify the client NWDAF(s) to perform ML training.
[0102] In some implementations, the ML preparation phase and the ML execution phase are improved by extending and reusing the Nnwdaf_MLPreparation service. The extended Nnwdaf_MLPreparation service is used to send a preparation request for selection of (one or more) client NWDAFs in the ML preparation phase and to notify (one or more) client NWDAFs to perform ML training in the ML execution phase.
[0103] Services for exchanging information and / or preparing parameters Taking the federated learning process as an example, in order to illustrate how the extended Nnwdaf_MLPreparation service is used to send preparation requests and messages from the server NWDAF to (one or more) client NWDAFs in the FL preparation phase and FL execution phase, respectively.
[0104] Figure 9 This is a sequence diagram of the process for extending services in the FL preparation phase and the FL execution phase. The following briefly describes each step.
[0105] In step 9-0, the server NWDAF discovers the client NWDAF via the NRF. The NRF provides the server NWDAF with a list of (one or more) client NWDAFs based on FL capability, analysis ID, interoperability indicator, time interval for supporting FL, and the like.
[0106] In step 9-1, the server NWDAF sends a federated learning preparation request to the client NWDAF(s) by calling the Nnwdaf_MLPreparation_Request service operation with interoperability information. The following parameters can be added to the preparation request: available data requirement, availability time requirement, message type (= "prepare") or FL execution flag (= "false") to indicate that the request is for preparation.
[0107] In step 9 - 2 , the client NWDAF(s) decide whether to join the federated learning process based on their availability, computing and communication capabilities, and interoperability information.
[0108] In step 9 - 3 , the client NWDAF(s) sends a response to the server NWDAF by calling the Nnwdaf_MLPreparation_Request response service operation with an indication about joining the FL process to indicate that it will join the FL process.
[0109] In step 9 - 4 , the server NWDAF performs selection of client NWDAF(s).
[0110] In step 9-5, the server NWDAF notifies the selected NWDAF (client NWDAF) containing the MTLF to perform federated learning by calling the Nnwdaf_MLPreparation_Request service operation with the initial ML model information and a message type ("execute") or FL execution flag ("true") indicating that this message is used to notify the client NWDAF to start FL training. The server NWDAF also includes the FL correlation ID and guidance information (e.g., the maximum response time for the FL client to provide temporary local ML model information) in the prepare message.
[0111] For the purpose of indicating Nnwdaf_MLPreparation_Request, for example, for sending a preparation request in the FL preparation phase or for notifying a preparation message in the FL execution phase, a message type or an FL execution flag may be used in steps 9-1 and 9-5.
[0112] If the FL Execute Flag is used in steps 9-1 and 9-5, it can be either an optional or a required parameter. If the FL Execute Flag is an optional parameter, it is only included in the Prepare message sent from the server NWDAF to the client NWDAF(s) during the FL Execute phase. If the FL Execute Flag is a required parameter, it will be set to FL Execute Flag = "false" in the Prepare Request during the FL Prepare phase and to FL Execute Flag = "true" in the Prepare Message during the FL Execute phase. More generally, a message type or flag can be used to indicate either the ML Prepare phase or the ML Execute phase.
[0113] At step 9 - 6 , the server NWDAF and the client NWDAF(s) begin performing FL training operations for the FL process.
[0114] Another joint learning process Figure 10A and Figure 10B This is a sequence diagram of another approach to performing the federated learning process in a network. The following briefly describes the individual steps.
[0115] In step 10-0a, the NWDAF registers with the NRF. In step 10-0b, the consumer (NWDAF containing AnLF) sends a subscription request to the NWDAF containing MTLF to retrieve the ML model, including the analysis ID and ML model filter information. As described in clause 7.5.2 of TS 23.288, the NWDAF containing MTLF can be a FL server (server NWDAF) with FL server capabilities or a MTLF without FL server capabilities. In step 10-0c, the server NWDAF discovers the client NWDAF via the NRF. The NRF provides the server NWDAF with a list of client NWDAFs based on FL capabilities, analysis ID, interoperability indicator, FL supported time intervals, and so on. Note that there are many possibilities for defining FL capabilities.
[0116] In step 10-1, the server NWDAF sends a federated learning preparation request to one or more client NWDAFs by calling the Nnwdaf_MLPreparation_Request service operation with interoperability information. In the preparation request, the following parameters may be added: available data requirement and availability time requirement.
[0117] In step 10 - 2 , the client NWDAF(s) decide whether to join the federated learning process based on their availability, computing and communication capabilities, and interoperability information.
[0118] In some implementations, the server NWDAF may use this request to check whether the NWDAF can meet the ML model training requirements (e.g., ML model interoperability information, analysis ID, service area, and / or data and time availability). In this case, the FL server NWDAF includes the ML-ready flag. In some implementations, when the ML-ready flag is present in the request, the service provider NWDAF only checks whether it can meet the ML model training requirements (e.g., ML model interoperability information, analysis ID, service area, and / or data and time availability) and / or whether it can successfully download the model if the model information is provided.
[0119] At step 10-3, the client NWDAF(s) send a response to the server NWDAF indicating whether it will join the FL process. This may be based on whether the ML model training requirements can be met.
[0120] In step 10-4, the server NWDAF selects one or more client NWDAFs based on the response received in step 10-3.
[0121] In step 10-5, the server NWDAF notifies the selected NWDAF (client NWDAF) containing the MTLF to perform federated learning by invoking the Nnwdaf_MLPreparation_Request service operation with initial ML model information and the FL execution flag. The server NWDAF includes the FL correlation ID and guidance information (e.g., the maximum response time for the FL client to provide temporary local ML model information) in the prepare message. As described in clause 6.2C.2.2 of TS 23.288, the server NWDAF and the client NWDAF(s) subscribe to each other for exchanging ML model information.
[0122] In step 10-6, each client NWDAF collects its local data by using the current mechanism in clause 6.2 of TS 23.288.
[0123] In step 10-7, during the federated learning training process, each client NWDAF further trains the ML model retrieved from the server NWDAF based on its own data and reports temporary local ML model information to the server NWDAF, as defined in clause 6.2C.2.2 of TS 23.288. During the FL training process, ML model information is exchanged between the client NWDAF(s) and the server NWDAF.
[0124] In step 10 - 8 , the server NWDAF aggregates all local ML model information retrieved in step 10 - 7 to update the global ML model.
[0125] In step 10-9a, based on the consumer request, the server NWDAF provides the consumer with the training status (i.e., accuracy level / information) by dynamically calling the Nnwdaf_MLModelProvision_Notify service operation periodically (one or more training rounds or every 10 minutes, etc.) or when a certain predetermined status (e.g., a certain accuracy level) is reached.
[0126] Optionally, at step 10-9b, the consumer determines whether the current model meets various requirements, such as accuracy and time. If the current model meets the requirements, the consumer modifies the ML model subscription. Note that the various requirements at step 10-9b are specific to the consumer and differ from the aforementioned available data requirements and availability time requirements for the client compute nodes. Note that there are many possibilities for the accuracy of the FL model provisioning process.
[0127] In step 10-9c, according to the request from the consumer, the server NWDAF updates or terminates the current FL training process.
[0128] In step 10-10, if the FL process continues, the server NWDAF sends the aggregated ML model information or ML model container to each client NWDAF for the next round of model training, as defined in clause 6.2C.2.2 of TS 23.288.
[0129] In steps 10 - 11 , each client NWDAF updates its own ML model based on the aggregated ML model information distributed by the server NWDAF in step 10 .
[0130] Note that steps 10-7 to 10-11 should be repeated until the training termination condition is reached (e.g., the maximum number of iterations, or the result of the loss function is lower than a threshold). After the training process is completed, the server NWDAF can send the global optimal ML model information to the consumer.
[0131] Model Information Exchange The Nnwdaf_MLModelProvision service is used for model sharing / parameter exchange during the execution phase of the ML training process between multiple NWDAFs.
[0132] The server NWDAF subscribes to (one or more) client NWDAFs for obtaining local model information by calling the Nnwdaf_MLModelProvision_Subscribe service operation. The client NWDAFs subscribe to the server NWDAF for obtaining global model information by calling the Nnwdaf_MLModelProvision_Subscribe service operation as described in 6.2A.1. The differences are as follows: The server NWDAF and the client NWDAF(s) include the following parameters in the Nnwdaf_MLModelProvision_Subscribe request: ● The identifier of the current ML process, that is, the FL correlation ID.
[0133] The server NWDAF and the client NWDAF(s) include the following parameters in the Nnwdaf_MLModelProvision_Notify message: ● The identifier of the current ML process, that is, the FL correlation ID.
[0134] ● The identifier of the current iteration round, i.e., IR ID (for example, IR ID = 1, 2, 3, ...).
[0135] Note that whether the server NWDAF and the client NWDAF share models or model parameters depends on the initial information in step 5 of clause 6.2C.2.1 of TS 23.288.
[0136] ML model subscription / unsubscription Figure 11 is a sequence diagram of the methods for subscribing and unsubscribing to ML models for analysis. Figure 11 The procedures in this section are used by an NWDAF service consumer (i.e., an NWDAF with AnLF or MTLF) that subscribes to / unsubscribes from another NWDAF (i.e., an NWDAF with MTLF) to be notified when ML model information becomes available for relevant analyses, or for model sharing / parameter exchange during the execution phase of a federated learning training process between multiple NWDAFs, using the Nnwdaf_MLModelProvision service, as defined in clause 7.5 of TS 23.288. ML model information is used by an NWDAF with AnLF to derive analyses or by an NWDAF with MTLF to update models (global or local). This service is also used by an NWDAF to modify existing ML model subscriptions. An NWDAF can be both a consumer of this service provided by other NWDAFs and a provider of this service to other NWDAFs. The following briefly describes the various steps.
[0137] In step 11-1, an NWDAF service consumer (i.e., an NWDAF containing AnLF or MTLF) subscribes to, modifies, or cancels its subscription to (a set of) trained ML models associated with (a set of) analysis IDs by calling the Nnwdaf_MLModelProvision_Subscribe / Nnwdaf_MLModelProvision_Unsubscribe service operations. The parameters that can be provided by an NWDAF service consumer are listed in clause 6.2A.2 of TS 23.288. If applicable, the service consumer can optionally indicate support for multiple ML models.
[0138] When receiving a subscription to a trained ML model associated with an analysis ID, the NWDAF containing the MTLF may: ● Determine whether an existing trained ML model is available for subscription; or ● Determine whether a subscription requires or is expected to trigger further training of an existing trained ML model.
[0139] As described in clause 6.2 of TS 23.288, if the NWDAF including the MTLF determines that further training is necessary, the NWDAF may initiate data collection from the NF (e.g., AMF / DCCF / ADRF), UE application (via AF), or OAM to generate the ML model.
[0140] If the service call is for subscription modification or subscription cancellation, the NWDAF service consumer includes the identifier to be modified (subscription correlation ID) in the call to Nnwdaf_MLModelProvision_Subscribe.
[0141] If the service call is for model sharing / parameter exchange in the execution phase of a federated learning training process between multiple NWDAFs, the NWDAF service consumer includes the identity of the current ML process (FL correlation ID) in the call to Nnwdaf_MLModelProvision_Subscribe.
[0142] In step 11-2, if the NWDAF service consumer subscribes to (the set of) trained ML models(s) associated with (the set of) analysis ID(s), the NWDAF including the MTLF notifies the NWDAF service consumer that: ● When the consumer does not support multiple ML models, trained ML model information (including (a set of) file addresses of (one or more) trained ML models); or ● When the consumer supports multiple ML models, a set of pairs of ML model information and unique ML model identifiers associated with the analysis ID.
[0143] Note that the structure and format of the ML model identifier and its uniqueness depends on (up to) Phase 3. Also note that the parameters defined for multiple models are to improve the accuracy of the analysis.
[0144] By calling the Nnwdaf_MLModelProvision_Notify service operation, clause 6.2A.2 of TS 23.288 specifies the content of the trained ML model information that can be provided by the NWDAF including the MTLF.
[0145] When the NWDAF including the MTLF determines in step 11 - 1 that the previously provided trained ML model should have retraining, the NWDAF including the MTLF also calls the Nnwdaf_MLModelProvision_Notify service operation to notify the available retrained ML model.
[0146] When step 11 - 1 is for subscription modification (ie, including subscription correlation ID), the NWDAF including the MTLF may provide a new trained ML model different from the previously provided trained ML model, or a retrained ML model, by calling the Nnwdaf_MLModelProvision_Notify service operation.
[0147] When step 11-1 is used for model sharing / parameter exchange in the execution phase of the federated learning training process between multiple NWDAFs, the FL correlation ID and identity (IR ID (e.g., IR ID = 1, 2, ...)) of the current iteration round should be included in the Nnwdaf_MLModelProvision_Notify message.
[0148] ML model preconfiguration content A consumer of the ML model provisioning service as described in clauses 7.5 and 7.6 of TS 23.288 (i.e., NWDAF including AnLF or MTLF) may provide the following input parameters: ● Information about the analysis that the requested ML model will be used for, including: ○ List of analysis ID(s): Identifies the analysis the ML model is used for.
[0149] [Optional] Use Case Context: Indicates the context in which the analysis is used to select the most relevant ML model. Note that when several ML models are available for the requested analysis ID(s), the NWDAF containing the MTLF can use the parameter "Use Case Context" to select the most relevant ML model. The value of this parameter is not standardized.
[0150] ○ [Optional] ML model interoperability information. This is vendor-specific information that conveys, for example, the requested model file format, model execution environment, etc. The encoding, format, and value of the ML model interoperability information are not specified, as it is vendor-specific information and is to be agreed upon between vendors, if necessary for sharing purposes.
[0151] [Optional] ML Model Filter Information: Allows selection of the ML model requested for analysis, e.g., S-NSSAI, Region of Interest. The parameter types in the ML Model Filter Information are the same as those in the Analysis Filter Information defined in the procedure.
[0152] ○ [Optional] Target of ML model report: Indicates the object(s) for which the ML model is requested, e.g. a specific UE, a group(s) of UEs, or any UE (i.e. all UEs).
[0153] ○ ML model reports information with the following parameters: ○ (For Nnwdaf_MLModelProvision_Subscribe only) ML model report information parameters that comply with the event report information parameters defined in Table 4.15.1-1 of 3rd Generation Partnership Project TS 23.502, “Procedures for the 5G System (5GS)” Release 18, Version 18.0.0 (2022-12) (“TS 23.502”).
[0154] ○ [Optional] ML Model Target Period: Indicates the time interval [start, end] for requesting the ML model for analysis. The time interval is expressed as an actual start time and an actual end time (e.g., via UTC time).
[0155] ○ The Notification Target Address (+Notification Correlation ID) as defined in clause 4.15.1 of TS 23.502, allowing notifications received from the NWDAF containing the MTLF to be associated with this subscription.
[0156] ○ (Only used for model sharing / parameter exchange in the execution phase of the federated learning training process between multiple NWDAFs) The identifier of the current ML process, i.e., the FL correlation ID.
[0157] ○ [Optional] Indication for supporting multiple ML models.
[0158] ○ [Optional] The level of precision of interest.
[0159] Note that for multi-model preconfiguration, there are many possibilities whether to utilize additional parameters.
[0160] As described in clauses 7.5 and 7.6 of TS 23.288, the NWDAF including the MTLF provides the following output information to consumers of the ML Model Provisioning service operation: ● (Only for Nnwdaf_MLModelProvision_Notify) Notify dependency information.
[0161] ML model information, which includes: ○ When multiple ML models are not supported, the ML model file addresses (e.g., URLs or FQDNs) for the analysis ID(s); or ○ In case multiple ML models are supported, a set of pairs of ML model file addresses (e.g. URLs or FQDNs) and unique ML model identifiers for the analysis ID(s).
[0162] ● Validity period: Indicates the time period during which the provided ML model information is applicable.
[0163] ○ [Optional] Spatial validity: Indicates the region to which the provided ML model information is applicable. Note that the spatial validity and validity period are determined by MTLF internal logic and are respectively a subset of the ML model target period and a subset of the AoI if provided in the ML model filter information.
[0164] ● (Only used for model sharing / parameter exchange in the execution phase of the joint learning and training process between multiple NWDAFs) The identifier of the current ML process, that is, the FL correlation ID.
[0165] ● (Only used for model sharing / parameter exchange in the execution phase of the federated learning training process between multiple NWDAFs) The identifier of the current iteration round (for example, for Nnwdaf_MLModelProvision_Notify, IR ID = 1, 2,...).
[0166] Table 1: Example NF services provided by NWDAF Table 2: Example analysis information provided by NWDAF Example details of the Nnwdaf_MLModelProvision_Subscribe service operation: ● Service operation name: Nnwdaf_MLModelProvision_Subscribe.
[0167] ● Description: Subscribe to a NWDAF ML model preconfiguration with specific parameters.
[0168] ● Example input: (set of) analysis ID(s) defined in Table 7.1-2, notification target address (+ notification correlation ID), FL correlation ID (when used for model sharing / parameter exchange in the execution phase of the federated learning training process between multiple NWDAFs).
[0169] ● Optional inputs: subscription correlation ID (in case of modifying ML model subscription), ML model filter information indicating the conditions for requesting ML models for analysis and target of ML model report indicating the object(s) for which the ML model is requested (e.g., a specific UE, a group of UEs, or any UE (i.e., all UEs)), ML model reporting information (including, for example, ML model target period), expiration time, use case context, indication of support for multiple ML models, multiple ML model filter information indicating the conditions for requesting multiple ML models.
[0170] • Example output: When subscription is accepted: Subscription correlation ID (used for management of this subscription), Expiration time (used if subscription is expirable based on operator's policy).
[0171] ● Optional output: None.
[0172] ● Example details of the Nnwdaf_MLModelProvision_Unsubscribe service operation: ● Service operation name: Nnwdaf_MLModelProvision_Unsubscribe.
[0173] ● Description: Unsubscribe from NWDAF ML model pre-configuration.
[0174] ● Exemplary input: subscription correlation ID, FL correlation ID (when used for model sharing / parameter exchange in the execution phase of the federated learning training process between multiple NWDAFs).
[0175] ● Optional input: FL correlation ID (when used for model sharing / parameter exchange in the execution phase of the federated learning training process between multiple NWDAFs).
[0176] ● Example output: Indication of the result of the operation execution.
[0177] ● Optional output: None.
[0178] Example details of the Nnwdaf_MLModelProvision_Notify service operation: ● Service operation name: Nnwdaf_MLModelProvision_Notify.
[0179] ● Description: NWDAF notifies ML model information to consumer instances that have subscribed to a specific NWDAF service.
[0180] ● Exemplary input: notification correlation information, FL correlation ID and the identifier of the current iteration round (IR ID), when used for model sharing / parameter exchange in the execution phase of the joint learning training process between multiple NWDAFs, the set of the following items ○ When multiple ML models are not supported, a tuple (address of the model file (e.g., URL or FQDN), analysis ID); or ○ Tuple (one or more tuples of the address of the model file (e.g., URL or FQDN) and a unique ML model identifier, analysis ID).
[0181] ● Optional inputs: validity period, spatial validity, FL correlation ID (when used for model sharing / parameter exchange in the execution phase of the joint learning training process between multiple NWDAFs).
[0182] ● Example output: Indication of the result of the operation execution.
[0183] ● Optional output: None.
[0184] Example details for the Nnwdaf_MLPreparation service: ● Service Description: This service enables consumers to request NWDAF with MTLF to prepare or execute ML model training.
[0185] Example details of the Nnwdaf_MLPreparation_Request service operation: ● Service operation name: Nnwdaf_MLPreparation_Request ● Description: Consumer requests NWDAF to prepare or execute ML model training.
[0186] ● Example input: ● Interoperability information, available data requirements, and availability time requirements (when no federated learning implementation is provided).
[0187] ● Initial ML model information (one of the following three types: ADRF ID with ML model identifier, or address of model file (e.g., URL or FQDN), or ML model container (when ML model container exists), FL correlation ID, and guide information (e.g., maximum response time for FL client to provide temporary local ML model information) (when federated learning execution flag is provided).
[0188] ● Optional input: FL execution flag.
[0189] ● Example output: None.
[0190] ● Optional output: indication to join federated learning (when FL execution flag is not provided).
[0191] Example Communication System Now refer to Figure 12 , shows a schematic diagram of an example cellular communication system 100 in which some embodiments of the present disclosure may be implemented. In the embodiments described herein, cellular communication system 100 is a 5G system (5GS) comprising a next-generation RAN (NG-RAN) and a 5G core (5GC). In this example, the RAN includes base stations 102-1 and 102-2 that control corresponding (macro) cells 104-1 and 104-2. Base stations 102-1 and 102-2, in the 5GS, include NR base stations (GNBs) and optional next-generation eNBs (ng-eNBs) (e.g., LTE RAN nodes connected to the 5GC). Base stations 102-1 and 102-2 are generally referred to herein collectively as base stations 102, and individually as base stations 102. Similarly, (macro) cells 104-1 and 104-2 are generally referred to herein collectively as (macro) cells 104, and individually as (macro) cells 104. The RAN may also include a plurality of low power nodes 106-1 to 106-4 that control corresponding small cells 108-1 to 108-4. The low power nodes 106-1 to 106-4 may be small base stations (such as pico or femto base stations) or remote radio heads (RRHs) or the like. It is noteworthy that, although not shown, one or more of the small cells 108-1 to 108-4 may alternatively be provided by the base station 102. The low power nodes 106-1 to 106-4 are generally referred to herein as low power nodes 106, and individually as low power nodes 106. Likewise, the small cells 108-1 to 108-4 are generally referred to herein as small cells 108, and individually as small cells 108. The cellular communication system 100 also includes a core network 130A, which is referred to as 5GC in a 5G system (5GS). Note that the core network 130A is Figure 1 An example implementation of network 130 is depicted in FIG. Base stations 102 (and optionally low power nodes 106 ) are connected to a core network 130A.
[0192] Base station 102 and low power node 106 provide services to wireless communication devices 112-1 through 112-5 in corresponding cells 104 and 108. Wireless communication devices 112-1 through 112-5 are generally referred to herein collectively as wireless communication devices 112, and individually as wireless communication devices 112. In the following description, wireless communication device 112 is generally a UE, but the present disclosure is not limited thereto.
[0193] Now refer to Figure 13, shows a block diagram of a wireless communication system represented as a 5G network architecture consisting of core network functions (NFs), where the interaction between any two NFs is represented by point-to-point reference points / interfaces. Figure 13 Can be considered as Figure 12 A specific implementation of the system 100 is provided.
[0194] From the access side, Figure 13 The 5G network architecture shown includes multiple UEs 112 connected to the RAN 102 or access network (AN) and the AMF 200. Typically, the RAN 102 includes a base station, such as an eNB or gNB. From the core network side, Figure 13 The 5GC NFs shown include NSSF 202, AUSF 204, UDM 206, AMF 200, SMF 208, PCF 210, Application Function (AF) 212, and NWDAF 220. NWDAF 220 may be used to implement server and client NWDAF in the FL process.
[0195] The reference point representation of the 5G network architecture is used to develop detailed call flows in the standardization of specifications. The N1 reference point is defined as carrying signaling between the UE 112 and the AMF 200. The reference points for connecting between the AN 102 and the AMF 200 and between the AN 102 and the UPF 214 are defined as N2 and N3, respectively. Reference point N11 exists between the AMF 200 and the SMF 208, meaning that the SMF 208 is at least partially controlled by the AMF 200. N4 is used by the SMF 208 and the UPF 214, enabling the UPF 214 to be configured using control signals generated by the SMF 208 and for the UPF 214 to report its status to the SMF 208. N9 is a reference point for connecting between different UPFs 214, while N14 is a reference point for connecting between different AMFs 200. N15 and N7 are defined because PCF 210 applies policies to AMF 200 and SMF 208 respectively. N12 is used by AMF 200 to perform authentication of UE 112. N8 and N10 are defined because subscription data of UE 112 is used by AMF 200 and SMF 208.
[0196] The 5GC network is designed to separate UP and CP. UP carries user services, while CP carries signaling in the network. Figure 13In this architecture, UPF 214 resides in the UP, while all other NFs—namely, AMF 200, SMF 208, PCF 210, AF 212, NSSF 202, AUSF 204, and UDM 206—are in the CP. Separating the UP and CP ensures that resources for each plane can be scaled independently. It also allows the UPF to be deployed in a distributed manner, separate from the CP functionality. In this architecture, for some applications requiring low latency, the UPF can be deployed very close to the UE to shorten the round-trip time (RTT) between the UE and the data network.
[0197] The core 5G network architecture consists of modular functions. For example, AMF 200 and SMF 208 are independent functions in CP. Separating AMF 200 and SMF 208 allows independent evolution and scaling. Other CP functions such as PCF 210 and AUSF 204 can be Figure 13 The modular functional design enables the 5GC network to flexibly support various services.
[0198] Each NF interacts directly with another NF. It is possible to use intermediate functions to route messages from one NF to another. In the CP, a set of interactions between two NFs is defined as a service, enabling reuse. This service enables modularization. The UP supports interactions such as forwarding operations between different UPFs.
[0199] Now refer to Figure 14 , shows the use of service-based interfaces between NFs in CP instead of Figure 13 The point-to-point reference points / interfaces used in the 5G network architecture are shown in the block diagram of the 5G network architecture. Figure 14 The NF described corresponds to Figure 13 The NF shown in . The (one or more) services provided by the NF to other authorized NFs can be opened to the authorized NFs through the service-based interface. Figure 14 In the NF, the service based interface is indicated by the letter “N” followed by the NF name, such as Namf for the service based interface of AMF 200 and Nsmf for the service based interface of SMF 208, etc. Figure 14 The NEF 300 and NRF 302 in the above discussion Figure 13 However, it should be clarified that although Figure 13 There is no clear indication, but Figure 13 All NFs depicted in the Figure 14 The NEF 300 and NRF 302 interact.
[0200] It can be described in the following ways Figure 13 and14 The AMF 200 provides UE-based authentication, authorization, mobility management, and other functions. Even UEs 112 using multiple access technologies are essentially connected to a single AMF 200, as the AMF 200 is independent of the access technology. The SMF 208 is responsible for session management and assigns Internet Protocol (IP) addresses to UEs. It also selects and controls the UPF 214 for data transfer. If a UE 112 has multiple sessions, a different SMF 208 can be assigned to each session to manage them independently, potentially providing different functionality for each session. The AF 212 provides information about packet flows to the PCF 210, which is responsible for policy control, to support QoS. Based on this information, the PCF 210 determines policies regarding mobility and session management to ensure proper operation of the AMF 200 and SMF 208. The AUSF 204 supports authentication functions for the UE or the like and stores data used for authentication of the UE or the like, while the UDM 206 stores subscription data for the UE 112. Data networks (DNs), which are not part of the 5GC network, provide Internet access or operator services and similar services.
[0201] NFs can be implemented as network elements on dedicated hardware, as software instances running on dedicated hardware, or as virtualized functions instantiated on a suitable platform (e.g., cloud infrastructure).
[0202] Any suitable steps, methods, features, functions, or benefits disclosed herein may be performed by one or more functional units or modules of one or more virtual devices. Each virtual device may include multiple of these functional units. These functional units may be implemented via processing circuitry, which may include one or more microprocessors or microcontrollers, as well as other digital hardware, such as digital signal processors (DSPs), dedicated digital logic, and the like. The processing circuitry may be configured to execute program code stored in a memory, which may include one or more types of memory, such as read-only memory (ROM), random access memory (RAM), cache memory, flash memory devices, optical storage devices, and the like. The program code stored in the memory includes program instructions for executing one or more telecommunications and / or data communication protocols, as well as instructions for implementing one or more of the techniques described herein. In some implementations, processing circuitry may be used to cause the corresponding functional units to perform corresponding functions according to one or more embodiments of the present disclosure.
[0203] Many modifications and variations of the present disclosure are possible in light of the above teachings.It is therefore to be understood that within the scope of the appended claims, the present disclosure may be practiced otherwise than as specifically described herein.
Claims
1. A method for performing a federated learning process in a network, comprising: Sending, by the first computing node, a request message to the plurality of client computing nodes for participating in the federated learning process, wherein the request message has a message type or flag indicating an ML (machine learning) preparation phase, and wherein the request message includes information indicating an available data requirement and / or an availability time requirement; receiving, by the first computing node, at least one response message in response to the request message; selecting, by the first computing node, which computing nodes of the plurality of client computing nodes to join the federated learning process based on the at least one response message; and The first computing node notifies the selected computing node to execute the joint learning process.
2. The method according to claim 1, wherein The information from the request message indicates both the availability data requirement and the availability time requirement.
3. The method according to claim 1 or 2, wherein: The request message further includes interoperability information, and wherein the information indicating the availability data requirement and / or the availability time requirement is additional information supplementing the interoperability information.
4. The method according to claim 3, wherein: The at least one response message includes a response message from each client computing node, the response message indicating whether the client computing node can join the federated learning process based on the interoperability information and the additional information.
5. The method according to claim 3, wherein The at least one response message includes response messages only from each client computing node that is able to join the federated learning process based on the interoperability information and the additional information.
6. The method according to claim 4 or 5, wherein: Selecting which computing nodes to join the federated learning process includes: The first computing node selects all client computing nodes that can join the joint learning process.
7. The method according to claim 4 or 5, wherein: Selecting which computing nodes to join the federated learning process includes: A subset of the client computing nodes that can join the federated learning process is selected by the first computing node.
8. The method according to any one of claims 1 to 7, wherein The first computing node is a server NWDAF (Network Data Analysis Function), and each client computing node is a client NWDAF.
9. The method according to any one of claims 1 to 8, wherein The request message is sent using the same service and notified to the selected computing node.
10. The method according to claim 9, wherein: The same service is the Nnwdaf_MLModelTraining_Subscribe or Nnwdaf_MLModelTrainingInfo_Request service; Sending the request message comprises invoking a service operation via the same service with a message type or flag indicating an ML preparation phase; as well as Notifying the selected computing nodes to execute the federated learning process includes invoking a service operation via the same service with a message type or flag indicating an ML execution phase.
11. The method according to claim 10, wherein: For the ML execution phase, the service operation is called with at least some of ML model information, FL correlation ID, and guideline information.
12. A method for performing a federated learning process in a network, comprising: Sending, by the first computing node, a request message for participating in the federated learning process to the plurality of client computing nodes, wherein the request message has a message type or a flag indicating an ML (machine learning) preparation phase; receiving, by the first computing node, at least one response message in response to the request message; selecting, by the first computing node, which computing nodes of the plurality of client computing nodes to join the federated learning process based on the at least one response message; and The first computing node notifies the selected computing node to execute the joint learning process; The same service is used to send the request message and notify the selected computing node.
13. The method according to claim 12, wherein: The same service is the Nnwdaf_MLModelTraining_Subscribe or Nnwdaf_MLModelTrainingInfo_Request service; Sending the request message comprises invoking a service operation via the same service with a message type or flag indicating an ML preparation phase; as well as Notifying the selected computing nodes to execute the federated learning process includes invoking a service operation via the same service with a message type or flag indicating an ML execution phase.
14. The method according to claim 13, wherein: For the ML execution phase, the service operation is called via the same service with at least some of ML model information, FL correlation ID, and guideline information.
15. A non-transitory computer-readable medium having recorded thereon statements and instructions that, when executed by a processor of a first computing node, configure the processor to implement the method of any one of claims 1 to 14.
16. A first computing node configured to perform a federated learning process in a network, the first computing node comprising: a network interface configured to communicate with other computing nodes of the network; as well as a federated learning circuit coupled to the network interface and configured to: sending, via the network interface, a request message to a plurality of client computing nodes for participating in the federated learning process, wherein the request message has a message type or flag indicating an ML (machine learning) preparation phase, and wherein the request message includes information indicating an available data requirement and / or an availability time requirement; receiving, via the network interface, at least one response message in response to the request message; selecting, based on the at least one response message, which of the plurality of client computing nodes are to join the federated learning process; and The selected computing node is notified via the network interface to execute the federated learning process.
17. The first computing node according to claim 16, wherein: The joint learning circuit is further configured to implement the method of any one of claims 2 to 11.
18. A first computing node configured to perform a federated learning process in a network, comprising: a network interface configured to communicate with other computing nodes of the network; as well as a federated learning circuit coupled to the network interface and configured to: sending, via the network interface, a request message to a plurality of client computing nodes for participating in the federated learning process, wherein the request message has a message type or a flag indicating an ML (machine learning) preparation phase; receiving, via the network interface, at least one response message in response to the request message; selecting, based on the at least one response message, which of the plurality of client computing nodes to join the federated learning process; and Notifying the selected computing node via the network interface to execute the joint learning process; The same service is used to send the request message and notify the selected computing node.
19. The first computing node according to claim 18, wherein: The joint learning circuit is further configured to implement the method of any one of claims 13 to 14.
20. A method for performing a federated learning process in a network, comprising: receiving, by a client computing node, a request message for participating in the federated learning process, wherein the request message has a message type or flag indicating an ML (machine learning) preparation phase, and wherein the request message includes information indicating an available data requirement and / or an availability time requirement; According to the information provided by the request message, the client computing node determines whether to join the federated learning process based on the availability and capabilities of the client computing node; and The client computing node sends a response message indicating whether to join the joint learning process.
21. The method according to claim 20, wherein The information from the request message indicates both the availability data requirement and the availability time requirement.
22. The method according to claim 20 or 21, wherein The request message also includes interoperability information, and wherein the information indicating the availability data requirement and / or the availability time requirement is additional information supplementing the interoperability information, and wherein determining whether to join the joint learning process is based on the interoperability information and the additional information.
23. A non-transitory computer-readable medium having recorded thereon statements and instructions that, when executed by a processor of a client computing node, configure the processor to implement the method of any one of claims 20 to 21.
24. A client computing node configured to perform a federated learning process in a network, comprising: a network interface configured to communicate with other computing nodes of the network; as well as a federated learning circuit coupled to the network interface and configured to: receiving, via the network interface, a request message for participating in the federated learning process, wherein the request message has a message type or a flag indicating an ML (machine learning) preparation phase, and wherein the request message includes information indicating an available data requirement and / or an availability time requirement; determining, based on the information provided by the request message, whether to join the federated learning process based on the availability and capabilities of the client computing node; and A response message indicating whether to join the joint learning process is sent via the network interface.
25. The client computing node of claim 24, wherein: The joint learning circuit is further configured to implement the method of any one of claims 21 to 22.