Method and apparatus for vertical federated learning in wireless communication network
Vertical Federated Learning addresses data privacy and optimization challenges in wireless networks by enabling collaborative, privacy-preserving machine learning across distributed entities, enhancing network performance and predictive analytics.
Patent Information
- Application Number
- PCT/KR2025/002194
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-01-23
- Filing Date
- 2025-02-14
- Publication Date
- 2025-08-21
AI Technical Summary
Existing wireless communication networks face challenges in efficiently managing data analytics and machine learning across distributed network entities while preserving data privacy and optimizing network performance.
Implementing Vertical Federated Learning (VFL) across network entities with advanced data alignment, model parameter optimization, and iterative learning to facilitate collaborative, privacy-preserving machine learning without direct data sharing.
Enhances network performance, improves predictive models and analytics, and ensures stringent data privacy standards, leading to optimized network management and service quality.
Smart Images

Figure KR2025002194_21082025_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR VERTICAL FEDERATED LEARNING IN WIRELESS COMMUNICATION NETWORK
[0001] Certain examples of the present disclosure provide one or more techniques for supporting Vertical federated Learning (VFL). For example, certain examples of the present disclosure provide one or more techniques for enhancing Network Exposure Function (NEF) to support VFL in a 3rdGeneration Partnership Project (3GPP) 5thGeneration (5G) New Radio (NR) network. Additionally, certain examples of the present disclosure provide methods and apparatus for supporting vertical federated learning (VFL) at Network Data Analytics Functions (NWDAFs)
[0002] 5G mobile communication technologies define broad frequency bands such that high transmission rates and new services are possible, and can be implemented not only in "Sub 6GHz" bands such as 3.5GHz, but also in "Above 6GHz" bands referred to as mmWave including 28GHz and 39GHz. In addition, it has been considered to implement 6G mobile communication technologies (referred to as Beyond 5G systems) in terahertz bands (for example, 95GHz to 3THz bands) in order to accomplish transmission rates fifty times faster than 5G mobile communication technologies and ultra-low latencies one-tenth of 5G mobile communication technologies.
[0003] At the beginning of the development of 5G mobile communication technologies, in order to support services and to satisfy performance requirements in connection with enhanced Mobile BroadBand (eMBB), Ultra Reliable Low Latency Communications (URLLC), and massive Machine-Type Communications (mMTC), there has been ongoing standardization regarding beamforming and massive MIMO for mitigating radio-wave path loss and increasing radio-wave transmission distances in mmWave, supporting numerologies (for example, operating multiple subcarrier spacings) for efficiently utilizing mmWave resources and dynamic operation of slot formats, initial access technologies for supporting multi-beam transmission and broadbands, definition and operation of BWP (BandWidth Part), new channel coding methods such as a LDPC (Low Density Parity Check) code for large amount of data transmission and a polar code for highly reliable transmission of control information, L2 pre-processing, and network slicing for providing a dedicated network specialized to a specific service.
[0004] Currently, there are ongoing discussions regarding improvement and performance enhancement of initial 5G mobile communication technologies in view of services to be supported by 5G mobile communication technologies, and there has been physical layer standardization regarding technologies such as V2X (Vehicle-to-everything) for aiding driving determination by autonomous vehicles based on information regarding positions and states of vehicles transmitted by the vehicles and for enhancing user convenience, NR-U (New Radio Unlicensed) aimed at system operations conforming to various regulation-related requirements in unlicensed bands, NR UE Power Saving, Non-Terrestrial Network (NTN) which is UE-satellite direct communication for providing coverage in an area in which communication with terrestrial networks is unavailable, and positioning.
[0005] Moreover, there has been ongoing standardization in air interface architecture / protocol regarding technologies such as Industrial Internet of Things (IIoT) for supporting new services through interworking and convergence with other industries, IAB (Integrated Access and Backhaul) for providing a node for network service area expansion by supporting a wireless backhaul link and an access link in an integrated manner, mobility enhancement including conditional handover and DAPS (Dual Active Protocol Stack) handover, and two-step random access for simplifying random access procedures (2-step RACH for NR). There also has been ongoing standardization in system architecture / service regarding a 5G baseline architecture (for example, service based architecture or service based interface) for combining Network Functions Virtualization (NFV) and Software-Defined Networking (SDN) technologies, and Mobile Edge Computing (MEC) for receiving services based on UE positions.
[0006] As 5G mobile communication systems are commercialized, connected devices that have been exponentially increasing will be connected to communication networks, and it is accordingly expected that enhanced functions and performances of 5G mobile communication systems and integrated operations of connected devices will be necessary. To this end, new research is scheduled in connection with eXtended Reality (XR) for efficiently supporting AR (Augmented Reality), VR (Virtual Reality), MR (Mixed Reality) and the like, 5G performance improvement and complexity reduction by utilizing Artificial Intelligence (AI) and Machine Learning (ML), AI service support, metaverse service support, and drone communication.
[0007] Furthermore, such development of 5G mobile communication systems will serve as a basis for developing not only new waveforms for providing coverage in terahertz bands of 6G mobile communication technologies, multi-antenna transmission technologies such as Full Dimensional MIMO (FD-MIMO), array antennas and large-scale antennas, metamaterial-based lenses and antennas for improving coverage of terahertz band signals, high-dimensional space multiplexing technology using OAM (Orbital Angular Momentum), and RIS (Reconfigurable Intelligent Surface), but also full-duplex technology for increasing frequency efficiency of 6G mobile communication technologies and improving system networks, AI-based communication technology for implementing system optimization by utilizing satellites and AI (Artificial Intelligence) from the design stage and internalizing end-to-end AI support functions, and next-generation distributed computing technology for implementing services at levels of complexity exceeding the limit of UE operation capability by utilizing ultra-high-performance communication and computing resources.
[0008] Various acronyms, abbreviations and definitions used in the present disclosure are defined at the end of this description.
[0009] The content of the following documents is referred to below and / or their content provides background information and context that the following disclosure should be considered in view of:
[0010] [1] SP-231800, Study on Core Network Enhanced Support for Artificial Intelligence (AI) / Machine Learning (ML), 3GPP TSG SA Meeting #102, 11-15 December 2023, Edinburgh, UK.
[0011] [2] 3GPP TS 23.288 V18.4.0, Architecture enhancements for 5G System (5GS) to support network data analytics services.
[0012] [3] 3GPP TR 23.700-84 V0.1.0, Study on Core Network Enhanced Support for Artificial Intelligence (AI) / Machine Learning (ML).
[0013] (Note: the example versions shown for each TS are non-limiting, other versions of the TS may be considered also)
[0014] Wireless or mobile (cellular) communications networks in which a mobile terminal (e.g., user equipment (UE), such as a mobile handset) communicates via a radio link with a network of base stations, or other wireless access points or nodes, have undergone rapid development through a number of generations. The 3rdGeneration Partnership Project (3GPP) design, specify and standardise technologies for mobile wireless communication networks. Fourth Generation (4G) and Fifth Generation (5G) systems are now widely deployed, and development of Sixth Generation (6G) Systems is in progress.
[0015] 3GPP standards for 4G systems include an Evolved Packet Core (EPC) and an Enhanced-UTRAN (E-UTRAN: an Enhanced Universal Terrestrial Radio Access Network). The E-UTRAN uses Long Term Evolution (LTE) radio technology. LTE is commonly used to refer to the whole system including both the EPC and the E-UTRAN, and LTE is used in this sense in the remainder of this document. LTE should also be taken to include LTE enhancements such as LTE Advanced and LTE Pro, which offer enhanced data rates compared to LTE.
[0016] In 5G systems a new air interface has been developed, which may be referred to as 5G New Radio (5G NR) or simply NR. NR is designed to support the wide variety of services and use case scenarios envisaged for 5G networks, though builds upon established LTE technologies. New frameworks and architectures are also being developed as part of 5G networks in order to increase the range of functionality and use cases available through 5G networks.
[0017] 3GPP has also started studying the benefits of introducing Artificial Intelligence (AI) / Machine Learning (ML) solutions to communications networks, for example, enhancement of management and orchestration, performance, resource allocation, in addition to reduction of complexity and overhead in the network
[0018] In AI / ML operation, AI / ML models and / or data might be transferred across the AI / ML applications (AFs), 5GC and UEs. The AI / ML works could be divided into two main phases: model training and inference. During model training and inference, multiple rounds of interaction may be required.
[0019] The above information is presented as background information only to assist with an understanding of the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above might be applicable as prior art with regard to the present invention.
[0020] In the rapidly evolving landscape of telecommunications, future 6G networks promise an era of unprecedented connectivity, supporting a wide array of devices and applications, from smartphones to Internet of Things (IoT) devices, and autonomous vehicles. This technological leap forward come with a surge in data volume, variety, and velocity, presenting both opportunities and challenges in data management, privacy, and utilization. Central to harnessing the potential of this data is the ability to analyze and derive actionable insights in real-time, across diverse network environments and geographies.
[0021] Various approaches of the present disclosure address these challenges by introducing an innovative approach to data analytics and machine learning model training within the context of Vertical Federated Learning (VFL) across multiple network entities. This approach is designed to optimize network performance, enhance user experience, and maintain / support stringent data privacy standards, all within the framework established by the 3rd Generation Partnership Project (3GPP) for networks and services.
[0022] More specifically, the present disclosure provides a novel method and system for implementing VFL among distributed network entities, enabling collaborative, privacy-preserving machine learning without or with a reduced need for direct data sharing. This is achieved through a sophisticated orchestration of data processing, model training, and optimization processes that leverage the unique capabilities of network entities / components across different network segments and service areas. By doing so, the proposed approach facilitates a more efficient and effective utilization of network data, leading to improved predictive models and analytics that are valuable for network management, service quality assurance, and the development of new services.
[0023] An important aspect of various approaches of the present disclosure is the introduction of advanced mechanisms for data alignment, model parameter optimization, and iterative learning, which collectively enhance the accuracy and reliability of the federated models developed through this process. Furthermore, the disclosed approaches outline a structured procedure for the participation of network entities in the VFL process, including registration, model request and provisioning, participant selection, data alignment and processing, as well as model evaluation and feedback mechanisms. These procedures help to ensure that each participating entity can contribute to and benefit from the federated learning process, despite the inherent challenges of distributed data and the need for privacy preservation.
[0024] Embodiments or examples disclosed in the description and / or figures falling outside the scope of the claims are to be understood as examples useful for understanding the present invention.
[0025] According to an aspect of the present disclosure, there is provided a method of a vertical federated learning (VFL) server for a VFL inference procedure in a wireless communication network, the method comprising: transmitting, to one or more VFL clients, a VFL inference request message including a VFL correlation identity (ID); receiving, from the one or more VFL clients, a response message including an intermediate inference result; and generating, a VFL inference result by aggregating the intermediate inference result based on the VFL correlation ID.
[0026] According to an aspect of the present disclosure, there is provided a vertical federated learning (VFL) server for a VFL inference procedure in a wireless communication network, the VFL server comprising: a transceiver; and a processor, wherein the processor is configured to: transmit, to one or more VFL clients, a VFL inference request message including a VFL correlation identity (ID); receive, from the one or more VFL clients, a response message including an intermediate inference result; and generate, a VFL inference result by aggregating the intermediate inference result based on the VFL correlation ID.
[0027] According to an aspect of the present disclosure, there is provided a method of a vertical federated learning (VFL) client for a VFL inference procedure in a wireless communication network, the method comprising: receiving, from a VFL server, a VFL inference request message including a VFL correlation identity (ID); generating an intermediate inference result based on the VFL correlation ID; and transmitting, to the VFL server, a response message including the intermediate inference result.
[0028] According to an aspect of the present disclosure, there is provided a server entity for performing a vertical federated learning (VFL) procedure, the server entity configured to: based on a message received from a service consumer, initiate the VFL procedure; transmit a request for VFL inference to at least one VFL client entity; receive intermediate inference results generated by the at least one VFL client entity; and perform inference computation on the received intermediate inference results.
[0029] According to various examples, the server entity is further configured to: discover or select the at least one VFL client entity to participate in the VFL procedure.
[0030] According to various examples, the server entity is further configured to: aggregate the received intermediate inference results; and generate inference results based on the aggregated intermediate inference results.
[0031] According to various examples, the server entity is further configured to generate the inference results for an analytics ID received from the service consumer.
[0032] According to various examples, the server entity is further configured to: collect local data and generate intermediate local inference results based on the local data; and aggregate the intermediate local inference results with the received intermediate inference results.
[0033] According to various examples, the server entity is further configured to: derive analytics according to the message, based on the generated inference results; and transmit the derived analytics to the service consumer.
[0034] According to various examples, based on the at least one VFL client entity being an application function (AF), the request is transmitted to the at least one VFL client entity via a network exposure function (NEF) entity, and the intermediate inference results are received from the at least one VFL client entity via the NEF entity.
[0035] According to various examples, based on the server entity being untrusted, the request is transmitted to a network exposure function (NEF) entity to be forwarded to the at least one VFL client entity, and the intermediate inference results are received from the NEF entity.
[0036] According to various examples, transmitting the request for VFL inference to be forwarded to the at least one VFL client entity comprises transmitting, to the NEF entity, a request for VFL inference for each one of the at least one VFL client entity, wherein each request is to be forwarded to a corresponding one of the at least one VFL client entity.
[0037] According to various examples, transmitting the request for VFL inference to be forwarded to the at least one VFL client entity comprises transmitting, to the NEF entity, a single request for VFL inference, wherein the request includes a correlation ID or an indication of the at least one VFL client entity.
[0038] According to various examples, the request comprises an ID (e.g. the correlation ID) indicating a VFL process including the VFL inference and / or VFL training.
[0039] According to various examples, wherein the message is an analytics request received from the service consumer.
[0040] According to various examples, the server entity is further configured to register, to network repository function (NRF), information including at least one of network function (NF) profile, analytics ID(s), service area, VFL capability information, and time interval supporting VFL.
[0041] According to various examples, the server entity is further configured to transmit a VFL preparation request to one or more VFL client, the VFL preparation request including at least one of analytics ID, machine learning (ML) model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met.
[0042] According to various examples, the server entity is further configured to transmit dataset identifiers to the one or more VFL client entity.
[0043] According to various examples, the server entity is further configured to: receive, from each of the one or more VFL client entity, information including an indication of whether said VFL client entity will participate; and select the at least one VFL client entity from among the one or more VFL client entity based on the received information.
[0044] According to various examples, the received information further comprises a reason why said VFL client entity cannot participate.
[0045] According to various examples, the server entity is further configured to: transmit, to the at least one VFL client entity, a request to perform ML model training; and receive, from the at least one VFL client entity, intermediate training results.
[0046] According to various examples, the server entity is further configured to perform VFL computation based on the received intermediate training results.
[0047] According to various examples, the intermediate training results relate to a ML model associated with the VFL inference.
[0048] According to another aspect of the present disclosure, there is provided a vertical federated learning (VFL) client entity for participating in a VFL procedure, the VFL client entity configured to: receive a request for VFL inference from a server entity; based on the request, obtain intermediate inference results using a local machine learning (ML) model; and transmit the intermediate inference results to the server entity.
[0049] According to various examples, based on the at least one VFL client entity being an application function (AF), the request is received from the server entity via a network exposure function (NEF) entity, and the intermediate inference results are transmitted to the server entity via the NEF entity.
[0050] According to various examples, the request is received from the server entity via a network exposure function (NEF) entity; and wherein the intermediate inference results are transmitted to the server entity via the NEF entity based on the server entity being untrusted.
[0051] According to various examples, the VFL client entity is further configured to: collect local data based on determining stored data is not sufficient to support the VFL inference of the request; and generate the intermediate inference results based on the collected local data.
[0052] According to various examples, the VFL client entity is further configured to determine the local data to collect based on an analytics ID included in the request for VFL inference or an indication received from the server entity via the NEF entity.
[0053] According to various examples, the VFL client entity is further configured to register, to network repository function (NRF), information including at least one of network function (NF) profile, analytics ID(s), service area, VFL capability information, and time interval supporting VFL.
[0054] According to various examples, the VFL client entity is further configured to receive a VFL preparation request from the server entity, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met.
[0055] According to various examples, the VFL client entity is further configured to: based on the VFL preparation request, determine capability to meet model training requirements; and transmit, to the server entity, information including an indication of whether the VFL client entity will participate in the VFL procedure.
[0056] According to various examples, the VFL client entity is further configured to receive dataset identifier from the server entity.
[0057] According to various examples, the transmitted information further includes a reason why the VFL client entity cannot participate.
[0058] According to various examples, the VFL client entity is further configured to: receive, from the server entity, a request to perform ML model training; and train the local ML model; wherein the trained local ML model is used to obtain the intermediate inference results.
[0059] According to various examples, the VFL client entity is further configured to: obtain intermediate training results of the trained local ML model, and share the intermediate training results with the server entity; or refine a previously trained model to obtain the trained local ML model and share the model refinement results with the server entity.
[0060] According to various examples, the intermediate training results or the model refinement results are shared using a same ML model training service as used for receiving the request to perform ML model training.
[0061] According to various examples, the VFL client entity is further configured to share the intermediate training results with another VFL client entity.
[0062] According to another aspect of the present disclosure, there is provided a network exposure function (NEF) entity configured to: receive, from a server entity, a request for vertical federated learning (VFL) inference to be forwarded to at least one VFL client entity; forward the request for VFL inference to the at least one VFL client entity; receive, from the at least one VFL client entity, intermediate inference results generated by the at least one VFL client entity; and forward the intermediate inference results to the server entity.
[0063] According to various examples, the server entity is untrusted and / or the at least one VFL client entity is an application function (AF).
[0064] According to various examples, receiving the request for VFL inference comprises receiving, from the server entity, at least one request for VFL inference each corresponding to one of the at least one VFL client entity; and wherein forwarding the request for VFL inference to the at least one VFL client entity comprises forwarding each of the at least one request for VFL inference to the corresponding one of the at least one VFL client entity.
[0065] According to various examples, receiving the request for VFL inference comprises receiving, from the server entity, a single request for VFL inference, wherein the request includes a correlation ID or an indication of the at least one VFL client entity; and wherein the NEF entity is configured to: select the at least one VFL client entity to forward the request for VFL inference to, based on the correlation ID or the indication of the at least one VFL client entity, wherein forwarding the request for VFL inference comprises forwarding the request for VFL inference to the selected at least one VFL client entity; or split the request for VFL inference into at least one request for VFL inference, each of the at least one request for VFL inference corresponding to one of the at least one VFL client entity, wherein forwarding the request for VFL inference comprises forwarding each of the at least one request for VFL inference to the corresponding one of the at least one VFL client entity.
[0066] According to various examples, the request comprises an ID (e.g. the correlation ID) indicating a VFL process including VFL inference and VFL training, and wherein selecting the at least one VFL client entity is based on the VFL inference and VFL training indicated by the ID.
[0067] According to various examples, the NEF entity is further configured to enable the server entity to subscribe, unsubscribe, notify and / or modify for artificial intelligence / machine learning (AI / ML) training toward the at least one VFL client entity via the NEF.
[0068] According to another aspect of the present disclosure there is provided a server entity for performing a vertical federated learning (VFL) procedure, the server entity configured to: transmit a VFL preparation request to one or more VFL client, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met; receive, from each of the one or more VFL client entity, information including an indication of whether said VFL client entity will participate; and select at least one VFL client entity to participate from among the one or more VFL client entity based on the received information .
[0069] According to various examples, the at least one VFL client entity is selected to participate in VFL training and / or VFL inference.
[0070] According to various examples, the server entity is further configured to transmit dataset identifier to the one or more VFL client entity.
[0071] According to various examples, the received information further comprises a reason why said VFL client entity cannot participate.
[0072] According to another aspect of the present disclosure there is provided a server entity for performing a vertical federated learning (VFL) procedure, the server entity configured to: transmit, to at least one VFL client entity, a request to perform machine learning (ML) model training; receive, from the at least one VFL client entity, intermediate training results or model refinement results; and perform VFL computation based on the received results.
[0073] According to various examples, the intermediate training results relate to a ML model associated with a VFL inference process.
[0074] According to another aspect of the present disclosure there is provided a vertical federated learning (VFL) client entity for participating in a VFL procedure, the VFL client entity configured to: receive a VFL preparation request from a server entity, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met; based on the VFL preparation request, determine capability to meet model training requirements; and transmit, to the server entity, information including an indication of whether the VFL client entity will participate in the VFL procedure.
[0075] According to various examples, the model training requirements relate to VFL training and / or VFL inference.
[0076] According to various examples, the transmitted information further includes a reason why the VFL client entity cannot participate.
[0077] According to various examples, the VFL client entity is further configured to receive dataset identifier from the server entity.
[0078] According to another aspect of the present disclosure, there is provided a vertical federated learning (VFL) client entity for participating in a VFL procedure, the VFL client entity configured to: receive, from a server entity, a request to perform machine learning (ML) model training; and train a local ML model; wherein training the local ML model comprises obtaining intermediate training results of the trained local ML model or refining a previously trained model to obtain the trained local ML model; and wherein VFL client entity is configured to share the intermediate training results or the model refinement results with the server entity.
[0079] According to another aspect of the present disclosure, there is provided a server entity for performing a vertical federated learning (VFL) procedure, the server entity configured to: register, to network repository function (NRF), information including at least one of network function (NF) profile, analytics ID(s), service area, VFL capability information, and time interval supporting VFL; and discover one or more VFL client entity via the NRF by invoking a discovery request service operation.
[0080] According to another aspect of the present disclosure, there is provided a vertical federated learning (VFL) client entity for participating in a VFL procedure, the VFL client entity configured to: register, to network repository function (NRF), information including at least one of network function (NF) profile, analytics ID(s), service area, VFL capability information, and time interval supporting VFL.
[0081] According to another aspect of the present disclosure, there is provided a method of a server entity for performing a vertical federated learning (VFL) procedure, the method comprising: based on a message received from a service consumer, initiating the VFL procedure; transmitting a request for VFL inference to at least one VFL client entity; receiving intermediate inference results generated by the at least one VFL client entity; and performing inference computation on the received intermediate inference results.
[0082] According to another aspect of the present disclosure, there is provided a method of a vertical federated learning (VFL) client entity for participating in a VFL procedure, the method comprising: receiving a request for VFL inference from a server entity; based on the request, obtaining intermediate inference results using a local machine learning (ML) model; and transmitting the intermediate inference results to the server entity.
[0083] According to another aspect of the present disclosure, there is provided a method of a network exposure function (NEF) entity, the method comprising: receiving, from a server entity, a request for vertical federated learning (VFL) inference to be forwarded to at least one VFL client entity; forwarding the request for VFL inference to the at least one VFL client entity; receiving, from the at least one VFL client entity, intermediate inference results generated by the at least one VFL client entity; and forwarding the intermediate inference results to the server entity.
[0084] According to another aspect of the present disclosure, there is provided a method of a server entity for performing a vertical federated learning (VFL) procedure, the method comprising: transmitting a VFL preparation request to one or more VFL client, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met; receiving, from each of the one or more VFL client entity, information including an indication of whether said VFL client entity will participate; and selecting at least one VFL client entity to participate from among the one or more VFL client entity based on the received information.
[0085] According to another aspect of the present disclosure, there is provided a method of a vertical federated learning (VFL) client entity for participating in a VFL procedure, the method comprising: receiving a VFL preparation request from a server entity, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met; based on the VFL preparation request, determining capability to meet model training requirements; and transmitting, to the server entity, information including an indication of whether the VFL client entity will participate in the VFL procedure.
[0086] According to another aspect of the present disclosure, there is provided a method of a server entity for performing a vertical federated learning (VFL) procedure, the method comprising: transmitting, to at least one VFL client entity, a request to perform machine learning (ML) model training; receiving, from the at least one VFL client entity, intermediate training results or model refinement results; and performing VFL computation based on the received results.
[0087] According to another aspect of the present disclosure, there is provided a method of a vertical federated learning (VFL) client entity for participating in a VFL procedure, the method comprising: receiving, from a server entity, a request to perform machine learning (ML) model training; and training a local ML model; wherein training the local ML model comprises obtaining intermediate training results of the trained local ML model or refining a previously trained model to obtain the trained local ML model; and wherein the method comprises sharing the intermediate training results or the model refinement results with the server entity.
[0088] It is an aim of certain examples of the present disclosure to address, solve and / or mitigate, at least partly, at least one of the problems and / or disadvantages associated with the related art, for example at least one of the problems and / or disadvantages described herein. It is an aim of certain examples of the present disclosure to provide at least one advantage over the related art, for example at least one of the advantages described herein.
[0089] Other aspects, advantages and salient features of the disclosure will become apparent to those skilled in the art from the following detailed description taken in conjunction with the accompanying drawings.
[0090] Embodiments / examples of the present disclosure are further described hereinafter with reference to the accompanying drawings, in which:
[0091] Figure 1 is a reproduction of Figure 6.2C.2.1-1: Registration and Discovery procedure for Federated Learning in 3GPP TS 23.288;
[0092] Figure 2 is a reproduction of Figure 6.2C.2.2-1: General procedure for Federated Learning among Multiple NWDAF in 3GPP TS 23.288;
[0093] Figure 3 is a reproduction of Figure 6.2C.2.3-1: Procedure of FL Server NWDAF reselects FL Client NWDAF(s), FL Client NWDAF(s) Join or Leave Federated Learning Process Dynamically in Federated Learning execution phase in 3GPP TS 23.288;
[0094] Figure 4A, 4B, and 4C illustrate an exemplary VFL model training procedure;
[0095] Figure 5 illustrates an exemplary VFL inference procedure;
[0096] Figure 6 illustrates exemplary intermediate training / inference results from VFL clients;
[0097] Figure 7A, and 7B provide a procedure for the support of Vertical Federated Learning (VFL) at NWDAF in accordance with an approach of the present disclosure;
[0098] Figure 8A, 8B, and 8C provide a procedure for VFL with NWDAF and Application Function (AF) as Participants; and
[0099] Figure 9 is a block diagram of an exemplary network entity / function that may be used in certain examples of the present disclosure.
[0100] Figure 10 is a flow diagram illustrating a method in accordance with an example of the present disclosure.
[0101] Figure 11 is a flow diagram illustrating a method in accordance with an example of the present disclosure.
[0102] Figure 12 is a flow diagram illustrating a method in accordance with an example of the present disclosure.
[0103] Figure 13 is a flow diagram illustrating a method in accordance with an example of the present disclosure.
[0104] Figure 14 is a flow diagram illustrating a method in accordance with an example of the present disclosure.
[0105] Figure 15 is a flow diagram illustrating a method in accordance with an example of the present disclosure.
[0106] Figure 16 is a flow diagram illustrating a method accordance with an example of the present disclosure.
[0107] The following description of examples of the present disclosure, with reference to the accompanying drawings, is provided to assist in a comprehensive understanding certain examples of the present disclosure. The description includes various specific details to assist in that understanding but these are to be regarded as merely exemplary. Accordingly, those of ordinary skill in the art will recognize that various changes and modifications of the examples described herein can be made without departing from the scope of the disclosure.
[0108] The same or similar components may be designated by the same or similar reference numerals, although they may be illustrated in different drawings.
[0109] Detailed descriptions of techniques, structures, constructions, functions, operations or processes known in the art may be omitted for clarity and conciseness, and to avoid obscuring the subject matter of the present disclosure.
[0110] The terms and words used herein are not limited to the bibliographical or standard meanings, but, are merely used to enable a clear and consistent understanding of the disclosure.
[0111] Throughout the description and claims of this specification, the words "comprise", "include" and "contain" and variations of the words, for example "comprising" and "comprises", means "including but not limited to", and is not intended to (and does not) exclude other features, elements, components, integers, steps, processes, operations, functions, characteristics, properties and / or groups thereof.
[0112] Throughout the description and claims of this specification, the singular form, for example "a", "an" and "the", encompasses the plural unless the context otherwise requires. For example, reference to "an object" includes reference to one or more of such objects.
[0113] Throughout the description and claims, the expression "at least one of A, B and / or C" (or the like) and the expression "one or more of A, B and / or C" (or the like) should be seen to separately include all possible combinations, for example: A, B, C, A and B, A and C, A and B and C.
[0114] Throughout the description and claims of this specification, language in the general form of "X for Y" (where Y is some action, process, operation, function, activity or step and X is some means for carrying out that action, process, operation, function, activity or step) encompasses means X adapted, configured or arranged specifically, but not necessarily exclusively, to do Y.
[0115] Features, elements, components, integers, steps, processes, operations, functions, characteristics, properties and / or groups thereof described or disclosed in conjunction with a particular aspect, embodiment, example or claim are to be understood to be applicable to any other aspect, embodiment, example or claim described herein unless incompatible therewith.
[0116] The skilled person will appreciate that the techniques described herein may be used in any suitable combination.
[0117] Certain examples of the present disclosure provide one or more techniques for supporting VFL. For example, certain examples of the present disclosure provide one or more techniques for enhancing NEF to support VFL in a 3GPP 5G NR network. However, the skilled person will appreciate that the present invention is not limited to these examples, and may be applied in any suitable system or standard, for example one or more existing and / or future generation wireless communication systems or standards, including any existing or future releases of the same standards specification, for example 3GPP 5G, 5G-advanced or 6thGeneration (6G).
[0118] The functionality of the various network entities and other features disclosed herein may be applied to corresponding or equivalent entities or features in the same or any other suitable communication systems or standards. Corresponding or equivalent entities or features may be regarded as entities or features that perform the same or similar role, function or purpose within the network.
[0119] For example, the functionality of a base station or the like (e.g. eNB, gNB, NB, RAN node, access point, wireless point, transmission / reception point, central unit, distributed unit, radio unit, remote radio head, etc.) in the examples below may be applied to any other suitable type of entity performing RAN functions, and the functionality of a UE or the like (e.g. electronic device, user device, mobile station, subscriber station, customer premises equipment, terminal, remote terminal, wireless terminal, vehicle terminal, etc.) in the examples below may be applied to any other suitable type of device.
[0120] A particular network entity may be implemented as a network element on a dedicated hardware, as a software instance running on a dedicated hardware, and / or as a virtualised function instantiated on an appropriate platform, e.g. on a cloud infrastructure.
[0121] The following examples are applicable to, and use terminology associated with, 3GPP 4G (e.g., LTE) and / or 5G (e.g., NR). However, the skilled person will appreciate that the techniques disclosed herein are not limited to these examples or to 3GPP 4G (e.g., LTE) and / or 5G (e.g., NR), and may be applied in any suitable system or standard, for example one or more existing and / or future generation wireless communication systems or standards (e.g., B5G, 5G-Advanced, 6G etc.). The skilled person will appreciate that the techniques disclosed herein may be applied in any existing or future releases of 3GPP 4G (e.g., LTE) and / or 5G (e.g., NR) and / or 5G Advanced and / or 6G, and / or (3GPP Release 17, 18, 19, 20, etc.) or any other relevant standard. For example, the functionality of the various network entities and other features disclosed herein may be applied to corresponding or equivalent entities or features in other communication systems or standards. Corresponding or equivalent entities or features may be regarded as entities or features that perform the same or similar role, function, operation or purpose within the network.
[0122] Furthermore. the following also applies to the present disclosure:
[0123] · The terms functionality / use-case / configuration / scenario / site may be used interchangeably.
[0124] · The terms model and model functionality may be used interchangeably.
[0125] · This disclosure also apply to non-3GPP entities.
[0126] · The concepts, proposals, solutions, methods, embodiments, figures, and / or examples, presented in this disclosure, would apply to various type of communication systems, such as 4G, 4G-Advanced, 5G, 5G-Advanced, and 6G. Moreover, the above may also apply (in full or part or modified) to systems of Non-Terrestrial Networks (NR-NTN and / or IoT-NTN and / or UAV, etc.), in addition to Terrestrial Networks (TN).
[0127] The skilled person will appreciate that the present disclosure is not limited to the specific examples disclosed herein. For example:
[0128] · The techniques disclosed herein are not limited to 3GPP 4G or 5G or 5G-Advanced and also apply to B5G and 6G systems.
[0129] · One or more entities in the examples disclosed herein may be replaced with one or more alternative entities performing equivalent or corresponding functions, processes or operations.
[0130] · One or more of the messages in the examples disclosed herein may be replaced with one or more alternative messages, signals or other type of information carriers that communicate equivalent or corresponding information.
[0131] · One or more further elements, entities and / or messages may be added to the examples disclosed herein.
[0132] · One or more non-essential elements, entities and / or messages may be omitted in certain examples.
[0133] · The functions, processes or operations of a particular entity in one example may be divided between two or more separate entities in an alternative example.
[0134] · The functions, processes or operations of two or more separate entities in one example may be performed by a single entity in an alternative example.
[0135] · Information carried by a particular message in one example may be carried by two or more separate messages in an alternative example.
[0136] · Information carried by two or more separate messages in one example may be carried by a single message in an alternative example.
[0137] · The order in which operations are performed and / or the order in which messages are transmitted may be modified, if possible, in alternative examples.
[0138] · The transmission of information between network entities is not limited to the specific form, type and / or order of messages described in relation to the examples disclosed herein.
[0139] Certain examples of the present disclosure may be provided in the form of an apparatus / device / network entity configured to perform one or more defined network functions and / or a method therefor. Such an apparatus / device / network entity may comprise one or more elements, for example one or more of receivers, transmitters, transceivers, processors, controllers, modules, units, and the like, each element configured to perform one or more corresponding processes, operations and / or method steps for implementing the techniques described herein. For example, an operation / function of X may be performed by a module configured to perform X (or an X-module). Certain examples of the present disclosure may be provided in the form of a system (e.g. network or wireless communication system) comprising one or more such apparatuses / devices / network entities, and / or a method therefor.
[0140] From the perspective of the operation types, AI / ML operation types may be categorised into three types: model splitting, model sharing, and distributed / federated learning.
[0141] Overview of Federated Learning Support at NWDAF
[0142] In current SA2 specifications, the Federated Learning (FL) among multiple NWDAFs (so called Horizontal Federated Learning) were supported since 3GPP Rel-18. High-level and detailed descriptions of supporting Federated Learning (FL) among multiple NWDAFs are mainly documented in clause 5.3 and clause 6.2c of TS 23.288 [2].
[0143] In Rel-16, a single instance or multiple instances of NWDAF may be deployed in a Public land mobile network (PLMN). In case multiple NWDAF instances are deployed, the architecture supports deploying the NWDAF as a central NF, as a collection of distributed NFs, or as a combination of both. When multiple NWDAFs exist, not all of them need to be able to provide the same type of analytics results. However, no specific requirement has been defined regarding how different NWDAFs could cooperate in Rel-16. In Rel-16, each NWDAF acts independently from the other NWDAFs.
[0144] In reality, some of the NWDAFs in one network may be providing the same type of analytics, and so may help each other for e.g. specific analytics for specific target UEs or specific analytics for specific area of interest. Although some of these NWDAFs may be providing different type of analytics, they may still be able to help each other if e.g. analytics are somehow related: one example is for expected UE behavioural parameters related network data analytics, which have a tight relation with UE mobility analytics and UE communication analytics. In another example, in order to build abnormal behaviour related network data analytics, the NWDAF would need to collect similar type of data to the data needed to build analytics for UE mobility pattern and for UE communication pattern.
[0145] In order to address the coordination among multiple NWDAFs for Federated Learning (FL), Rel-18 study and normative work were carried out by SA2. As documented in clause 5.3 of TS 23.288 [2]:
[0146] Federated learning among multiple NWDAFs is a machine learning technique in core network that trains an ML Model across multiple decentralized entities holding local data set, without exchanging / sharing local data set. This approach stands in contrast to traditional centralized machine learning techniques where all the local datasets are uploaded to one server, thus allowing to address critical issues such as data privacy, data security, data access rights.
[0147] For Federated Learning supported by multiple NWDAFs containing MTLF, there is one NWDAF containing MTLF acting as FL server (called FL server NWDAF for short) and multiple NWDAFs containing MTLF acting as FL client (called FL client NWDAF for short).
[0148] The FL server NWDAF and FL client NWDAF have different functionalities (in clause 5.3 of TS 23.288 [2]):
[0149] FL server NWDAF:
[0150] -discovers and selects FL client NWDAFs to participant in an FL procedure
[0151] -requests FL client NWDAFs to do local model training and to report local model information.
[0152] -generates global ML model by aggregating local model information from FL client NWDAFs.
[0153] -sends the global ML model back to FL client NWDAFs and repeats training iteration if needed.
[0154] FL client NWDAF:
[0155] -locally trains ML model that tasked by the FL server NWDAF with the available local data set, which includes the data that is not allowed to share with others due to e.g. data privacy, data security, data access rights.
[0156] -reports the trained local ML model information to the FL server NWDAF.
[0157] -receives the global ML model feedback from FL server NWDAF and repeats training iteration if needed.
[0158] Either the NWDAF containing MTLF or the NWDAF containing AnLF can trigger the ML model training, as a consumer. The NWDAF containing MTLF determines to train an ML model either based on local configuration or when it receives the request from NWDAF containing AnLF. The NWDAF containing MTLF may further determine whether the ML model should be trained via FL mechanism based on different aspects, e.g. Analytic ID, Service Area / DNAI or data cannot be obtained directly from data producer NF (e.g. due to data privacy, data security). However, the NWDAF containing AnLF is not aware whether the ML model is trained based on FL or not.
[0159] In order to perform the FL among multiple NWDAFs, before FL procedure is initiated, appropriate NWDAFs containing MTLF that can act as an FL server and FL clients should be discovered and chosen based on specified criteria and interactions between the FL server and FL clients.
[0160] When starting an FL procedure, the FL server NWDAF provides an initial model to each FL client NWDAFs, and then each FL client NWDAFs perform local model training using their local data set based on the request from FL server.
[0161] During the FL execution phase, in order to maintain a Federation Learning process, considering the performance and capability of FL Client NWDAF(s), the FL Server NWDAF may trigger reselection, addition, or removal of FL Client NWDAF(s), discovers new FL Client NWDAF(s) via NRF and FL Client NWDAF(s) joins or leaves Federated Learning process dynamically.
[0162] The detailed procedures of Registration and Discovery procedure for Federated Learning, General procedure for Federated Learning among Multiple NWDAF Instances, Procedures for Maintaining Federated Learning Processes are documented in clause 6.2C.2.1, 6.2C.2.2 and 6.2C.2.3 of TS 23.288 [2].
[0163] Registration and Discovery procedure for Federated Learning
[0164] As mentioned above, before FL procedure is initiated, appropriate FL server and FL clients discovered and chosen based on specified criteria, according to the procedures illustrated in Figure 1 (Figure 6.2C.2.1-1 in clause 6.2C.2.1 of TS 23.288 [2]):
[0165] Steps 1 to 3 are the NWDAF registration procedure.
[0166] 1-3.NWDAF containing MTLF as FL Server NWDAF or FL Client NWDAF registers to NRF with its NF profile, which includes NWDAF NF Type, Analytics ID(s), Address information of NWDAF, Service Area, FL capability type information (i.e. FL server or FL client) and Time interval supporting FL as described in clause 5.2.
[0167] Steps 4 to 6 are the NWDAF Discovery procedure.
[0168] 4-6.NWDAF containing MTLF determines ML model requires FL based on operator policy (e.g. pre-configured list of ML models), Analytic ID, Service Area / DNAI or data can not be obtained directly from data producer NF (e.g. due to privacy reasons).
[0169] If the NWDAF containing MTLF can not perform as FL Server NWDAF, the MTLF first discovers and selects FL Server NWDAF from NRF by invoking the Nnrf_NFDiscovery_Request service operation. The following criteria might be used: Analytic ID of the ML model required, Model filter information as defined in TS 23.288 [5], FL capability Type (i.e. FL server), Time Period of Interest, Service Area.
[0170] Once the FL Server NWDAF (the requested or the selected one) is determined, the FL Server NWDAF discovers and selects other NWDAF(s) containing MTLF as FL Client NWDAF(s) from NRF by invoking the Nnrf_NFDiscovery_Request service operation. The following criteria might be used: Analytic ID of the ML model required, FL capability Type (i.e. FL client), Service Area, NF type(s) of data sources from which the FL Client NWDAF is able to collect data for local model training, Time Period of Interest, ML Model Interoperability Indicator.
[0171] 7.FL Server NWDAF sends Federated Learning preparation request to the FL Client NWDAF(s), using Nnwdaf_MLModelTraining_Subscribe or Nnwdaf_MLModelTrainingInfo_Request service with the ML Preparation Flag, to check if the FL Client NWDAF(s) can meet the ML model training requirement (e.g. Analytics ID, ML Model Interoperability information, Available data requirement, Availability time requirement (time span needed for the FL process), etc.). Available data requirement includes a list of Event IDs of the local data for training, and may also include the dataset statistical properties, the time window of the data samples and the minimum number of data samples.
[0172] NOTE:Federated Learning preparation procedure (i.e. steps 7-9) can be skipped if the FL Server NWDAF can decide that the FL Client NWDAF(s) supports the FL procedure to be performed, e.g. based on information acquired from previous FL procedures or from the NRF, or based on local configuration.
[0173] 8.FL Client NWDAF(s) checks if it can meet the ML model training requirement and / or can successfully download the model if the model information is provided in the request and decides whether to join the Federated Learning process based on operator policy (e.g. pre-configured list of ML models) and / or implementation. Example criteria used by FL Client NWDAF(s) may be based on its data availability and time availability, computation and communication capability and ML Model Interoperability information.
[0174] 9.FL Client NWDAF(s) invokes Nnwdaf_MLModelTraining_Notify or Nnwdaf_MLModelTraining_Subscribe response service operation or Nnwdaf_MLModelTrainingInfo_Request response service operation to indicate to the FL Server NWDAF whether it will join the FL procedure and may include the reason in the response message if it cannot join the FL process.
[0175] 10.FL Server NWDAF determines the final list of FL Client NWDAF(s) to be involved in the FL procedures based on the information received in step 6 and other information received in step 9 (if available).
[0176] General procedure for Federated Learning among Multiple NWDAF Instances
[0177] The general procedure for Federated Learning among Multiple NWDAF is illustrated in Figure 2 (Figure 6.2C.2.2-1 in clause 6.2C.2.2 of TS 23.288 [2]):
[0178] 0.The consumer (NWDAF containing AnLF or NWDAF containing MTLF) sends a subscription request to FL server NWDAF to retrieve an ML model, using Nnwdaf_MLModelProvision service as defined in clause 7.5 including Analytics ID, ML model metric (e.g., ML model Accuracy), Accuracy reporting interval, pre-determined status (ML model Accuracy threshold or Time when the ML model is needed).
[0179] If the consumer (i.e. the NWDAF containing AnLF or NWDAF containing MTLF) provides the Time when the ML model is needed, the FL Server NWDAF can take this information into account to decide the maximum response time for its FL Client NWDAF(s).
[0180] 1.FL Server NWDAF selects NWDAF(s) containing MTLF (FL Client NWDAF(s)) as described in clause 6.2C.2.1.
[0181] 2.FL Server NWDAF sends a Nnwdaf_MLModelTraining_Subscribe or Nnwdaf_MLModelTrainingInfo_Request to the selected NWDAF containing MTLF (FL Client NWDAF(s)), which participates in the Federated learning to perform the local model training and determine the interim local ML model information based on the input parameter in the request from FL Server NWDAF. The request includes ML model metric and initial ML model and also includes the maximum response time, the FL Client NWDAF has to report the interim local ML model information to the FL Server NWDAF before the maximum response time elapses.
[0182] 3.[Optional] Each FL Client NWDAF collects its local data by using the current mechanism in clause 6.2 if the Client NWDAF has not local data available already.
[0183] 4.During Federated Learning training procedure, each FL Client NWDAF further trains the ML model provided by the FL Server NWDAF based on its own data and reports the interim local ML model information to the FL Server NWDAF in Nnwdaf_MLModelTraining_Notify or Nnwdaf_MLModelTrainingInfo_Request response. The Nnwdaf_MLModelTraining_Notify or Nnwdaf_MLModelTrainingInfo_Request response may also include the Status report of FL training that includes local ML model metric computed by the FL Client NWDAF and Training Input Data Information (e.g. areas covered by the data set, sampling ratio, maximum / minimum of value of each dimension of data, etc.) in the FL Client NWDAF. The Nnwdaf_MLModelTraining_Notify or Nnwdaf_MLModelTrainingInfo_Response also includes the global ML Model Accuracy when the ML Model Accuracy Check Flag was included in the Nnwdaf_MLModelTraining_Subscribe or Nnwdaf_MLModelTrainingInfo_Request (as described in step 7), the global ML Model Accuracy is calculated by the FL Client NWDAF using the local training data as the testing dataset.
[0184] The local ML model, which is sent from the FL Client NWDAF(s) to the FL Server NWDAF during the FL training process, is the information needed by the FL Server NWDAF to build the aggregated model.
[0185] If the FL Client NWDAF is not able to complete the training of the interim local ML model within the maximum response time provided by the FL Server NWDAF, the FL Client NWDAF shall send the Delay Event Notification that include the delay event indication, an optional cause code (e.g. local ML model training failure, more time necessary for local ML model training) and the expected time to complete the training if available to the FL Server NWDAF before the maximum response time elapses.
[0186] 4a.[Optional]If FL Server NWDAF receives notification / response that the FL Client NWDAF is not able to complete the training within the maximum response time, the FL Server NWDAF may send to the FL Client NWDAF a new maximum response time in Nnwdaf_MLModelTraining_Subscribe or Nnwdaf_MLModelTrainingInfo_Request, before which the FL Client NWDAF has to report the interim local ML model information to the FL Server NWDAF. Otherwise, the FL Server NWDAF may indicate FL Client NWDAF to skip reporting for this iteration. FL Server NWDAF includes the current iteration round ID in the message to indicate that the request is to modify the training parameters of the current iteration round.
[0187] Alternatively, the FL Server NWDAF may inform the FL Client NWDAF to cease the ML model training by sending termination request and to report back the current local ML model updates.
[0188] 5.The FL Server NWDAF aggregates all the local ML model information retrieved at step 4, to update the global ML model. The FL Server NWDAF may also compute the global ML model metric, e.g. based on the local ML model metric(s) or by applying the global model on the validation dataset (if available). The FL Server NWDAF may update the global ML model each time a FL Client NWDAF provides updated local ML model information, or the FL Server NWDAF may decide to wait for local ML model information from all FL Client NWDAFs before updating the global ML model.
[0189] If the FL Server NWDAF provides the maximum response time for the FL Client NWDAF(s) to provide the interim local ML model information in step 2, or the new maximum response time in step 4a, the FL Server NWDAF decides either to wait for the FL Client NWDAF(s) which have not yet provided their interim local ML model within the new maximum response time or to aggregate only the retrieved local ML model information instances to update global ML model. The FL Server NWDAF makes this decision, considering the notification / response from the FL Client NWDAF or, if the notification is not received, based on local configuration.
[0190] 6a.[Optional] Based on the consumer request in step 0, the FL Server NWDAF sends a Nnwdaf_MLModelProvision_Notify message to update the ML model metric to the consumer periodically (e.g. a certain number of training rounds or every 10 min) or dynamically when some pre-determined status is achieved (e.g. the ML Model Accuracy threshold is achieved or training time expires).
[0191] 6b.[Optional] The consumer decides whether the current model can fulfil the requirement, e.g. global ML model metric is satisfactory for the consumer and determines to stop or continue the training process. The consumer re-invokes Nnwdaf_MLModelProvision_Subscribe service operation as used in step 0 to continue the training process or invokes Nnwdaf_MLModelProvision_Unsubscribe service operation to stop the training process.
[0192] 6c.[Optional] Based on the subscription request sent from the consumer in step 6b, the FL Server NWDAF updates or terminates the current FL training process.
[0193] If the FL Server NWDAF received a request in step 6b to stop the Federated Training process, steps 7 and 8 are skipped.
[0194] 7.If the FL procedure continues, FL Server NWDAF may determine FL Client NWDAF as described in clause 6.2C.2.3 and sends Nnwdaf_MLModelTraining_Subscribe or Nnwdaf_MLModelTrainingInfo_Request that includes the aggregated ML model information to selected FL Client NWDAF(s) for next round of Federated Training. The request may also include the ML Model Accuracy Check Flag, that indicates the FL Client NWDAF(s) to use the local training data as the testing dataset to calculate the Model Accuracy of the global ML model provided by the FL Server NWDAF.
[0195] 8.Each FL Client NWDAF updates its own ML model based on the aggregated ML model information distributed by the FL Server NWDAF at step 7.
[0196] When the Federated Training procedure is complete, the FL Server NWDAF requests the FL client NWDAF(s) to terminate the FL procedure by invoking Nnwdaf_MLModelTraining_Unsubscribe service with a cause code that the FL process has finished and optionally with the final aggregated ML model information. Then the FL client NWDAF(s) terminate the local model training and if the final aggregated ML model information is received from the FL server NWDAF, the FL client NWDAF(s) can store it for further use.
[0197] 9.After the training process is complete, the FL Server NWDAF may send Nnwdaf_MLModelProvision_Notify that includes the globally optimal ML model information to the consumer.
[0198] Procedures for Maintaining Federated Learning Processes
[0199] In order to maintain a Federation Learning process during the FL execution phase, the FL Server NWDAF may trigger reselection, addition, or removal of FL Client NWDAF(s), discovers new FL Client NWDAF(s) via NRF and FL Client NWDAF(s) joins or leaves Federated Learning process dynamically, considering the performance and capability of FL Client NWDAF(s). The detailed procedures are shown in Figure 3 (Figure 6.2C.2.3-1 in 6.2C.2.3 of TS 23.288 [2]):
[0200] 1a.FL Server NWDAF may get the updated status of current FL Client NWDAF(s) via NRF by using Nnrf_NFManagement service (as in clause 5.2.7.2 of TS 23.502 [3]) in the Federated Learning execution phase.
[0201] FL Server NWDAF may subscribe to NRF for notifications of status changes of the current NWDAF(s) (FL Client NWDAFs 1 ... N) by invoking an Nnrf_NFManagement_NFStatusSubscribe service operation. NRF notifies the FL Server NWDAF the status changes of the current FL Client NWDAF(s) by invoking Nnrf_NFManagement_NFStatusNotify service operation(s).
[0202] The status of a current FL Client NWDAF could be availability changes, capability changes (e.g. it will not support FL anymore, etc.).
[0203] 1b.The current FL Client NWDAF(s) may inform FL Server NWDAF that it is leaving the Federated Learning process by invoking Nnwdaf_MLModelTraining_Notify service operation with Termination Request and cause code (reason for leaving, e.g. high NF load, time availability changes).
[0204] 1c.FL Server NWDAF may get the information of the new FL Client NWDAF(s) dynamically via NRF by subscribing to the event that a new FL Client NWDAF registers (Nnrf_NFManagement_NFStatusSubscribe service as in clause 5.2.7.2 of TS 23.502 [3]).
[0205] 1d.NWDAF may subscribe for NF load analytics of the FL Client NWDAF(s).
[0206] 1e.FL Client NWDAF(s) may send Status report of FL training and Global ML Model Accuracy Information by invoking Nnwdaf_MLModelTraining_Notify service.
[0207] 2.FL Server NWDAF checks FL Client NWDAF(s) status based on the received information and may determine whether reselection of FL Client NWDAF(s) for the next round(s) of Federated Learning is needed based on the received information from step 1.
[0208] 3.[If re-selection is needed as judged in step 2] If step 1c is not performed, FL Server NWDAF may discover new candidate FL Client NWDAF(s) via NRF by using Nnrf_NFDiscovery services as in clause 5.2.7.3 of TS 23.502 [3]. FL Server NWDAF reselects FL Client NWDAF(s) from the current FL Client NWDAF(s) and the new candidate FL Client NWDAF(s) (found in steps 1c or 3). For the new candidate FL Client NWDAF(s), the interaction between FL Server NWDAF and FL Client NWDAF(s) is same as the selection procedure described in clause 6.2C.2.1. The adding / deleting FL Client NWDAF(s) may happen at the end of each iteration.
[0209] 4.FL Server NWDAF sends termination request by invoking Nnwdaf_MLModelTraining_Unsubscribe service operation or Nnwdaf_MLModelTrainingInfo_Request service operation with Correlation Termination Flag to the FL Client NWDAF(s), optionally indicating the reason, e.g. FL Client NWDAF is unselected by the FL Server NWDAF for the FL process, or the FL process is suspended, etc. And FL server may also send the updated global ML model information to the unselected FL client NWDAF. FL Client NWDAF(s) terminates operations for the Federated Learning process if receive termination request from the FL Server NWDAF and may perform further action to be qualified in participation of FL training in the next cycles.
[0210] Overview of Vertical Federated Learning
[0211] 3GPP SA2 Rel-19 Study in AIML (FS_AIML_CN)
[0212] New SID on Core Network Enhanced Support for Artificial Intelligence (AI) / Machine Learning (ML) was approved in SP-231800 [1] in TSG SA Meeting #102 (Dec 2023). In WT#2 in the SID:
[0213] in SP-231800:
[0214] -WT2: Study whether and what potential enhancements are needed to enable 5G system to assist in collaborative AI / ML operation involving 5GC / NWDAF and / or AF for "Vertical Federated Learning (VFL)". The work will be based only on and limited to the scope of justified use cases.
[0215] NOTE 7: RAN and UE aspects are out of scope. Solutions based on interactions between the application client and 5GS are out of scope. The necessary communication between AF and UE application client to support the collaborative AI / ML operation is understood as no normative procedure impact. Horizontal FL procedure defined in R18 should be taken into account and reused whenever possible.
[0216] NOTE 8: coordination with SA6 is required
[0217] In WG SA2 Meeting #160-Ad Hoc-e meeting (Jan 2024), the Key Issue (KI) description of WT#2 was agreed in S2-2401830 as KI#2. The detailed description of KI#2: 5GC Support for Vertical Federated Learning was documented in clause 5.2.2 of TR 23.700-84 [3]. The issues to be addressed for KI#2 by SA2 during Rel-19 study phase include:
[0218] in S2-2401830:
[0219] This key issue aims to provide solutions for enabling 5GC support for vertical federated learning (VFL) involving NWDAF and / or AF, where no raw data need to be exchanged but some level of coordination is still required when training and inference are performed on local models. In particular, datasets used for each local model need to share the same samples while holding different features.
[0220] In Rel-18, ML model sharing between NWDAFs has been studied as a part of Horizontal Federated Learning. However, Federated learning between NWDAF and AF has not been studied (e.g. when the NWDAFs and / or AFs are in different domains, locations, regions etc).
[0221] Vertical Federated Learning (VFL) can be considered as an alternative mechanism for distributed functionalities of an ML model. Note that, as scoped in Rel-19, NWDAF and / or AF may be involved for VFL.
[0222] This Key Issue aims to study architecture enhancement to support VFL, which allows the cooperative AI / ML training and inference with the following aspects:
[0223] -Identify VFL use cases and under which conditions, and for which entities these VFL use cases show that VFL is justified to train ML models.
[0224] -Whether and how to support architecture enhancement for supporting VFL for model training and / or inference. In particular:
[0225] -Whether and how the existing NF discovery and selection needs to be enhanced.
[0226] -Whether and how ML Model training and / or inference related procedures need to be enhanced to support VFL
[0227] -Whether and how to do performance monitoring for the ML model trained via VFL
[0228] -Whether and how to provide ML Models to the participants in the VFL training process.
[0229] -How to support sample and feature alignment among the participating network entities when performing VFL
[0230] NOTE 1: Application layer-based VFL requiring communication between AFs and / or UEs application client, is out of scope.
[0231] NOTE 2: During the study on this KI, consultation with SA3 is required for handling security aspects.
[0232] NOTE 3: RAN and UE aspects are out of scope.
[0233] In order to clarify the scenarios of using Vertical Federated Learning, VFL use cases are to be identified. One possible use case of implementing VFL is to deploy NWDAF to Support for Sample and Feature Alignment in VFL:
[0234] In the AI / ML literature, VFL is a federated learning setting where multiple parties perform training on data sets that share the same sample space but differ in feature space. Because of this, an alignment in sample and feature spaces among participating entities is usually required before applying VFL. VFL further allows to perform joint training without exposing raw data or model parameters, the latter being a way in which VFL differs from HFL. TS 23.288 [X] provides NWDAF specification support for HFL but no VFL support is available.
[0235] This use case proposes NWDAF support for VFL in analytics derivation by means of sample and feature alignment between the entities participating in VFL, where the main entity facilitating the VFL operation is NWDAF and other entities may be other NWDAF instances and / or AF(s). The motivation for this use case is mainly two-fold: i) in a multi-vendor scenario, VFL may be more suitable than HFL for multiple NWDAF deployments since such accuracy increase may be achieved without the need to share model parameters among the participating NWDAF from different vendors, and ii) VFL allows an enhanced accuracy of the NWDAF predictions as models trained via VFL usually generalize better by learning from a broader feature set.
[0236] In PLMNs where multiple NWDAFs are deployed, each NWDAF instance may perform data collection locally according to their suitable data sources. Depending on the Analytics ID, the different NWDAF instances may share the sample ID space (e.g. S-NSSAI) or train on different sample ID spaces (e.g. UE IDs within their corresponding Area of Interest). Furthermore, the NWDAF instances are not all obliged to collect the same input data for the same Analytics ID as most input data is optional, thus their feature spaces may range from full to little overlap. Finally, while an AF may also participate on VFL supported by NWDAF, an alignment of samples would still be needed between the two entities, and feature alignment may also prove beneficial.
[0237] Depending on the range of overlap in sample and feature spaces of the participating entities, VFL may be a more or less suitable technique to combine models at NWDAF. Hence, support for sample and feature alignment would allow the VFL supported by NWDAF to be more effective for those scenarios that are suitable.
[0238] Based on the above background information and analysis of the give use case, compared to HFL, VFL is able to increase the accuracy of FL without sharing model parameters among the participating NWDAF from different vendors and also allows an enhanced accuracy of the NWDAF predictions as models.
[0239] In SA2 163 meeting (May 2024), the following conclusions have been documented in clause 8.2 of TR 23.700-84:
[0240] P#2.4.1: Either the NWDAF or the AF can act as VFL server and initiate VFL training process with the VFL client(s).
[0241] P#2.4.2: If an untrusted AF is involved in the VFL training process, their interactions with the NWDAF(s) are via NEF. When the NWDAF acts as VFL Server, the NWDAF can receive labels from an AF.
[0242] P#2.4.3: An identifier is allocated by a VFL server, which is used to correlate the participants during the VFL training and subsequent VFL inference processes and it is associated with the distributed ML Models in the VFL joint model training process.
[0243] P#2.4.4: VFL Clients compute the intermediate results for their local ML models involved in the VFL training and provide reports with the intermediate results to the AF or NWDAF acting as VFL server.
[0244] P#2.4.5: VFL clients may also provide intermediate results (e.g. gradient information, loss information) to other VFL clients as instructed by the VFL server.
[0245] P#2.4.6: An AF or NWDAF acting as VFL server aggregates intermediate results from VFL client(s), trains a local model, computes intermediate results based on its local ML model, and sends the intermediate results towards VFL clients involved in the joint VFL training process.
[0246] P#2.4.7: VFL server may compute different intermediate training information (e.g. gradient information, loss information) for updating its own local model and the models of VFL clients during the VFL training process after processing the received intermediate results (that may include convergence reports), sends the updates to the VFL client(s), and the VFL server / client(s) update their local ML model based on the received information. The VFL server determines when the VFL training process terminates, then it will inform the VFL Clients that the training ends.
[0247] As it has been agreed by SA2, the following issues should be addressed during SA2 Rel-19 study to support Vertical Federated Learning:
[0248] in S2-2401830:
[0249] This Key Issue aims to study architecture enhancement to support VFL, which allows the cooperative AI / ML training and inference with the following aspects:
[0250] -Identify VFL use cases and under which conditions, and for which entities these VFL use cases show that VFL is justified to train ML models.
[0251] -Whether and how to support architecture enhancement for supporting VFL for model training and / or inference. In particular:
[0252] -Whether and how the existing NF discovery and selection needs to be enhanced.
[0253] -Whether and how ML Model training and / or inference related procedures need to be enhanced to support VFL
[0254] Whether and how to do performance monitoring for the ML model trained via VFL
[0255] And also it has been concluded in SA2 163 meeting and document in clause 8.2 of TR 23.700-84 [3]:
[0256] P#2.4: For VFL training process:
[0257] P#2.4.1: Either the NWDAF or the AF can act as VFL server and initiate VFL training process with the VFL client(s).
[0258] P#2.4.2: If an untrusted AF is involved in the VFL training process, their interactions with the NWDAF(s) are via NEF. When the NWDAF acts as VFL Server, the NWDAF can receive labels from an AF.
[0259] P#2.4.3: An identifier is allocated by a VFL server, which is used to correlate the participants during the VFL training and subsequent VFL inference processes and it is associated with the distributed ML Models in the VFL joint model training process.
[0260] P#2.4.4: VFL Clients compute the intermediate results for their local ML models involved in the VFL training and provide reports with the intermediate results to the AF or NWDAF acting as VFL server.
[0261] P#2.4.5: VFL clients may also provide intermediate results (e.g. gradient information, loss information) to other VFL clients as instructed by the VFL server.
[0262] P#2.4.6: An AF or NWDAF acting as VFL server aggregates intermediate results from VFL client(s), trains a local model, computes intermediate results based on its local ML model, and sends the intermediate results towards VFL clients involved in the joint VFL training process.
[0263] P#2.4.7: VFL server may compute different intermediate training information (e.g. gradient information, loss information) for updating its own local model and the models of VFL clients during the VFL training process after processing the received intermediate results (that may include convergence reports), sends the updates to the VFL client(s), and the VFL server / client(s) update their local ML model based on the received information. The VFL server determines when the VFL training process terminates, then it will inform the VFL Clients that the training ends.
[0264] NOTE 6:Whether further specification on how the AF provides the label to NWDAF is needed in P#2.4.2 will be discussed in normative work.
[0265] NOTE 7:Whether and how P#2.4.5 is supported will be determined in normative phase.
[0266] NOTE 8:How the initial model is provided to VFL clients is to be discussed in normative phase.
[0267] P#2.5: For VFL inference process:
[0268] P#2.5.1: NWDAF acting as VFL server scenario, the NF consumer of analytics will obtain the required output based on the VFL inference process coordinated by VFL server, which is generated via the VFL inference process between VFL sever and corresponding AFs or NWDAF acting as VFL client(s):
[0269] -For NWDAF acting as VFL server, NWDAF triggers the VFL inference phase after receiving the subscription or request for analytics and delivers the analytics to the NF consumer.
[0270] P#2.5.2: For the AF acting as a VFL server, the AF acting as VFL server can also start an inference process with corresponding NWDAF acting as VFL client(s). The AF may be triggered to start the inference by a 5GC consumer (i.e. NWDAF containing AnLF). Interactions are via NEF if AF is untrusted.
[0271] P#2.5.3: Before performing the VFL inference, the NWDAF acting as VFL server determines corresponding AF(s) and / or NWDAF(s) as VFL client(s) for the inference process based on the same identifier used in the VFL training process. Related details will be specified in the normative phase.
[0272] P#2.5.4: The VFL inference process may be controlled by a set of requirements, i.e. whether the determination if all VFL participants associated with the same identifier are needed in the VFL inference process may be based e.g. on accuracy requirements, the VFL signalling and load cost, contribution weights of each client, and temporal availability of output from VFL participants.
[0273] P#2.5.5: When performing VFL model performance monitoring, inference data can be used for model retraining, which is aligned with R18.
[0274] In the following agreed principles:
[0275] P#2.4.2: If an untrusted AF is involved in the VFL training process, their interactions with the NWDAF(s) are via NEF. When the NWDAF acts as VFL Server, the NWDAF can receive labels from an AF.
[0276] P#2.5.2: For the AF acting as a VFL server, the AF acting as VFL server can also start an inference process with corresponding NWDAF acting as VFL client(s). The AF may be triggered to start the inference by a 5GC consumer (i.e. NWDAF containing AnLF). Interactions are via NEF if AF is untrusted.
[0277] When the AF acting as a VFL server, the interaction between AF and VFL client NWDAF(s) should be via NEF. However, for VFL training and inference, a large number of VFL client NWDAF(s) might be involved, which will result in high work load at NEF for mapping the information and forwarding information between untrusted AF and each VFL client NWDAF(s).
[0278] Before the SA2 164 meeting, some companies proposed to allow the NEF to integrate the VFL intermediate results from multiple NWDAF before exposing to AF. However, the extra integration at NEF may create latency and lower the efficiency of VFL training and inference.
[0279] How to minimize the work load and optimise the information forwarding mechanism by NEF for VFL to avoid significant delay should be discussed and specified in SA2 specifications.
[0280] In certain examples, the VFL server may be, for example: NWDAF, trusted AF, untrusted AF; the VFL client may be, for example: NWDAF, trusted AF, untrusted AF.
[0281] Certain examples of the present disclosure provide techniques defining how and when the NEF should expose the integrated VFL intermediate results reported from multiple VFL clients (e.g. NWDAF). Certain examples define new NEF services to support VFL operation for the scenario if the VFL server is untrusted AF.
[0282] Various examples of the present disclosure will now be described in more detail.
[0283] In certain examples, the VFL server could be either AF (untrusted or trusted) or NWDAF. The procedure illustrated in Figure 4 to support VFL model training with using untrusted AF as VFL server as an example to describe the VFL training procedure step by step below.
[0284] 0. VFL server (i.e. untrusted AF) and VFL clients (i.e. NWDAF) register to NRF. The registration may include their NF profiles, Analytics ID(s), Address information of NWDAF, Service Area, VFL capability type information (i.e. VFL server and VFL client type), VFL client coordination capability, VFL computational capability (i.e. acting as VFL server and / or VFL client ), and Time interval supporting VFL. The latter parameter can be the same as Time interval supporting FL described in clause 5.2 of TS 23.288.
[0285] The VFL server and clients are discovered via NRF by invoking the Nnrf_NFDiscovery_Request service operation. During the discovery, the VFL server may include the requirements on the VFL clients and server in the discovery request, e.g. the location of the VFL clients and server, the (minimum) available capacity of the VFL clients and server, VFL client coordination capability.
[0286] 1. The VFL server determines to initiate the VFL model training based on its internal logic and sends a VFL preparation request to the VFL client NWDAF(s) (via NEF if the VFL server AF is untrusted AF).
[0287] Alternative 1:
[0288] 1-1a. For untrusted VFL server AF, the AF invokes multiple service operation / request Nnef_MLModelTrainingInfo_Request or Nnef_MLModelTraining_Subscribe request towards the NEF.
[0289] 1-1b. Then the NEF forwards the multiple model preparation request to the corresponding VFL client NWDAF(s) by invoking Nnwdaf_MLModelTrainingInfo_Request or Nnwdaf_MLModelTraining_Subscribe service operation.
[0290] The multiple service operation / request is the number of VFL client NWDAF(s) to be involved in the VFL process. E.g. if there are 100 VFL client NWDAF(s) selected by the VFL server to participate in the VFL process (training and / or inference), the AF will invoke 100 Nnef_MLModelTrainingInfo_Request or Nnef_MLModelTraining_Subscribe request towards the NEF, then the NEF will forward the 100 requests to the 100 selected VFL client NWDAF(s) one by one.
[0291] Alternative 2:
[0292] 1-2a. For untrusted VFL server AF, the AF invokes one / single service operation / request Nnef_MLModelTrainingInfo_Request or Nnef_MLModelTraining_Subscribe request towards the NEF, no matter how many VFL client NWDAF(s) are selected to perform the VFL process (training and / or inference). The VFL server AF include the full list of the VFL clients in the request toward NEF.
[0293] 1-2b. Then the NEF forwards the request to the corresponding VFL client NWDAF(s) in the full list from VFL server AF by invoking multiple Nnwdaf_MLModelTrainingInfo_Request or Nnwdaf_MLModelTraining_Subscribe service operation towards the selected VFL client NWDAF(s).
[0294] in 1-2b the multiple service operation / request is the number of VFL client NWDAF(s) to be involved in the VFL process. E.g. if there are 100 VFL client NWDAF(s) selected by the VFL server to participate in the VFL process (training and / or inference), the AF will invoke one Nnef_MLModelTrainingInfo_Request or Nnef_MLModelTraining_Subscribe request towards the NEF, then the NEF will map the information and forward the request to the 100 selected VFL client NWDAF(s) one by one.
[0295] In alternative 2, the NEF needs to split / map the one request from VFL server AF into multiple requests and sends to the multiple VFL client NWDAF(s) selected and indicated by the AF in the AF request, e.g. sent via Nnef_MLModelTrainingInfo_Request or Nnef_MLModelTraining_Subscribe request.
[0296] 2. [optional] upon receiving the VFL training request from the VFL server (via NEF) / from the NEF, the VFL clients may collect local data for local VFL computation / model training. Whether to collect data for VFL training or not can be decided by the VFL clients based on its internal logic, e.g. the VFL client determines the existing / stored / retrieved data from storage (e.g. from ADRF) is not sufficient to support VFL training, it may determine to collect data. What type of data to collect can be decided by the VFL clients based on its internal logic (e.g. using the analytics ID received in the VFL training request, the VFL client can determine what data are needed based on the defined input data of the corresponding analytics ID), or indicated by the VFL server (via NEF), or determined based on the feature indicated by the VFL server (via NEF).
[0297] 3. VFL training / local computation at each VFL clients, triggered by the VFL server (via NEF), to calculate the intermediate results.
[0298] 4. The VFL clients NWDAF notify the intermediate results of VFL training to the VFL server.
[0299] Alternative 1:
[0300] 4-1a. If the VFL client is untrusted AF, each VFL clients (1..N) NWDAF notify the intermediate results to the VFL server via NEF. Each VFL client (1...N) NWDAF invokes Nnwdaf_MLModelTrainingInfo_Request response (if Nnwdaf_MLModelTrainingInfo_Request request is used in step 1) or Nnwdaf_MLModelTraining_Notify service operation (if Nnwdaf_MLModelTraining_Subscribe is used in step 1), based on the service operation in step 1.
[0301] 4-1b. The NEF forwards the intermediate results to the VFL server by invoking multiple NEF service operation Nnef_MLModelTrainingInfo_Request response (if Nnef_MLModelTrainingInfo_Request is used in step 1 ) or Nnef_MLModelTraining_Notify (if Nnef_MLModelTraining_Subscribe service operation is used in Step 1), based on the service operation used in step 1.
[0302] The NEF may forward the intermediate results to the VFL server once it receives the intermediate results from any of the VFL client NWDAF.
[0303] The multiple service operation / request is the number of VFL client NWDAF(s) involved in the VFL process / training. E.g. if there are 100 VFL client NWDAF(s) participate in the VFL process (training and / or inference), the NEF will invoke 100 Nnef_MLModelTrainingInfo_Request or Nnef_MLModelTraining_Subscribe request towards the VFL server.
[0304] Alternative 2:
[0305] 4-2a. If the VFL client is untrusted AF, each VFL clients (1..N) NWDAF notify the intermediate results to the VFL server via NEF. Each VFL client (1...N) NWDAF invokes Nnwdaf_MLModelTrainingInfo_Request response (if Nnwdaf_MLModelTrainingInfo_Request request is used in step 1) or Nnwdaf_MLModelTraining_Notify service operation (if Nnwdaf_MLModelTraining_Subscribe is used in step 1), based on the service operation in step 1.
[0306] 4-2b. The NEF integrates the intermediate results from one or more VFL clients NWDAF, then forward the integrated intermediate results from the one or more VFL clients NWDAF to the VFL server by invoking one NEF service operation Nnef_MLModelTrainingInfo_Request response (if Nnef_MLModelTrainingInfo_Request is used in step 1 ) or Nnef_MLModelTraining_Notify (if Nnef_MLModelTraining_Subscribe service operation is used in Step 1), based on the service operation used in step 1.
[0307] 4-2c. after a while, another one or more VFL clients NWDAF may finish the local VFL computation / training and notify the intermediate results to VFL server via the NEF.
[0308] 4-2d. The NEF integrates the intermediate results from another one or more VFL clients NWDAF in step 4-2c, then forward the integrated intermediate results from the another one or more VFL clients NWDAF to the VFL server by invoking one NEF service operation Nnef_MLModelTrainingInfo_Request response (if Nnef_MLModelTrainingInfo_Request is used in step 1 ) or Nnef_MLModelTraining_Notify (if Nnef_MLModelTraining_Subscribe service operation is used in Step 1), based on the service operation used in step 1.
[0309] 4-2a. - 4-2b. / 4-2a. - 4-2d. will repeat until the VFL training iteration is finished (e.g. all the VFL clients have reported the intermediate results) or the VFL server / NEF determines to terminate a iteration or drop the intermediate results that haven't been received for a VFL training iteration (e.g. if the configure waiting / response time is up, the time window expires, etc.)
[0310] 5. The VFL server performs further VFL computation and aggregation for local model (e.g. calculates the loss using the labels) by using the intermediate results received from VFL clients. The VFL server updates its local model, and generate intermediate results.
[0311] Alternatively, the VFL server may perform further VFL computation and aggregation after step 4-2b, once it receives the exposure of intermediate results from NEF which is the integrated intermediate results of one or more VFL clients NWDAF in step 4-2a.
[0312] The VFL server may perform further VFL computation and aggregation every time it receives the exposure of intermediate results from NEF. Alternatively, the VFL server may perform further VFL computation and aggregation once it receive a number of exposure of intermediate results from NEF. This may be that there is a configured or a set number of exposures that initiate further VFL computation and aggregation. It may also be based on a time window, where a set of exposures may be used for further VFL computation and aggregation based on the time window. In other words, all of the exposures that are received within a time window are used for further VFL computation and aggregation. This may also be considered to be periodic VFL computation and aggregation. The age of an exposure may also be used to determine whether to pursue further VFL computation and aggregation. In other words, the VFL computation and aggregation is performed within a certain time from receiving the exposure from a client. The above principles applies to any type of VFL server, e.g. the VFL server could be AF (both trusted and untrusted), or NWDAF.
[0313] Steps 1-5 are for forward computation. After the forward computation, the VFL server will trigger the subsequent interaction of VFL training towards the VFL clients to perform the backwards propagation.
[0314] Steps 6-10 are for backwards propagation for model refinement. The service operations used in Step 6-10 are the same as those in steps 1-5.
[0315] 11. The VFL server determines to terminate the VFL model training process based on its internal logic, e.g. if the model is converged.
[0316] Step 11 may also happen after Step 5. In this case, Step 6-10 will be skipped.
[0317] Alternative 1:
[0318] 11-1a. For untrusted VFL server AF, the AF invokes multiple service operation / request Nnef_MLModelTraining_Unsubscribe request towards the NEF.
[0319] 11-1b. Then the NEF forwards the multiple VFL training termination request to the corresponding VFL client NWDAF(s) by invoking Nnwdaf_MLModelTraining_Unsubscribe.
[0320] The multiple service operation / termination request is the number of VFL client NWDAF(s) to be involved in the VFL process. E.g. if there are 100 VFL client NWDAF(s) selected by the VFL server to participate in the VFL process (training and / or inference), the AF will invoke 100 Nnef_MLModelTraining_Unsubscribe request towards the NEF, then the NEF will forward the 100 requests to the 100 selected VFL client NWDAF(s) one by one by invoking Nnef_MLModelTraining_Unsubscribe.
[0321] Alternative 2:
[0322] 11-2a. For untrusted VFL server AF, the AF invokes one / single service operation / request for VFL training termination by Nnef_MLModelTraining_Unsubscribe request towards the NEF, no matter how many VFL client NWDAF(s) are selected to perform the VFL process (training and / or inference). The VFL server AF include the full list of the VFL clients in the request toward NEF.
[0323] 11-2b. Then the NEF forwards the request to the corresponding VFL client NWDAF(s) in the full list from VFL server AF by invoking multiple Nnwdaf_MLModelTraining_Unsubscribe request towards the selected VFL client NWDAF(s).
[0324] in 11-2b the multiple service operation / request is the number of VFL client NWDAF(s) involved in the VFL process. E.g. if there are 100 VFL client NWDAF(s) selected by the VFL server to participate in the VFL process (training and / or inference), the AF will invoke one Nnef_MLModelTraining_Unsubscribe request towards the NEF, then the NEF will map the information and forward the request to the 100 selected VFL client NWDAF(s) one by one.
[0325] In alternative 2, the NEF needs to split / map the one request from VFL server AF into multiple requests and sends to the multiple VFL client NWDAF(s) selected and indicated by the AF in the AF request, e.g. sent via Nnef_MLModelTraining_Unsubscribe request.
[0326] Description of VFL inference procedure
[0327] Figure 5 illustrates an exemplary VFL inference procedure. The operations of Figure 5 are described below.
[0328] 1. The VFL server determines to perform VFL model inference based on its internal logic (e.g. triggered by a service consumer) and sends a VFL inference request to the VFL client NWDAF(s) (via NEF if the VFL server AF is untrusted AF).
[0329] Alternative 1:
[0330] 1-1a. For untrusted VFL server AF, the AF invokes multiple service operation / request Nnef_AnalyticsExposure_Subscribe request towards the NEF.
[0331] 1-1b. Then the NEF forwards the multiple model preparation request to the corresponding VFL client NWDAF(s) by invoking Nnwdaf_AnalyticsSubscription_Subscribe request / service operation.
[0332] The multiple service operation / request is the number of VFL client NWDAF(s) to be involved in the VFL inference. E.g. if there are 100 VFL client NWDAF(s) selected by the VFL server to participate in the VFL inference, the AF will invoke 100 Nnef_AnalyticsExposure_Subscribe request towards the NEF, then the NEF will forward the 100 requests to the 100 selected VFL client NWDAF(s) one by one .
[0333] Alternative 2:
[0334] 1-2a. For untrusted VFL server AF, the AF invokes one / single service operation / request Nnef_AnalyticsExposure_Subscribe request towards the NEF, no matter how many VFL client NWDAF(s) are selected to perform the VFL inference. Optionally, the VFL server AF include the full list of the VFL clients in the request toward NEF.
[0335] The VFL server may include an ID for the VFL process (e.g. the correlation ID), including both VFL training and inference, to the NEF. Based on the ID for the VFL process, the NEF reuse the list of VFL clients perform the VFL training as the VFL clients to perform the VFL inference associated to the same ID.
[0336] 1-2b. Then the NEF forwards the request to the corresponding VFL client NWDAF(s) in the full list from VFL server AF by invoking multiple Nnwdaf_AnalyticsSubscription_Subscribe request service operation towards the selected VFL client NWDAF(s).
[0337] in 1-2b the multiple service operation / request is the number of VFL client NWDAF(s) to be involved in the VFL inference. E.g. if there are 100 VFL client NWDAF(s) selected by the VFL server to participate in the VFL inference, the AF will invoke one Nnef_AnalyticsExposure_Subscribe request towards the NEF, then the NEF will map the information and forward the request to the 100 selected VFL client NWDAF(s) one by one.
[0338] In alternative 2, the NEF needs to split / map the one request from VFL server AF into multiple requests and send to the multiple VFL client NWDAF(s) selected and indicated by the AF in the AF request, e.g. sent via Nnef_AnalyticsExposure_Subscribe request.
[0339] 2. [optional] upon receiving the VFL inference request from the VFL server (via NEF) / from the NEF, the VFL clients may collect local data for local VFL computation / inference. Whether to collect data for VFL training or not can be decided by the VFL clients based on its internal logic, e.g. the VFL client determines the existing / stored / retrieved data from storage (e.g. from ADRF) is not sufficient to support VFL inference, it may determine to collect data. To collect what type of data can be decided by the VFL clients based on its internal logic (e.g. using the analytics ID received in the VFL training request, the VFL client can determine what data are needed based on the defined input data of the corresponding analytics ID), or indicated by the VFL server (via NEF), or determined based on the feature indicated by the VFL server (via NEF).
[0340] 3. VFL inference / local computation at each VFL clients, triggered by the VFL server (via NEF), to calculate the intermediate inference results.
[0341] 4. The VFL clients NWDAF notify the intermediate inference results of VFL inference to the VFL server.
[0342] Alternative 1:
[0343] 4-1a. If the VFL client is untrusted AF, each VFL clients (1..N) NWDAF notify the intermediate inference results to the VFL server via NEF. Each VFL client (1...N) NWDAF invokes Nnwdaf_AnalyticsSubscription_Notify towards the NEF.
[0344] 4-1b. The NEF forwards the intermediate inference results to the VFL server by invoking multiple NEF service operation Nnef_AnalyticsExposure_Notify.
[0345] The NEF may forward the intermediate inference results to the VFL server once it receives the intermediate results from any of the VFL client NWDAF.
[0346] The multiple service operation / request is the number of VFL client NWDAF(s) involved in the VFL inference. E.g. if there are 100 VFL client NWDAF(s) participate in the VFL inference, the NEF will invoke 100 Nnef_AnalyticsExposure_Notify towards the VFL server.
[0347] Alternative 2:
[0348] 4-2a. If the VFL client is untrusted AF, each VFL clients (1..N) NWDAF notify the intermediate results to the VFL server via NEF. Each VFL client (1...N) NWDAF invokes Nnwdaf_AnalyticsSubscription_Notify towards the NEF.
[0349] 4-2b. The NEF integrates the intermediate inference results from one or more VFL clients NWDAF, then forward the integrated intermediate inference results from the one or more VFL clients NWDAF to the VFL server by invoking one NEF service operation Nnef_AnalyticsExposure_Notify.
[0350] 4-2c. after a while, another one or more VFL clients NWDAF may finish the local VFL computation / inference and notify the intermediate results to VFL server via the NEF.
[0351] 4-2d. The NEF integrates the intermediate results from another one or more VFL clients NWDAF in step 4-2c, then forward the integrated intermediate results from the another one or more VFL clients NWDAF to the VFL server by invoking one NEF service operation Nnef_AnalyticsExposure_Notify towards the AF.
[0352] 4-2a. - 4-2b. / 4-2a. - 4-2d. will repeat until the VFL inference is finished (e.g. all the VFL clients have reported the intermediate inference results) or the VFL server / NEF determines to terminate the inference or drop the intermediate inference results of the NWDAF VFL clients that haven't been received for the VFL training inference (e.g. if the configure waiting / response time is up, the time window expires, etc.)
[0353] 5. The VFL server performs VFL computation and aggregation for inference by using the intermediate inference results and / or its local data received from VFL clients.
[0354] The VFL server generates the inference results for the required analytics ID.
[0355] Alternatively, the VFL server may perform further VFL computation and aggregation after step 4-2b, once it receives the exposure of intermediate inference results from NEF which is the integrated intermediate inference results of one or more VFL clients NWDAF in step 4-2a.
[0356] The VFL server may perform further VFL computation and aggregation every time once it receives the exposure of intermediate inference results from NEF. Alternatively, the VFL server may perform further VFL computation and aggregation once it receive a number of exposure of intermediate results from NEF. This may be that there is a configured or a set number of exposures that initiate further VFL computation and aggregation. It may also be based on a time window, where a set of exposures may be used for further VFL computation and aggregation based on the time window. In other words, all of the exposures that are received within a time window are used for further VFL computation and aggregation. This may also be considered to be periodic VFL computation and aggregation. The age of an exposure may also be used to determine whether to pursue further VFL computation and aggregation. In other words, the VFL computation and aggregation is performed within a certain time from receiving the exposure from a client. The VFL server could be AF (both trusted and untrusted), or NWDAF.
[0357] Enhanced NEF exposure mechanism to support VFL
[0358] In the procedures for VFL training and VFL inference described above under the headings "Description of VFL model training procedure" and "Description of VFL inference procedure", Alternative 2 in different steps may require the NEF to integrate the information reporting / exposure from the VFL clients before exposing the information to the VFL sever AF, if the AF is untrusted.
[0359] For VFL inference and training, different VFL clients may perform different inference and training (e.g. using the VFL model with different layers, nodes, etc.), use different local data, etc. The VFL clients may also have different computational / process capabilities, available resources, work load, task arrangement / organisation, etc. Therefore, for model training and inference, different VFL clients may finish the training for the same interaction at different times. As a result, the NEF will receive the intermediate training and inference from different NWDAF at different times, as shown in Figure 6 intermediate training / inference results from VFL clients. Then the NEF needs to determine when to forward the integrated / consolidated results to the AF to enhance the VFL process and avoid delay / latency.
[0360] In this disclosure, it could potentially be one or more NEF(s) are involved into the AIML (VFL) process (including model training and inference).
[0361] The NEF may determine to expose the intermediate results of VFL training (for instance by invoking Nnef_MLModelTraining_Notify or Nnef_MLModelTrainingInfo_Request response) or VFL inference (for instance by invoking Nnef_AnalyticsExposure_Notify) based on different triggering conditions, e.g.
[0362] · The NEF exposure / notification / reporting is triggered by VFL server, e.g. the VFL server (e.g. untrusted AF) may send a request to the NEF to request the information exposure for VFL training or inference.
[0363] o E.g. the VFL server sends one or more requests to the NEF for VFL information reporting, e.g. by invoking Nnef_MLModelTraining_Request / Nnef_MLModelTraining_Subscribe request, an indication / flag (e.g. intermediate results reporting) might be included in the service operation / request to clarify the request is to acquire the VFL intermediate results for training / inference.
[0364] o Or a VFL server sends one or more request / service operation that is only used for the VFL server to acquire the information exposure from NEF for VFL training / inference related information.
[0365] · The NEF performs periodic reporting, e.g. based on a time window / timers / duration. For example the NEF sends the intermediate results to AF every N minutes.
[0366] o The time for periodic reporting could be configured by the VFL server, e.g. the AF, by indicating the time information in the service operation / request for VFL training (e.g. in Nnef_MLModelTraining_Subscribe request / service operation or Nnef_MLModelTrainingInfo_Request ) or VFL inference request (e.g. in Nnef_AnalyticsExposure_Subscribe)
[0367] o The time for periodic reporting could be configured by the network operator, or determined based on NEF internal logic.
[0368] o If there is no intermediate results from the VFL clients, then the NEF may not report the information.
[0369] · Based on the number of the reports / notification received by the NEF from the VFL clients (e.g. NWDAF) for VFL model training or VFL inference.
[0370] o It can be defined as a maximum number of reports received from the VFL clients
[0371] o The number of reports / notification received by the NEF to trigger the information exposure could be configured by the VFL server, e.g. the AF, by indicating the time information in the service operation / request for VFL training (e.g. in Nnef_MLModelTraining_Subscribe request / service operation or Nnef_MLModelTrainingInfo_Request ) or VFL inference request (e.g. in Nnef_AnalyticsExposure_Subscribe)
[0372] o The number of reports / notifications received by the NEF to trigger could be configured by the network operator, or determined based on NEF internal logic.
[0373] · Based on maximum wait time / time window / maximum response time
[0374] o This can be used by the NEF to send all of the information it has received from VFL clients, but has not been exposed to VFL server. The NEF exposes / reports the VFL related information (e.g. intermediate results of VFL training or inference) that is stored at the NEF, when the maximum wait time / maximum response time is reached, the time window expires.
[0375] o The starting point / time of the wait time / response time could be when the NEF forwards the VFL training / inference request to the corresponding VFL clients.
[0376] · Minimum wait time
[0377] o The report will be stored at the NEF for a minimum wait time before being sent along with other reports from other VFL clients.
[0378] · By NEF itself, based on e.g. current load of NEF, used / remaining storage of the NEF, energy consumption (EC) of the NEF, if the load / storage / EC exceeds a configured threshold / limit / constraint.
[0379] · (remaining) Response time
[0380] o If there is a maximum response time for the VFL server request, then the VFL client may include a remaining response time. Similarly, the VFL client may also have a pre-configured max response time. This can be used by the NEF to ensure that the NEF buffering does not cause a delay of the reports. In other words, the NEF ensure to deliver the report to the VFL client within the remaining response time.
[0381] · Based on the configured reporting / exposure requirement on group reporting
[0382] o the VFL server may require the NEF to integrate the intermediate results from different VFL clients into different groups for exposure. E.g. the VFL server requires the NEF to integrate the results from VFL client #1,3,5 to report together, and integrate the results from VFL client #2,4,6 to report together.
[0383] For the above triggers, e.g. periodicity of NEF reporting / exposure, number of the reports / notification received by NEF, maximum wait time / time window / maximum response time, Minimum wait time, etc., the values of the triggers could be different for different VFL training iterations. E.g. more time / computation resources might be required for the 1stiteration than the later iteration. Therefore, the VFL sever may configured longer periodicity, less number of the reports / notification, longer maximum wait time / time window / maximum response time and Minimum wait time for the1stiteration than the later iteration. To configure the different values of the triggers in different iteration, the VFL server indicates the values of the triggers in the VFL training request every time the training is triggered.
[0384] Alternatively, the values of the triggers for NEF exposure / reporting might be the same for different iteration of the same VFL process (including all iterations of VFL training). In this case, VFL server only indicates the values of the triggers in the 1strequest of VFL training. For the later iteration (after the 1stiteration), the NEF uses the values of the triggers configured for the VFL process. The NEF may identify whether the request for VFL training from the VFL server is for the same VFL process or not by checking the VFL process ID (e.g. correlation ID) in the server's request. If the ID is the same, it mean the request for the VFL training is for the same VFL process, but just different iteration. In another word, if values of the triggers are not indicted by the VFL server, the NEF just use previous values received in the VFL training request with same VFL process ID / correlation ID, or use the default values.
[0385] NEF service operations to support VFL
[0386] General description on Nnef_MLModelTraining service
[0387] Service Description:This service enables the consumer to subscribe / unsubscribe / notify / modify for AI / ML (e.g. vertical federated learning (VFL)) Model training, model training preparation, sample alignment, feature alignment, etc. toward the VFL clients via NEF.
[0388] The consumer of this service could be AF, e.g. when the VFL server is untrusted AF, the untrusted AF is the consumer of this service.
[0389] This service include the Nnef_MLModelTraining_Subscribe, Nnef_MLModelTraining_Unsubscribe, Nnef_MLModelTraining_Notify, optional Nnef_MLModelTrainingInfo_Request.
[0390] When used for (vertical) Federated Learning, this service enables (V)FL server (e.g. trusted or untrusted AF, or NWDAF) to enable (vertical) Federated Learning while providing global / partial / initial / untrained / per VFL client specific ML Model information to (v)FL Client (e.g. NWDAF or AF) via NEF, and getting local ML Model information and status report of (v)FL training, including the report of intermediate results (for VFL training and / or inference), loss information, gradient information, integrated intermediate results of one or more VFL clients (e.g. NWDAF), indication on whether the VFL clients will join the VFL process / training / inference or not, the cause value indicates the reason if the VFL clients can and / or cannot join the VFL process / training / inference.
[0391] This service may also be used by the consumer (i.e. VFL Server NWDAF / untrusted AF / trusted AF) to check if the service provider (i.e. VFL Client NWDAF / AF) can meet the ML Model training requirement or not via NEF, e.g. whether the VFL Client has the capability / available work load / computation resources and ability / support the required sample (e.g. UE ID, slice ID, NF ID, etc.) or not / support the required features or not / coordination with other VFL clients is supported or not, etc.
[0392] This service may also be used by the consumer (e.g. trusted or untrusted AF, or NWDAF) to request the service provider (i.e. VFL Client NWDAF / AF) to calculate and provide Model Accuracy of the global / partial / initial / VFL client local ML Model.
[0393] This service may also be used to inform the NEF about the requirements and criteria of NEF exposure and integration for VFL intermediate results. E.g. if the AF requires the NEF to integrate the exposure for VFL intermediate results of one or more VFL clients, or if the NEF is able to / allowed to integrate the exposure for VFL intermediate results of one or more VFL clients.
[0394] Nnef_MLModelTraining_Subscribe service operation
[0395] Service operation name:Nnef_MLModelTraining_Subscribe
[0396] Description:Subscribes to NEF for VFL / ML Model training with specific parameters.
[0397] Inputs may include:
[0398] - Analytics ID as defined in Table 7.1-2 of TS 23.288;
[0399] - ML Model Interoperability information;
[0400] - Notification Target Address (+ Notification Correlation ID).
[0401] - VFL / ML Model identifier: identifies the provided ML Model.
[0402] - VFL / ML Model Information, could be global / partial / initial / untrained / per VFL client specific model
[0403] - VFL / ML Model file, could be global / partial / initial / untrained / per VFL client specific model;
[0404] - Subscription Correlation ID (in the case of modification of the ML Model Training subscription), VFL process ID, analytics ID;
[0405] - VFL / ML Training Information, i.e. data availability requirement, time availability requirement.
[0406] - VFL / ML Preparation Flag;
[0407] - VFL / ML Model Accuracy Check Flag;
[0408] - VFL / ML Correlation ID;
[0409] - VFL Training Filter Information;
[0410] - Target of VFLTraining Reporting;
[0411] - VFL Training Reporting Information
[0412] - VFL Use case context;
[0413] - VFL Iteration round ID;
[0414] - VFL Expiry time.
[0415] - List of VFL clients to participate the VFL process / training;
[0416] - optional, Label information and / or the associated VFL clients that will received the label. The VFL server may indicate the label information to some of the (active) VFL clients for VFL training, e.g. to calculate the loss for the (active) VFL client itself and also for the other (passive) VFL clients that do not have the label information. This might be indicated together with VFL / ML Preparation Flag.
[0417] - flag / indication of whether NEF need to integrate the intermediate results or not. E.g. if indicated, the NEF need to integrate the intermediate results from one or more VFL clients; or if indicated, the NEF does not need to integrate the intermediate results from one or more VFL clients;
[0418] - flag / indication of integration capability is required at NEF or not, e.g. whether the NEF can integrate the intermediate results reported by one or more VFL client (NWDAF) or not.
[0419] - remaining List of VFL clients to participate the VFL process / training, in the case if the VFL client share intermediate results between each other, the previous VFL client only inform to the next VFL client of the remaining VFL clients in the list who have not performed model training, e.g. in this iteration round.
[0420] - order / sequence of the VFL clients to participate the VFL process / training, e.g. 1stVFL client perform VFL training, then share the results to the 2ndVFL client, the VFL training will be performed successively until the last VFL client in the list.
[0421] - order / sequence of the remaining VFL clients to participate the VFL process / training, e.g. 1stVFL client perform VFL training, then share the results to the 2ndVFL client, the VFL training will be performed successively until the last VFL client in the list.
[0422] - Indication for VFL training or VFL inference.
[0423] - Indication of skipping the current (v)FL round.
[0424] - Indication of VFL client coordination capability is required or not. Some the VFL process may require the VFL clients to share intermediate results between each other. During the VFL preparation stage the VFL serve may need to check whether the coordination between VFL clients could be supported by the candidate VFL clients to be selected for VFL process or not. This indication might be used together with VFL / ML Preparation Flag.
[0425] - Timer / time window for NEF exposure / reporting of intermediate results. The intermediate results could be the integrated intermediate results which is gathered by the NEF from one or more VFL clients (NWDAF). The timer / time window is used to NEF periodic exposure / reporting to the service consumer (e.g. VFL server AF) as described above under the heading "Enhanced NEF exposure mechanism to support VFL".
[0426] - Indication / flag to require the NFF to exposure of (integrated) intermediate results. Can be optional input. Might be only used when the consumer (e.g. VFL server AF) requires the NFF to exposure of (integrated) intermediate results, but not to trigger (an iteration of) VFL model training.
[0427] - Indication of the number of the reports / notification received by the NEF from the VFL clients for NEF exposure / reporting. The number of the reports / notification received by the NEF to trigger NEF exposure is described above under the heading "Enhanced NEF exposure mechanism to support VFL".
[0428] - Indication of the maximum wait time / time window / maximum response time counted by the NEF for NEF exposure / reporting. The maximum wait time / time window / maximum response time to trigger NEF exposure is described above under the heading "Enhanced NEF exposure mechanism to support VFL".
[0429] - Indication of the (remaining) Response time counted by the NEF for NEF exposure / reporting. (remaining) Response time to trigger NEF exposure is described above under the heading "Enhanced NEF exposure mechanism to support VFL".
[0430] Outputs Required:When the request is accepted: Subscription Correlation ID (required for management of this subscription) or VFL process ID, analytics ID. When the request is not accepted, an error response with cause code (e.g. NWDAF does not meet the ML training requirements, ML training is not complete, NWDAF overload, not available for the FL process anymore, required samples are not supported, required features are not supported, coordination between VFL clients are not supported, intermediate results integration at NEF is not supported, etc.).
[0431] NOTE: The detail reasons in the cause code are up to Stage 3.
[0432] Outputs, Optional:ML / VFL Correlation ID / process ID / analytics ID (e.g. confirm of the subscription for this (V)FL process / training).
[0433] Nnef_MLModelTraining_Notify service operation
[0434] Service operation name:Nnef_MLModelTraining_Notify
[0435] Description:the VFL clients, e.g. NWDAF / AF, notifies the consumer (instance) of the intermediate results of VFL (loss related information, gradient related information, etc.), or trained VFL / ML Model, that has subscribed to the specific NEF / NWDAF service or VFL process, via NEF. The NEF can also use this service to indicate to consumer that the VFL server will terminate the VFL / ML Model training.
[0436] Inputs, Required and / or optional:
[0437] - Notification Correlation Information: this parameter indicates the Notification Correlation ID / VFL process ID / analytics ID that has been assigned by the consumer during VFL training / ML Model training and / or preparation stage.
[0438] - Set of the tuple (Analytics ID, VFL / ML Model Information, label information, loss information, gradient information)
[0439] - VFL / ML Correlation ID, when for vertical Federated Learning, e.g. the training process;
[0440] - Corresponding Use case context;
[0441] - Termination Request: this parameter indicates that NWDAF requests to terminate the VFL / ML Model training, i.e. VFL client (e.g. NWDAF) will not provide further notifications related to this request / VFL (training) process, with cause code (e.g. NWDAF overload, not available for the VFL process anymore, has more urgent task, no enough computation resources, cannot coordinate with other VFL client (NWDAF), VFL client will offline / shut down / switch off, cannot finish the model training within required time, etc.);
[0442] - VFL / ML Model identifier: this parameter identifies the provisioned VFL / ML Model, if any;
[0443] - Global VFL / ML Model Accuracy information: The model accuracy of the global VFL / ML Model, which is calculate by the FL Client NWDAF using the local training data as the testing dataset;
[0444] - Local VFL / ML Model Accuracy information: The model accuracy of the local model / partial global ML Model, which is calculate by the FL Client NWDAF using the local training data as the testing dataset;
[0445] - Status report of VFL training: local ML Model accuracy and monitoring information and Training Input Data Information (e.g. areas covered by the data set, sampling ratio, maximum / minimum of value of each dimension, etc.), which are generated by the FL Client NWDAF during FL procedure,
[0446] - Delay Event Notification.
[0447] - Iteration round ID.
[0448] NOTE: The detail reasons in the cause code are up to stage 3.
[0449] Outputs, Required:Operation execution result indication.
[0450] Outputs, Optional:None.
[0451] Nnef_MLModelTraining_Unsubscribe service operation
[0452] Service operation name:Nnef_MLModelTraining_Unsubscribe
[0453] Description:the VFL server terminates VFL / AIML Model training at VFL clients, via NEF.
[0454] Inputs, Required and or optional:Subscription Correlation ID, VFL process / training ID, analytics ID, list of the information / ID of the VFL clients (if those VFL clients are unselected for the VFL training / process).
[0455] Outputs, Required:Operation execution result indication.
[0456] Outputs, Optional:Cause code (e.g. VFL Client NWDAF is unselected by the VFL Server for the VFL process (e.g. because the VFL client does not upload results within required time / accuracy cannot meet the requirement / is too low, reduce the number of VFL clients for VFL process as some of the clients are not needed any longer to improve the VFL efficiency, etc.), or the VFL process is suspended or finished, etc.). Final aggregated vfl / ML Model / final intermediate results information / final loss information / final gradient information (if VFL has finished) or updated aggregated information / updated intermediate results / updated loss information / updated gradient information (if VFL is suspended).
[0457] Nnef_MLModelTrainingInfo_Request service operation
[0458] Service Description: This service enables the consumer to request for the information about VFL / ML Model training, e.g. based on the VFL / ML Model file or VFL / ML Model information, VFL intermediate results (e.g. loss and / or gradient related information) provided by the consumer.
[0459] When used for vertical Federated Learning, this service enables VFL server (e.g. NWDAF, trusted AF, untrusted AF) to enable Federated Learning towards the corresponding VFL clients (indicated in the service request) while providing local model / partial global ML Model information to VFL Client (e.g. NWDAF) and getting local VFL / ML Model information, local computation result of VFL intermediate results (e.g. loss and / or gradient related information) from the FL Client (NWDAF).
[0460] Description:Request information about NWDAF ML Model training with specific parameters.
[0461] Inputs, Required or optional
[0462] - Termination Request, when terminating the vertical Federated Learning identified by the VFL / ML Correlation ID / VFL process ID and optionally indicating the reason, e.g. VFL Client NWDAF is unselected by the VFL Server NWDAF for the VFL process, or the VFL process is suspended, etc.
[0463] - Analytics ID as defined in Table 7.1-2 of TS 23.288;
[0464] - ML Model Interoperability information;
[0465] - Notification Target Address (+ Notification Correlation ID).
[0466] - VFL / ML Model identifier: identifies the provided ML Model.
[0467] - VFL / ML Model Information, could be global / partial / initial / untrained / per VFL client specific model
[0468] - VFL / ML Model file, could be global / partial / initial / untrained / per VFL client specific model;
[0469] - Subscription Correlation ID (in the case of modification of the ML Model Training subscription), VFL process ID, analytics ID;
[0470] - VFL / ML Training Information, i.e. data availability requirement, time availability requirement.
[0471] - VFL / ML Preparation Flag;
[0472] - VFL / ML Model Accuracy Check Flag;
[0473] - VFL / ML Correlation ID;
[0474] - VFL Training Filter Information;
[0475] - Target of VFLTraining Reporting;
[0476] - VFL Training Reporting Information
[0477] - VFL Use case context;
[0478] - VFL Iteration round ID;
[0479] - VFL Expiry time.
[0480] - List of VFL clients to participate the VFL process / training;
[0481] - optional, Label information and / or the associated VFL clients that will received the label. The VFL server may indicate the label information to some of the (active) VFL clients for VFL training, e.g. to calculate the loss for the (active) VFL client itself and also for the other (passive) VFL clients that do not have the label information. This might be indicated together with VFL / ML Preparation Flag.
[0482] - flag / indication of integration capability is required at NEF or not, e.g. whether the NEF can integrate the intermediate results reported by one or more VFL client (NWDAF) or not.
[0483] - remaining List of VFL clients to participate the VFL process / training, in the case if the VFL client share intermediate results between each other, the previous VFL client only inform to the next VFL client of the remaining VFL clients in the list who have not performed model training, e.g. in this iteration round.
[0484] - order / sequence of the VFL clients to participate the VFL process / training, e.g. 1stVFL client perform VFL training, then share the results to the 2ndVFL client, the VFL training will be performed successively until the last VFL client in the list.
[0485] - order / sequence of the remaining VFL clients to participate the VFL process / training, e.g. 1stVFL client perform VFL training, then share the results to the 2ndVFL client, the VFL training will be performed successively until the last VFL client in the list.
[0486] - Indication for VFL training or VFL inference.
[0487] - Indication of skipping the current (v)FL round.
[0488] - Indication of VFL client coordination capability is required or not. Some the VFL process may require the VFL clients to share intermediate results between each other. During the VFL preparation stage the VFL serve may need to check whether the coordination between VFL clients could be supported by the candidate VFL clients to be selected for VFL process or not. This indication might be used together with VFL / ML Preparation Flag.
[0489] - Timer / time window for NEF exposure / reporting of intermediate results. The intermediate results could be the integrated intermediate results which is gathered by the NEF from one or more VFL clients (NWDAF). The timer / time window is used to NEF periodic exposure / reporting to the service consumer (e.g. VFL server AF) as described above under the heading "Enhanced NEF exposure mechanism to support VFL".
[0490] - Indication / flag to require the NFF to exposure of (integrated) intermediate results. Can be optional input. Might be only used when the consumer (e.g. VFL server AF) requires the NFF to exposure of (integrated) intermediate results, but not to trigger (an iteration of) VFL model training.
[0491] - Indication of the number of the reports / notification received by the NEF from the VFL clients for NEF exposure / reporting. The number of the reports / notification received by the NEF to trigger NEF exposure is as described above under the heading "Enhanced NEF exposure mechanism to support VFL".
[0492] - Indication of the maximum wait time / time window / maximum response time counted by the NEF for NEF exposure / reporting. The maximum wait time / time window / maximum response time to trigger NEF exposure is as described above under the heading "Enhanced NEF exposure mechanism to support VFL".
[0493] - Indication of the (remaining) Response time counted by the NEF for NEF exposure / reporting. (remaining) Response time to trigger NEF exposure is as described above under the heading "Enhanced NEF exposure mechanism to support VFL".
[0494] Outputs Required and positional:
[0495] - When the request is accepted: Operation execution result indication. When the request is not accepted, an error response with cause code (e.g. NWDAF does not meet the ML training requirements, ML training is not complete, NWDAF overload, not available for the FL process anymore, required samples are not supported, required features are not supported, coordination between VFL clients are not supported, intermediate results integration at NEF is not supported, etc.).
[0496] - ML Model identifier.
[0497] - VFL label identifier, intermediate results identifier
[0498] - Set of the tuple (Analytics ID, VFL / ML Model Information).
[0499] - VFL / ML Correlation ID, when for vertical Federated Learning.
[0500] - Corresponding Use case context.
[0501] - Global ML Model Accuracy information: The model accuracy of the global ML Model, which is calculate by the FL Client NWDAF using the local training data as the testing dataset.
[0502] - Local VFL / ML Model Accuracy information: The model accuracy of the local model / partial global ML Model, which is calculate by the FL Client NWDAF using the local training data as the testing dataset;
[0503] - Status report of VFL training: local ML Model accuracy and monitoring information and Training Input Data Information (e.g. areas covered by the data set, sampling ratio, maximum / minimum of value of each dimension, etc.), which are generated by the FL Client NWDAF during FL procedure,
[0504] - Delay Event Notification.
[0505] - Iteration round ID.
[0506] Further aspects, examples and embodiments of the present disclosure are now described.
[0507] Vertical Federated Learning (VFL)
[0508] Unlike traditional centralized learning approaches, where data is pooled together in a single location, or Horizontal Federated Learning (HFL), where different entities contribute similar types of data about different samples, VFL allows for the collaborative training of machine learning models across entities that hold different types of information about the same entities or events. For example, collaboration is not only limited to a single type of entity e.g. only NWDAFs but instead may make use of dissimilar entities that performed different functions and which may be at different levels / layer of the network, for example, a UE, a NWDAF and a an AMF. This approach is particularly relevant and beneficial for telecommunication networks for several reasons.
[0509] VFL involves multiple parties collaboratively learning a model while keeping their data localized, only exchanging model intermediate results or gradients / learning parameters, not the raw data itself. Each participant in the VFL process contributes a different set of features about the same sample set, enhancing the model's learning capacity without compromising data privacy. This method is especially useful in scenarios where data cannot be shared freely due to privacy concerns, regulatory restrictions, or commercial competitiveness.
[0510] The process of VFL involves several steps such as those below, but is not only limited to these:
[0511] Data Alignment:To start the training, all parties should be referencing the same set of samples. This alignment is typically achieved through Private Set Intersection (PSI), a technique that identifies common samples across datasets without disclosing any actual data.
[0512] Feature Encoding:Parties locally encode their features to maintain data privacy throughout the training process. Techniques such as homomorphic encryption are employed to ensure that data remains secure, even when being processed.
[0513] Local Model Initialization:Each participant initializes a local model tailored to its unique set of features. These models can vary across parties, reflecting the diverse data each holds.
[0514] Collaborative Training:
[0515] 1.Insight Synthesis Sharing:In this phase, each party processes its data to generate insights or intermediate representations based on its local model and encrypted data. These insights are then shared with the next party in the VFL chain, serving as inputs for further processing. This iterative process continues until the last party integrates all insights to compute an overall outcome or loss, reflecting the collective learning from all parties.
[0516] 2.Model Refinement:Following the synthesis of insights, the process of model refinement begins. Here, adjustments or gradients are calculated to improve the model based on the collective outcome. These adjustments are securely passed back through the chain, allowing each party to refine its local model parameters. This secure aggregation of adjustments ensures that the model evolves to more accurately represent the data without compromising the privacy of any party's information.
[0517] The usefulness of VFL for communications networks comes from:
[0518] Enhanced Privacy and Security:Telecommunication networks handle vast amounts of sensitive user data that are subject to strict privacy regulations and standards. VFL enables the utilization of this data for network optimization and service improvement without exposing individual user data, thus maintaining privacy and security.
[0519] Richer Insights from Diverse Data Sources:Telecommunication networks are inherently complex, with data collected from a wide range of sources, including user devices, network equipment, and service platforms. VFL allows for the integration of diverse data types across these sources, leading to richer insights and more accurate network analytics to, for example, predict network demand, detect anomalies, or enhancing user experience, among others.
[0520] Operational Efficiency:By enabling collaborative model training across different network entities and domains without centralizing data, VFL can significantly reduce the bandwidth and storage requirements typically associated with large-scale data analytics. This efficiency is especially useful for real-time or near-real-time applications such as dynamic network slicing, congestion management, and service quality optimization.
[0521] Cross-Vendor Collaboration:The telecommunication industry often involves multiple vendors and operators working within the same ecosystem. VFL facilitates collaboration across these entities, allowing them to jointly develop and refine models that improve network performance and service offerings without sharing sensitive or proprietary data.
[0522] Customization and Personalization:By leveraging data from various sources, VFL enables the development of customized services and personalized user experiences. This can lead to improved customer satisfaction and new revenue opportunities for operators.
[0523] 3GPP SA2 Rel-19 Study in AIML (FS_AIML_CN)
[0524] A new Study Item (SID) on Core Network Enhanced Support for Artificial Intelligence (AI) / Machine Learning (ML) was approved in SP-231800 [1] in TSG SA Meeting #102 (Dec 2023). In WT#2 in the SID:
[0525] in SP-231800 [1]:-WT2: Study whether and what potential enhancements are needed to enable 5G system to assist in collaborative AI / ML operation involving 5GC / NWDAF and / or AF for "Vertical Federated Learning (VFL)". The work will be based only on and limited to the scope of justified use cases.NOTE 7: RAN and UE aspects are out of scope. Solutions based on interactions between the application client and 5GS are out of scope. The necessary communication between AF and UE application client to support the collaborative AI / ML operation is understood as no normative procedure impact. Horizontal FL procedure defined in R18 should be taken into account and reused whenever possible.NOTE 8: coordination with SA6 is required.
[0526] The detailed description of KI#2: 5GC Support for Vertical Federated Learning was documented in clause 5.2.2 of TR 23.700-84 [3]. The issues to be addressed for KI#2 by SA2 during Rel-19 study phase include:
[0527] This key issue aims to provide solutions for enabling 5GC support for vertical federated learning (VFL) involving NWDAF and / or AF, where no raw data need to be exchanged but some level of coordination is still required when training and inference are performed on local models. In particular, datasets used for each local model need to share the same samples while holding different features.In Rel-18, ML model sharing between NWDAFs has been studied as a part of Horizontal Federated Learning. However, Federated learning between NWDAF and AF has not been studied (e.g. when the NWDAFs and / or AFs are in different domains, locations, regions etc).Vertical Federated Learning (VFL) can be considered as an alternative mechanism for distributed functionalities of an ML model. Note that, as scoped in Rel-19, NWDAF and / or AF may be involved for VFL.This Key Issue aims to study architecture enhancement to support VFL, which allows the cooperative AI / ML training and inference with the following aspects:-Identify VFL use cases and under which conditions, and for which entities these VFL use cases show that VFL is justified to train ML models.-Whether and how to support architecture enhancement for supporting VFL for model training and / or inference. In particular:-Whether and how the existing NF discovery and selection needs to be enhanced.-Whether and how ML Model training and / or inference related procedures need to be enhanced to support VFL-Whether and how to do performance monitoring for the ML model trained via VFL-Whether and how to provide ML Models to the participants in the VFL training process.-How to support sample and feature alignment among the participating network entities when performing VFLNOTE 1: Application layer-based VFL requiring communication between AFs and / or UEs application client, is out of scope.NOTE 2: During the study on this KI, consultation with SA3 is required for handling security aspects.NOTE 3: RAN and UE aspects are out of scope.NOTE 4: The existing procedures defined for Horizontal FL in 3GPP TS 23.288 [x] will be taken into account when studying the procedure for VFL.
[0528] In order to clarify the scenarios of using Vertical Federated Learning, VFL use cases are to be identified. One possible use case of implementing VFL is to deploy NWDAF to Support for Sample and Feature Alignment in VFL:
[0529] In the AI / ML literature, VFL is a federated learning setting where multiple parties perform training on data sets that share the same sample space but differ in feature space. Because of this, an alignment in sample and feature spaces among participating entities is usually required before applying VFL. VFL further allows to perform joint training without exposing raw data or model parameters, the latter being a way in which VFL differs from HFL. TS 23.288 [2] provides NWDAF specification support for HFL but no VFL support is available.
[0530] This use case proposes NWDAF support for VFL in analytics derivation by means of sample and feature alignment between the entities participating in VFL, where the main entity facilitating the VFL operation is NWDAF and other entities may be other NWDAF instances and / or AF(s). The motivation for this use case is mainly two-fold: i) in a multi-vendor scenario, VFL may be more suitable than HFL for multiple NWDAF deployments since such accuracy increase may be achieved without the need to share model parameters among the participating NWDAF from different vendors, and ii) VFL allows an enhanced accuracy of the NWDAF predictions as models trained via VFL usually generalize better by learning from a broader feature set.
[0531] In PLMNs where multiple NWDAFs are deployed, each NWDAF instance may perform data collection locally according to their suitable data sources. Depending on the Analytics ID, the different NWDAF instances may share the sample ID space (e.g. S-NSSAI) or train on different sample ID spaces (e.g. UE IDs within their corresponding Area of Interest). Furthermore, the NWDAF instances are not all obliged to collect the same input data for the same Analytics ID as most input data is optional, thus their feature spaces may range from full to little overlap. Finally, while an AF may also participate on VFL supported by NWDAF, an alignment of samples would still be needed between the two entities, and feature alignment may also prove beneficial.
[0532] Depending on the range of overlap in sample and feature spaces of the participating entities, VFL may be a more or less suitable technique to combine models at NWDAF. Hence, support for sample and feature alignment would allow the VFL supported by NWDAF to be more effective for those scenarios that are suitable.
[0533] Based on the above background information and analysis of the given use case, compared to HFL, VFL is able to increase the accuracy of FL without sharing model parameters among the participating NWDAF from different vendors and also allows an enhanced accuracy of the NWDAF predictions as models.
[0534] As it has been agreed by SA2, the following issues should be addressed during SA2 Rel-19 study to support Vertical Federated Learning as documented in TR 23.700-84 [3]:
[0535] This Key Issue aims to study architecture enhancement to support VFL, which allows the cooperative AI / ML training and inference with the following aspects:-Identify VFL use cases and under which conditions, and for which entities these VFL use cases show that VFL is justified to train ML models.-Whether and how to support architecture enhancement for supporting VFL for model training and / or inference. In particular:-Whether and how the existing NF discovery and selection needs to be enhanced.-Whether and how ML Model training and / or inference related procedures need to be enhanced to support VFL-Whether and how to do performance monitoring for the ML model trained via VFL-Whether and how to provide ML Models to the participants in the VFL training process.-How to support sample and feature alignment among the participating network entities when performing VFL
[0536] VFL is not able to increase the accuracy of FL without sharing model parameters among the participating NWDAF from different vendors, but also allows an enhanced accuracy of the NWDAF predictions as models. However, in the current 3GPP specifications, it is lack of solutions to address the above issues in the KI description on KI#2: 5GC Support for Vertical Federated Learning.
[0537] In order to more fully support the Vertical Federated Learning by 5GC in different scenarios, new solutions are required to solve the above issues during SA2 Rel-19 study and normative phase.
[0538] Vertical Federated Learning Enhancements
[0539] Overview
[0540] Below is an overview of the steps / stages involved in the proposed implementation / approaches for enhancing vertical federated learning. This list of steps / stages is not exhaustive and only some of the steps / stages below may be used, and the steps / stages may be used in any combination.
[0541] 1.NWDAF registration:NWDAF containing Model Training Logical Function (MTLF) as VFL Server NWDAF and / or FL Client NWDAF registers to Network Repository Function (NRF) with its Network Function (NF) profile, which includes NWDAF NF, Analytics ID(s), Address information of NWDAF, Service Area, FL capability type information (i.e. FL server or FL client) and Time interval supporting FL as described in clause 5.2.
[0542] 2.Model Request:The model consumer (e.g. NWDAF containing Analytics Logical Function (AnLF) or NWDAF containing MTLF) sends a subscription / request to the VFL server NWDAF to retrieve an ML model, invoking the Nnwdaf_MLModelProvision_Subscribe / Nnwdaf_MLModelInfo_Request service (if consumer is NWDAF containing AnLF) or by invoking the Nnwdaf_MLModelTraining_Subscribe service / Nnwdaf_MLModelTrainingInfo_Request (if consumer is NWDAF containing MTLF) operation. This may include specifying the Analytics ID, desired ML model metrics (e.g., accuracy), accuracy reporting interval, and pre-determined status (accuracy threshold or time when the model is needed).P
[0543] 3.Participant Selection
[0544] 3.1 Once the VFL Server NWDAF is determined, the FL Server NWDAF discovers and selects other NWDAF(s) containing MTLF as VFL Client NWDAF(s) from NRF by invoking the Nnrf_NFDiscovery_Request service operation. NWDAF(s) containing MTLFs may also or alternatively be selected based on a predetermined list or other predetermined configuration as opposed to using a discovery process. For example, if a previous discovery process has been performed, the results may be reused.
[0545] 3.2 VFL Server NWDAF sends Federated Learning preparation request to the VFL Client NWDAF(s), using Nnwdaf_MLModelTraining_Subscribe or Nnwdaf_MLModelTrainingInfo_Request service with the ML Preparation Flag, to check if the VFL Client NWDAF(s) can meet the ML model training requirement (e.g. Analytics ID, ML Model Interoperability information, Available data requirement, Availability time requirement (time span needed for the FL process), etc.). In some examples, such a step may have been preformed at a previous time and / or the capabilities of VFL Client NWDAFs may have been pre-registered and thus known to the VFL Server NWDAF. In other examples, Federated Learning preparation requests may only be sent to VFL Client NWDAFs whose capabilities are known to be adequate given the model requirements.
[0546] 3.3 Data Alignment: Each VFL Client NWDAF aligns its dataset with the common samples identified across all tentative participating NWDAFs using the Nnwdaf_MLModelTraining_DataAlignment service.
[0547] 3.4 Each VFL Client NWDAF undertakes a detailed evaluation to ascertain its capability to meet the specified ML model training requirements. This multifaceted assessment includes one or more of 3.4.1 to 3.4.3 below. This capability evaluation may be performed in response to participant selection or may be performed at an earlier time and the result of the evaluation stored and / distributed to other entities such as VFL Server NWDAFs.
[0548] 3.4.1 Model Training Feasibility: Evaluating the technical feasibility to meet the ML model's training demands, considering the available computational resources and data processing capabilities.
[0549] 3.4.2 Model Access and Compatibility: Determining the ability to access or download the ML model, particularly when specific model details or parameters are shared within the request. This step helps to ensure that the FL Client NWDAF can effectively integrate and utilize the model within its existing infrastructure.
[0550] 3.4.3 Data Alignment Verification: Confirming the existence of common sample IDs with other NWDAFs, as established during the data alignment phase. This verification ensures that the FL Client NWDAF holds relevant and aligned data, making it a valuable participant in the federated learning endeavour.
[0551] 3.5 Based on this thorough evaluation, the VFL Client NWDAF makes an informed decision regarding its participation in the Federated Learning process. VFL Client NWDAF(s) invokes Nnwdaf_MLModelTraining_Notify or Nnwdaf_MLModelTraining_Subscribe response service operation or Nnwdaf_MLModelTrainingInfo_Request response service operation to indicate to the FL Server NWDAF whether it will join the FL procedure, the successful sample IDs identified (if any) and may include the reason in the response message if it cannot join the FL process. In some examples, capability information of the VFL Client NWDAF may already be known and thus such information can merely be compared to the criteria for determining whether to participate. In other examples, this assessment may be performed at the VFL Server NWDAF or other entity if the details of the VFL Client NWDAF are known to the VFL Server NWDAF or other entity.
[0552] 3.6 FL Server NWDAF determines the final list of VFL Client NWDAF(s) to be involved in the FL procedures based on the information received or otherwise determined in step 3.5
[0553] 4.Data Processing and Information Sharing:Each VFL Client NWDAF processes its dataset to generate predictive insights or intermediate data representations. These insights are then securely transmitted to the next designated VFL Client NWDAF(s), as orchestrated by the VFL Server NWDAF, using the Nnwdaf_MLModelTraining_DataProcessing_Sharing operation. This step ensures that the insights from one entity are integrated with or inform the processing by other entities, facilitating a collaborative approach to model refinement across the network. In the event of processing delays, the FL Client NWDAF may proactively notify the FL Server NWDAF to manage expectations and adjust timelines accordingly. In some examples, insights may be securely transmitted to one or more other designated VFL Client NWDAFs or one or more VFL Client NWDAFs may be skipped under control of the FL Server NWDAF.
[0554] 5.Insight Synthesis and Model Refinement:Beginning with the VFL Server NWDAF, each participating client synthesizes the received insights and applies them to refine their local model, using the Nnwdaf_MLModelTraining_InsightSynthesis_Refinement operation. This involves analyzing the shared data representations or insights to improve the model's accuracy or performance, adjusting model parameters as necessary. Each NWDAF contributes to this iterative refinement process by passing enhanced insights or adjustments back to preceding participants, fostering a cumulative improvement in the model's efficacy. This collaborative refinement process ensures that the collective intelligence of all participants is leveraged, leading to a more robust and accurate federated model.
[0555] 6.Model Evaluation and Feedback:If the VFL procedure continues, VFL Server NWDAF sends Nnwdaf_MLModelTraining_Subscribe or Nnwdaf_MLModelTrainingInfo_Request notifying for a next round of Federated Training.
[0556] 7.End of Procedure:When the Federated Training procedure is complete, the VFL Server NWDAF requests the VFL client NWDAF(s) to terminate the FL procedure by invoking Nnwdaf_MLModelTraining_Unsubscribe service with a cause code that the FL process has finished. Then the VFL client NWDAF(s) terminates.
[0557] 8. After the training process is complete, the VFL Server NWDAF may send Nnwdaf_MLModelProvision_Notify that includes the combination of NWDAF clients involved in the process and needed for collaborative intelligence to the consumer.
[0558] Procedure
[0559] A call flow diagram for the proposed procedure is shown in Figure 7. The steps of the procedure are not limited to those illustrated and described below and may include fewer or further steps. A subset including any combination of the steps described below may also be used.
[0560] The procedure below introduces support for VFL operation at NWDAF by means of enabling data alignment (i.e. sample and feature) among entities participating in the VFL training process with a new service. This new service allows guaranteeing that training samples are the same in the participating entities even though their training features are different. The solution requires an VFL server guiding the training process and VFL consumers following the VFL server instructions. NWDAF is the 5GC NF acting as VFL server.
[0561] The procedure in Figure 7 to support VFL operation at NWDAF is described step by step below.
[0562] 1a. VFL server (i.e. NWDAF) and VFL consumer (e.g. NWDAF) entities register to NRF. The registration may include one or more of their NF profiles, Analytics ID(s), Address information of NWDAF, Service Area, VFL capability type information (i.e. VFL server or VFL consumer) and Time interval supporting VFL. The latter parameter can be the same as Time interval supporting FL described in clause 5.2 of TS 23.288 [2].
[0563] 1b. The NWDAF instance with the capability of acting as VFL server is discovered upon a request by the consumer NF (i.e. NF that subscribes to / requires the analytics) or by an NWDAF containing AnLF.
[0564] 2a. [CONDITIONAL] If the VFL server does not include inference capabilities (i.e. NWDAF does not contain AnLF), the consumer NF sends an analytics subscription or request to an NWDAF containing AnLF supporting the requested Analytics ID.
[0565] 2b. [CONDITIONAL] If step 2a is executed, the NWDAF containing AnLF sends a subscription or request to the VFL server NWDAF to retrieve an AI / ML model, invoking the Nnwdaf_MLModelProvision_Subscribe or Nnwdaf_MLModelInfo_Request service. This includes specifying the Analytics ID, desired ML model metrics (e.g., accuracy), accuracy reporting interval, and pre-determined status (accuracy threshold or time when the model is needed).
[0566] 3. [CONDITIONAL] If the VFL server NWDAF contains AnLF, the consumer NF may directly subscribe or request to analytics from the VFL server NWDAF.
[0567] 4. The VFL server NWDAF discovers and selects other VFL consumers (e.g. NWDAF containing MTLF) from NRF by invoking the Nnrf_NFDiscovery_Request service operation. In some examples, VFL consumers may be considered to subscribe to model training services from a VFL server NWDAF.
[0568] 5. The VFL server NWDAF sends a VFL preparation request to the VFL consumers, using Nnwdaf_MLModelTrainingInfo_Request (or Nnwdaf_MLModelTraining_Subscribe) service operation with the ML Preparation Flag, to check if the VFL consumer(s) can meet the ML model training requirement (e.g. Analytics ID, ML Model Interoperability information, Available data requirement, Availability time requirement, etc.). The VFL consumer(s) may respond to the VFL server NWDAF indicating whether they will join the VFL operation and may include the reason in the response message if it cannot join the VFL operation.
[0569] NOTE: The selection of VFL consumer(s) by the VFL server NWDAF may happen after step 5, fully or partially. The selection of VFL consumer(s) by the VFL server NWDAF may also be performed or refined after step 8.
[0570] 6. The VFL server NWDAF sends a request for data alignment (i.e. samples and features) to the VFL consumer(s). This new service / service operation is required to facilitate the alignment of datasets from different sources in a VFL environment. It ensures that participating entities work with a common set of samples without revealing sensitive data. In addition, these samples are the intersection of different datasets where each entity has different features (information or attributes) for the same set of entities or individuals.
[0571] Required inputs of this service operation include dataset identifiers (i.e. unique identifiers for the datasets held by each VFL participating entity), alignment technique (i.e. specific methods or algorithms to be used for data alignment, such as Private Set Intersection (PSI), feature hashing, or other techniques that ensure data privacy and integrity) and Notification Target Address. Optional inputs may include an alignment correlation ID, expiry time, and additional data alignment information such as challenges or discrepancies encountered in the data alignment process.
[0572] 7. Each VFL consumer performs data alignment pre-processing. This is required for each VFL to ascertain its capability to meet the AI / ML model training requirements. Data alignment pre-processing may include model training demands feasibility assessment, model access and compatibility, data alignment verification, etc.
[0573] 8. The VFL consumer(s) make(s) a decision regarding its participation in VFL and notifies the VFL server NWDAF whether it will join the VFL operation along the successful sample IDs identified (if any). The VFL consumer(s) may also include the reason in the response message if it cannot join the VFL operation. The VFL server NWDAF may then select / finalise the VFL consumers to involve in the VFL.
[0574] 9. In the first training iteration, the VFL server NWDAF triggers the VFL training process by invoking the service operation Nnwdaf_MLModelTraining_Subscribe. Subsequent iterations of the VFL training process where intermediate training results are shared are also coordinated by the VFL server NWDAF via the same service operation, facilitating a collaborative approach to model refinement across the participating entities.
[0575] 10. VFL consumers may interact with each other during the VFL training process.
[0576] 11. VFL consumers notify of training results to the VFL server NWDAF via Nnwdaf_MLModelTraining_Notify service operation.
[0577] 12. Once training is complete, participating entities in the VFL process collaborate with each other to perform joint inference.
[0578] 13. [CONDITIONAL] If steps 2a and 2b were executed, NWDAF containing AnLF provides analytics to the NF consumer.
[0579] 14. [CONDITIONAL] If step 3 was executed, VFL server NWDAF provides analytics to the NF consumer.
[0580] Regarding impacts on services, entities and interfaces, NWDAF is required to support the new data alignment service and the existing service Nnwdaf_MLModelTraining needs enhancements to support sharing of intermediate training results for VFL.
[0581] In another embodiment, the following procedure, with reference to Figure 8, can be followed to support VFL with NWDAF and AF as Participants.
[0582] NOTE 1: In this solution, the VFL server coordinates the VFL operation and acts as active participant with access to labels.
[0583] NOTE 2: Participants in this solution can be either active participants with access to labels or passive participants without access to labels.
[0584] NOTE 3: VFL Participants may also be called VFL Clients.
[0585] The procedure in Figure 8 to support VFL operation at NWDAF is described step by step below.
[0586] 1. VFL server (i.e. NWDAF) and VFL participant entities (NWDAF, AF) entities register to NRF. The registration may include their NF profiles, Analytics ID(s), Address information of NWDAF, Service Area, VFL capability type information (i.e. VFL server or VFL participant type) and Time interval supporting VFL. The latter parameter can be the same as Time interval supporting FL described in clause 5.2 of TS 23.288 [5].
[0587] NOTE 3: VFL participant type parameter can have the value of 'active' or 'passive'.
[0588] 2. The VFL server and participants are discovered via NRF by invoking the Nnrf_NFDiscovery_Request service operation.
[0589] NOTE 4: Details of the discovery mechanism are not within the scope of this solution.
[0590] 3. The VFL server sends a VFL preparation request to the VFL consumers. For NWDAF participants, the existing service operations Nnwdaf_MLModelTrainingInfo_Request or Nnwdaf_MLModelTraining_Subscribe may be reused and enhanced, or a new service may be defined. For AF participants, a new AF / NEF service is required. Either way, the ML Preparation Flag is provided to check if the VFL participants can meet the ML model training requirement (e.g. Analytics ID, ML Model Interoperability information, Available data requirement, Availability time requirement, etc.). The VFL participants may respond to the VFL server indicating whether they will join the VFL operation and may include the reason in the response message if it cannot join the VFL operation.
[0591] 4. The selection of VFL participants by the VFL server may happen here, fully or partially. The selection of VFL participants by the VFL server may also be performed or refined in step 8.
[0592] 5. The VFL server sends a request for data alignment (i.e. samples and features) to the VFL participants, via NEF for AF participants. This new service / service operation is required to facilitate the alignment of datasets from different sources in a VFL environment. It ensures that participating entities work with a common set of samples without revealing sensitive data. In addition, these samples are the intersection of different datasets where each entity has different features (information or attributes) for the same set of entities or individuals.
[0593] Required inputs of this service operation include dataset identifiers (i.e. unique identifiers for the datasets held by each VFL participant), alignment technique (i.e. specific methods or algorithms to be used for data alignment, such as Private Set Intersection (PSI), feature hashing, or other techniques that ensure data privacy and integrity) and Notification Target Address. Optional inputs may include an alignment correlation ID, expiry time, and additional data alignment information such as challenges or discrepancies encountered in the data alignment process.
[0594] 6. Each VFL participant performs data alignment pre-processing. This is required for each VFL participant to ascertain its capability to meet the AI / ML model training requirements. Data alignment pre-processing may include model training demands feasibility assessment, model access and compatibility, data alignment verification, etc.
[0595] 7. The VFL participants notify the VFL server the result of the data alignment, via NEF for AF participants. Furthermore, the VFL participants may provide the VFL server with a decision regarding its participation in the VFL operation along the successful sample IDs identified (if any). The VFL participants may also include the reason in the response message if it cannot join the VFL operation.
[0596] 8. If not completed at step 4, the selection of VFL participants by the VFL server may be performed or refined based on the inputs received by the VFL server from the VFL participants.
[0597] 9. The VFL training process starts on this step, where intermediate training results are shared and coordinated by the VFL server, facilitating a collaborative approach to model refinement across the VFL participating entities. In the first iteration, the VFL server triggers the VFL training process by invoking a ML Model Training service subscription operation from the VFL Participant #1 (e.g. NWDAF). In the subsequent iteration, the VFL server may need to use the ML Model Training service notification operation with the Participant #2 (or last participant if more than two) (e.g. AF via NEF) to enable model refinements in each participant. Subsequent iterations may require the VFL server to subscribe or notify from / to either participant using the ML Model Training service. While the existing Nnwdaf_MLModelTraining service can be used for NWDAF participants, a new service at the NEF / AF is required to support this functionality. In the last iteration, the VFL server informs the VFL participants that the VFL training process is completed via a suitable flag, steps 10 through 12 are skipped, and the VFL training loop is terminated.
[0598] 10a.The VFL Participant #1 may perform computation to locally train its model. The computation may lead either to intermediate training results to be shared to the next participant in step 11, or to refine a previously trained model and notify the VFL server in step 13.
[0599] 10b.The VFL Participant #1 may share its intermediate training results via NEF with the AF participant, using the same service operations as in step 9.
[0600] 10c.Participant #2 (i.e. last VFL participant) performs computation to locally train its model. The computation may lead either to intermediate training results to be shared with the VFL server in step 11, or to refine a previously trained model and the previous VFL participant in step 10d.
[0601] 10d.The VFL Participant #1 performs computation to refine its local model.
[0602] 11. Intermediate training results and / or model refinement results are shared with the VFL server, via NEF if from AF, using the same ML model training service used in step 9.
[0603] 12. The VFL server performs further VFL computation.
[0604] 13. A consumer NF subscribes to or requests analytics from the NWDAF hosting VFL server functionality.
[0605] 14. A distributed VFL inference process is triggered by the VFL server invoking a ML model inference request from Participant #1. This step can be executed with a new service or by enhancing the existing ML Model Training service since the required VFL computation is essentially the same for inference and some training cycles.
[0606] 15a.Participant #1 performs local inference computation.
[0607] 15b.Participant #1 shares intermediate inference results with Participant #2 via NEF when Participant #2 is an AF. In that case, the NEF / AF service used may be new or an enhanced version of the new NEF / AF service already used in step 9.
[0608] 15c.Participant #2 performs local inference computation.
[0609] 16. Participant #2 (i.e. the last participant) notifies the result of the distributed inference process to the VFL server using the same service as in step 14 via the notification operation.
[0610] 17. The VFL server performs further inference computation and derives the requested analytics.
[0611] 18. The derived analytics are delivered to the NF consumer.
[0612] Services -Nnwdaf_MLModelTraining Service
[0613] Explained below are service operations that may be used to request or implement any of the steps described above with respect to the overview and the procedure.
[0614] The service operations described below can also be standalone services or service operations part of a different service (e.g.Nnwdaf _VFL service,Nnwdaf _DataAlignment service, etc.). Furthermore, although information is described as required and / or optional, embodiments are not limited to including all the required information. For example, with respect to the Nnwdaf_MLModelTraining_DataAlignment Service operation (and the other service operations), the message may be mandated to include only some of theInputs / Outputs, Required.
[0615] Nnwdaf_MLModelTraining_DataAlignment Service Operation
[0616] Service Operation Name:Nnwdaf_CollaborativeTraining_DataAlignment
[0617] Description:This operation facilitates the alignment of datasets from different sources in a vertical FL environment. It ensures that participating entities work with a common set of samples without revealing sensitive data. In the context of VFL, "common samples" refer to the data points or records that are shared across the datasets of different participating entities, based on a common identifier. These samples are the intersection of different datasets where each entity has different features (information or attributes) for the same set of entities or individuals. The service employs techniques such as, but not limited to, Private Set Intersection (PSI), Private Set Union (PSU) or Federated Learning withOut Revealing InterSecTions (FLORIST), etc. to securely identify overlapping data points across datasets.
[0618] When a request for data alignment is accepted by the NWDAF containing MTLF, the consumer NF (e.g., another NWDAF instance) receives a unique identifier (Alignment Correlation ID) for managing this alignment process. The NWDAF may modify the data alignment parameters based on operator policy and configuration.
[0619] Inputs, Required:
[0620] · Alignment Request: A request from the VFL Server NWDAF to the VFL Client NWDAFs, instructing them to align their datasets. This request includes guidelines or parameters for alignment, such as identifying common entities or samples.
[0621] · Dataset Identifiers: Unique identifiers for the datasets held by each VFL Client NWDAF, to facilitate the alignment process.
[0622] - Alignment Technique Specification: specific methods or algorithms to be used for data alignment, such as Private Set Intersection (PSI), feature hashing, or other techniques that ensure data privacy and integrity.
[0623] - Notification Target Address: Address for sending notifications related to the data alignment process.
[0624] Inputs, Optional:
[0625] · Alignment Correlation ID: Used for modifying an existing data alignment subscription.
[0626] · Data Alignment Information: Detailed information about the alignment process, including any challenges or discrepancies encountered and how they were resolved. This includes a summary of the aligned dataset, such as the number of common samples identified.
[0627] · Expiry Time: Time after which the data alignment request expires.
[0628] · Data Privacy Settings: Specifications for handling sensitive data during alignment.
[0629] Outputs Required:
[0630] · When the subscription is accepted: Alignment Correlation ID (for managing the alignment process), Expiry Time (if applicable based on operator policy).
[0631] Outputs, Optional:None.
[0632] Nnwdaf_MLModelTraining_DataProcessing_Sharing Service Operation
[0633] Service Operation Name:Nnwdaf_MLModelTraining_DataProcessing_Sharing .
[0634] Description: This operation facilitates the collaborative integration and sharing phase within the VFL framework. It enables each participating NWDAF to process its dataset, integrate insights, and share the resulting data representations with subsequent participants in the VFL ecosystem. This mechanism is pivotal for combining diverse data insights across the network, enhancing the federated model's predictive capabilities while upholding the stringent privacy requirements of each participant's data.
[0635] Inputs, Required:
[0636] · Local Model Insights: The insights or data representations derived from processing the local dataset with the current model parameters. These insights are crucial for enriching the federated model's learning context.
[0637] · Local Data Features: Specific features from the local dataset that have been processed to generate the aforementioned insights. This includes any transformations or feature engineering steps applied to the data.
[0638] Inputs, Optional:
[0639] · Privacy Enhancement Techniques: Parameters or methods employed to safeguard the privacy of shared data insights, such as homomorphic encryption, secure multi-party computation, or differential privacy mechanisms.
[0640] · Preceding Insights: Insights or data representations received from previous participants in the VFL chain. These are optionally incorporated into the local dataset to enrich the model's learning context.
[0641] Outputs, Required:
[0642] · Integrated Data Representations: The enhanced data representations ready to be shared with the next participant. These representations are a synthesis of local insights and, optionally, insights received from preceding models, tailored for the next phase of collaborative learning.
[0643] Outputs, Optional:
[0644] · Integration Metadata: Supplementary information regarding the data integration process, including computational metrics, resource utilization, and any challenges or anomalies encountered. This metadata can provide valuable feedback for optimizing the VFL process.
[0645] Nnwdaf_MLModelTraining _BackwardPropagation Service Operation
[0646] Service Operation Name:Nnwdaf_MLModelTraining _BackwardPropagation
[0647] Description:This operation orchestrates the model optimization phase within the VFL framework, focusing on the refinement of local models through collaborative insights. It encompasses the iterative adjustment of model parameters by leveraging gradient, or other learning information, derived from the collective learning process. This phase is critical for enhancing the model's accuracy and performance by incorporating feedback from across the VFL network, ensuring that each NWDAF's model evolves in alignment with shared objectives while maintaining data privacy.
[0648] Inputs, Required:
[0649] · Received Insights: Insights or gradient information received from subsequent participants in the VFL chain, or the aggregate loss insights in the case of the concluding participant. These insights are pivotal for recalibrating the local model's parameters.
[0650] · Local Processing Insights: Insights derived from the local model's processing, utilized to calculate adjustments for the model's parameters.
[0651] Inputs, Optional:
[0652] · Privacy Enhancement Mechanisms: Techniques or parameters implemented to safeguard the privacy of shared insights during the model optimization phase, such as secure multi-party computation or noise addition techniques.
[0653] Outputs, Required:
[0654] · Model Parameter Adjustments: Adjustments calculated for the local model's parameters, aimed at optimizing the model's performance based on the collaborative insights. These adjustments are foundational for the iterative enhancement of the model and may be shared with preceding participants to foster collective model refinement.
[0655] Outputs, Optional:
[0656] · Optimization Metadata: Supplementary information regarding the optimization process, including details on computational efficiency, resource allocation, and any obstacles or anomalies identified. This metadata is instrumental in diagnosing the optimization process, facilitating continuous improvement of the VFL methodology.
[0657] It will be appreciated that examples of the present disclosure may be realized in the form of hardware, software or a combination of hardware and software. Certain examples of the present disclosure may provide a computer program comprising instructions or code which, when executed, implement a method, system and / or apparatus in accordance with any aspect, example and / or embodiment disclosed herein. Certain embodiments of the present disclosure provide a machine-readable storage storing such a program.
[0658] The Annex to this description (found below) discloses one or more further techniques according to the present disclosure. The skilled person will appreciate that the techniques disclosed in the Annex may be used together with the techniques disclosed herein in any suitable combination.
[0659] Figure 9 is a block diagram of an exemplary network entity / function that may be used in examples of the present disclosure, such as the techniques disclosed in relation to any of the preceding figures. For example, one or more of the network entities, network functions etc. in the examples of Figures 1-8, 10-16 (or other examples herein) may comprise or be provided in the form of the entity illustrated in Figure 9. The skilled person will appreciate that a network entity / function may be implemented, for example, as a network element on a dedicated hardware, as a software instance running on a dedicated hardware, and / or as a virtualised function instantiated on an appropriate platform, e.g. on a cloud infrastructure.
[0660] The entity 900 comprises at least one of a processor (or controller) 901, a transmitter 903, a receiver 905, or a transceiver. The receiver 905 or the transceiver is configured for receiving one or more messages from one or more other network entities, for example as described above. The transmitter 903 or the transceiver is configured for transmitting one or more messages to one or more other network entities, for example as described above. The processor 901 is configured for performing one or more operations, for example according to the operations as described above.
[0661] Figure 10 illustrates a method according to an example of the present disclosure. The method is performed by a server entity, e.g. a VFL server entity.
[0662] In operation 1010, the server entity, based on a message received from a service consumer, initiates the VFL procedure.
[0663] In operation 1020, the server entity transmits a request for VFL inference to at least one VFL client entity.
[0664] In operation 1030, the server entity receives intermediate inference results generated by the at least one VFL client entity.
[0665] In operation 1040, the server entity performs inference computation on the received intermediate inference results.
[0666] Figure 11 illustrates a method according to an example of the present disclosure. The method is performed by a VFL client entity.
[0667] In operation 1110, the VFL client entity receives a request for VFL inference from a server entity.
[0668] In operation 1120, the VFL client entity, based on the request, obtains intermediate inference results using a local ML model.
[0669] In operation 1130, the VFL client entity transmits the intermediate inference results to the server entity.
[0670] Figure 12 illustrates a method according to an example of the present disclosure. The method is performed by a NEF entity.
[0671] In operation 1210, the NEF entity receives, from a server entity, a request for VFL inference to be forwarded to at least one VFL client entity.
[0672] In operation 1220, the NEF entity forwards the request for VFL inference to the at least one VFL client entity.
[0673] In operation 1230, the NEF entity receives from the at least one VFL client entity, intermediate inference results generated by the at least one VFL client entity.
[0674] In operation 1240, the NEF entity forwards the intermediate inference results to the server entity.
[0675] Figure 13 illustrates a method according to an example of the present disclosure. The method is performed by a server entity, e.g. a VFL server entity.
[0676] In operation 1310, the server entity transmits a VFL preparation request to one or more VFL client, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met.
[0677] In operation 1320, the server entity receives, from each of the one or more VFL client entity, information including an indication of whether said VFL client entity will participate.
[0678] In operation 1330, the server entity selects at least one VFL client entity to participate from among the one or more VFL client entity based on the received information.
[0679] Figure 14 illustrates a method according to an example of the present disclosure. The method is performed by a VFL client entity.
[0680] In operation 1410, the VFL client entity receives a VFL preparation request from a server entity, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met.
[0681] In operation 1420, the VFL client entity, based on the VFL preparation request, determines capability to meet model training requirements.
[0682] In operation 1430, the VFL client entity transmits, to the server entity, information including an indication of whether the VFL client entity will participate in the VFL procedure.
[0683] Figure 15 illustrates a method according to an example of the present disclosure. The method is performed by a server entity, e.g. a VFL server entity.
[0684] In operation 1510, the server entity transmits, to at least one VFL client entity, a request to perform ML model training.
[0685] In operation 1520, the server entity receives, from the at least one VFL client entity, intermediate training results or model refinement results.
[0686] In operation 1530, the server entity performs VFL computation based on the received results.
[0687] Figure 16 illustrates a method according to an example of the present disclosure. The method is performed by a VFL client entity.
[0688] In operation 1610, the VFL client entity receives, from a server entity, a request to perform ML model training.
[0689] In operation 1620, the VFL client entity trains a local ML model. Training the local ML model comprises obtaining intermediate training results of the trained local ML model or refining a previously trained model to obtain the trained local ML model.
[0690] In operation 1630, the VFL client entity shares the intermediate training results or the model refinement results with the server entity.
[0691] Further examples in accordance with the present disclosure are set out below, where the examples may be combined in any appropriate form and also combined with any of the approaches set out above.
[0692] According to an aspect of the present disclosure, there is provided a server entity for performing a vertical federated learning (VFL) procedure. The server entity is configured to: based on a message received from a service consumer, initiate the VFL procedure; transmit a request for VFL inference to at least one VFL client entity; receive intermediate inference results generated by the at least one VFL client entity; and perform inference computation on the received intermediate inference results.
[0693] The server entity may be further configured to discover or select the at least one VFL client entity to participate in the VFL procedure.
[0694] The server entity may be further configured to: aggregate the received intermediate inference results; and generate inference results based on the aggregated intermediate inference results.
[0695] The server entity may be further configured to generate the inference results for an analytics ID received from the service consumer.
[0696] The server entity may be further configured to: collect local data and generate intermediate local inference results based on the local data; and aggregate the intermediate local inference results with the received intermediate inference results.
[0697] The server entity may be further configured to: derive analytics according to the message, based on the generated inference results; and transmit the derived analytics to the service consumer.
[0698] Wherein, based on the at least one VFL client entity being an application function (AF), the request is transmitted to the at least one VFL client entity via a network exposure function (NEF) entity, and the intermediate inference results are received from the at least one VFL client entity via the NEF entity.
[0699] Wherein, based on the server entity being untrusted, the request is transmitted to a network exposure function (NEF) entity to be forwarded to the at least one VFL client entity, and the intermediate inference results are received from the NEF entity.
[0700] Wherein the transmitting the request for VFL inference to be forwarded to the at least one VFL client entity comprises transmitting, to the NEF entity, a request for VFL inference for each one of the at least one VFL client entity, wherein each request is to be forwarded to a corresponding one of the at least one VFL client entity.
[0701] Wherein the transmitting the request for VFL inference to be forwarded to the at least one VFL client entity comprises transmitting, to the NEF entity, a single request for VFL inference, wherein the request includes a correlation ID or an indication of the at least one VFL client entity.
[0702] Wherein the request comprises an ID (e.g. the correlation ID) indicating a VFL process including the VFL inference and / or VFL training.
[0703] Wherein the message is an analytics request received from the service consumer.
[0704] The server entity may be further configured to: register, to network repository function (NRF), information including at least one of network function (NF) profile, analytics ID(s), service area, VFL capability information, and time interval supporting VFL.
[0705] The server entity may be further configured to: transmit a VFL preparation request to one or more VFL client, the VFL preparation request including at least one of analytics ID, machine learning (ML) model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met.
[0706] The server entity may be further configured to: transmit dataset identifiers to the one or more VFL client entity.
[0707] The server entity may be further configured to: receive, from each of the one or more VFL client entity, information including an indication of whether said VFL client entity will participate; and select the at least one VFL client entity from among the one or more VFL client entity based on the received information.
[0708] Wherein the received information further comprises a reason why said VFL client entity cannot participate.
[0709] The server entity may be further configured to: transmit, to the at least one VFL client entity, a request to perform ML model training; and receive, from the at least one VFL client entity, intermediate training results.
[0710] The server entity may be further configured to: perform VFL computation based on the received intermediate training results.
[0711] Wherein the intermediate training results relate to a ML model associated with the VFL inference.
[0712] According to an aspect of the present disclosure, there is provided a vertical federated learning (VFL) client entity for participating in a VFL procedure. The VFL client entity is configured to: receive a request for VFL inference from a server entity; based on the request, obtain intermediate inference results using a local machine learning (ML) model; and transmit the intermediate inference results to the server entity.
[0713] Wherein, based on the at least one VFL client entity being an application function (AF), the request is received from the server entity via a network exposure function (NEF) entity, and the intermediate inference results are transmitted to the server entity via the NEF entity.
[0714] Wherein the request is received from the server entity via a network exposure function (NEF) entity; and wherein the intermediate inference results are transmitted to the server entity via the NEF entity based on the server entity being untrusted.
[0715] The VFL client entity may be further configured to: collect local data based on determining stored data is not sufficient to support the VFL inference of the request; and generate the intermediate inference results based on the collected local data.
[0716] The VFL client entity may be further configured to: determine the local data to collect based on an analytics ID included in the request for VFL inference or an indication received from the server entity via the NEF entity.
[0717] The VFL client entity may be further configured to: register, to network repository function (NRF), information including at least one of network function (NF) profile, analytics ID(s), service area, VFL capability information, and time interval supporting VFL.
[0718] The VFL client entity may be further configured to: receive a VFL preparation request from the server entity, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met.
[0719] The VFL client entity may be further configured to: based on the VFL preparation request, determine capability to meet model training requirements; and transmit, to the server entity, information including an indication of whether the VFL client entity will participate in the VFL procedure.
[0720] The VFL client entity may be further configured to: receive dataset identifier from the server entity.
[0721] Wherein the transmitted information further includes a reason why the VFL client entity cannot participate.
[0722] The VFL client entity may be further configured to: receive, from the server entity, a request to perform ML model training; and train the local ML model; wherein the trained local ML model is used to obtain the intermediate inference results.
[0723] The VFL client entity may be further configured to: obtain intermediate training results of the trained local ML model, and share the intermediate training results with the server entity; or refine a previously trained model to obtain the trained local ML model and share the model refinement results with the server entity.
[0724] Wherein the intermediate training results or the model refinement results are shared using a same ML model training service as used for receiving the request to perform ML model training.
[0725] The VFL client entity may be further configured to: share the intermediate training results with another VFL client entity.
[0726] According to an aspect of the present disclosure, there is provided a network exposure function (NEF) entity configured to: receive, from a server entity, a request for vertical federated learning (VFL) inference to be forwarded to at least one VFL client entity; forward the request for VFL inference to the at least one VFL client entity; receive, from the at least one VFL client entity, intermediate inference results generated by the at least one VFL client entity; and forward the intermediate inference results to the server entity.
[0727] Wherein the server entity is untrusted and / or the at least one VFL client entity is an application function (AF).
[0728] Wherein receiving the request for VFL inference comprises receiving, from the server entity, at least one request for VFL inference each corresponding to one of the at least one VFL client entity; and wherein forwarding the request for VFL inference to the at least one VFL client entity comprises forwarding each of the at least one request for VFL inference to the corresponding one of the at least one VFL client entity.
[0729] Wherein receiving the request for VFL inference comprises receiving, from the server entity, a single request for VFL inference, wherein the request includes a correlation ID or an indication of the at least one VFL client entity; and wherein the NEF entity is configured to: select the at least one VFL client entity to forward the request for VFL inference to, based on the correlation ID or the indication of the at least one VFL client entity, wherein forwarding the request for VFL inference comprises forwarding the request for VFL inference to the selected at least one VFL client entity; or split the request for VFL inference into at least one request for VFL inference, each of the at least one request for VFL inference corresponding to one of the at least one VFL client entity, wherein forwarding the request for VFL inference comprises forwarding each of the at least one request for VFL inference to the corresponding one of the at least one VFL client entity.
[0730] Wherein the request comprises an ID (e.g. the correlation ID) indicating a VFL process including VFL inference and VFL training, and wherein selecting the at least one VFL client entity is based on the VFL inference and VFL training indicated by the ID.
[0731] The NEF entity may be further configured to enable the server entity to subscribe, unsubscribe, notify and / or modify for artificial intelligence / machine learning (AI / ML) training toward the at least one VFL client entity via the NEF.
[0732] According to an aspect of the present disclosure, there is provided a server entity for performing a vertical federated learning (VFL) procedure. The server entity is configured to: transmit a VFL preparation request to one or more VFL client, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met; receive, from each of the one or more VFL client entity, information including an indication of whether said VFL client entity will participate; and select at least one VFL client entity to participate from among the one or more VFL client entity based on the received information.
[0733] Wherein the at least one VFL client entity is selected to participate in VFL training and / or VFL inference.
[0734] The server entity may be further configured to transmit dataset identifier to the one or more VFL client entity.
[0735] Wherein the received information further comprises a reason why said VFL client entity cannot participate.
[0736] According to an aspect of the present disclosure, there is provided a server entity for performing a vertical federated learning (VFL) procedure. The server entity is configured to: transmit, to at least one VFL client entity, a request to perform machine learning (ML) model training; receive, from the at least one VFL client entity, intermediate training results or model refinement results; and perform VFL computation based on the received results.
[0737] Wherein the intermediate training results relate to a ML model associated with a VFL inference process.
[0738] According to an aspect of the present disclosure, there is provided a vertical federated learning (VFL) client entity for participating in a VFL procedure. The VFL client entity is configured to: receive a VFL preparation request from a server entity, the VFL preparation request including at least one of analytics ID, ML model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met; based on the VFL preparation request, determine capability to meet model training requirements; and transmit, to the server entity, information including an indication of whether the VFL client entity will participate in the VFL procedure.
[0739] Wherein the model training requirements relate to VFL training and / or VFL inference.
[0740] Wherein the transmitted information further includes a reason why the VFL client entity cannot participate.
[0741] The VFL client entity may be further configured to: receive dataset identifier from the server entity.
[0742] According to an aspect of the present disclosure, there is provided a vertical federated learning (VFL) client entity for participating in a VFL procedure. The VFL client entity is configured to: receive, from a server entity, a request to perform machine learning (ML) model training; and train a local ML model; wherein training the local ML model comprises obtaining intermediate training results of the trained local ML model or refining a previously trained model to obtain the trained local ML model; and wherein VFL client entity is configured to: share the intermediate training results or the model refinement results with the server entity.
[0743] According to an aspect of the present disclosure, there is provided a method of a server entity for performing a vertical federated learning (VFL) procedure, the method comprising: based on a message received from a service consumer, initiating the VFL procedure; transmitting a request for VFL inference to at least one VFL client entity; receiving intermediate inference results generated by the at least one VFL client entity; and performing inference computation on the received intermediate inference results.
[0744] According to an aspect of the present disclosure, there is provided a method of a vertical federated learning (VFL) client entity for participating in a VFL procedure, the method comprising: receiving a request for VFL inference from a server entity; based on the request, obtaining intermediate inference results using a local machine learning (ML) model; and transmitting the intermediate inference results to the server entity.
[0745] According to an aspect of the present disclosure, there is provided a method of a network exposure function (NEF) entity, the method comprising: receiving, from a server entity, a request for vertical federated learning (VFL) inference to be forwarded to at least one VFL client entity; forwarding the request for VFL inference to the at least one VFL client entity; receiving, from the at least one VFL client entity, intermediate inference results generated by the at least one VFL client entity; and forwarding the intermediate inference results to the server entity.
[0746] According to an aspect of the present disclosure, there is provided a method of a server entity for performing a vertical federated learning (VFL) procedure, the method comprising: transmitting a VFL preparation request to one or more VFL client, the VFL preparation request including at least one of analytics ID, machine learning (ML) model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met; receiving, from each of the one or more VFL client entity, information including an indication of whether said VFL client entity will participate; and selecting at least one VFL client entity to participate from among the one or more VFL client entity based on the received information.
[0747] According to an aspect of the present disclosure, there is provided a method of a vertical federated learning (VFL) client entity for participating in a VFL procedure, the method comprising: receiving a VFL preparation request from a server entity, the VFL preparation request including at least one of analytics ID, machine learning (ML) model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met; based on the VFL preparation request, determining capability to meet model training requirements; and transmitting, to the server entity, information including an indication of whether the VFL client entity will participate in the VFL procedure.
[0748] According to an aspect of the present disclosure, there is provided a method of a server entity for performing a vertical federated learning (VFL) procedure, the method comprising: transmitting, to at least one VFL client entity, a request to perform machine learning (ML) model training; receiving, from the at least one VFL client entity, intermediate training results or model refinement results; and performing VFL computation based on the received results.
[0749] According to an aspect of the present disclosure, there is provided a method of a vertical federated learning (VFL) client entity for participating in a VFL procedure, the method comprising: receiving, from a server entity, a request to perform machine learning (ML) model training; and training a local ML model; wherein training the local ML model comprises obtaining intermediate training results of the trained local ML model or refining a previously trained model to obtain the trained local ML model; and wherein the method comprises: sharing the intermediate training results or the model refinement results with the server entity.
[0750] For all of the examples / aspects / embodiments etc. described above / herein, it should be considered that the corresponding features / operations apply in any order or combination, and that furthermore there exists the possibility to omit one or more features / operations.
[0751] Moreover, for all of the examples, embodiments, aspects etc. above, these apply to at least LTE, NR, NR NTN or IoT NTN (note this list is merely to give some examples and should not be seen as limiting), including any related signalling / messages on any of the inferences X2, Xn, NG, S1, F1, etc (again, this list is merely to give some examples and should not be seen as limiting).It will be appreciated that, in each example / embodiment / aspect etc. described above, one or more features or operations may be omitted, modified or moved (e.g., to change the order of the features or the operations), if desired and appropriate.
[0752] Additionally, where the figures illustrating example method flows include text in relation to a specific step / operation, it will be appreciated that this text is simply an example of the corresponding step / operation, where a more general definition (such as may be found in the description of the corresponding step) may apply for the step / operation.
[0753] Additionally, regarding all of the above, one or more features or operations etc. from any example / embodiment may be combined with features or operations from any other example / embodiment. That is, the present disclosure should be considered to include all combinations of examples / embodiments disclosed herein, as appropriate, as well as combinations of individual features within and between each example / embodiment, as appropriate.
[0754] The techniques described herein may be implemented using any suitably configured apparatus and / or system. Such an apparatus and / or system may be configured to perform a method according to any aspect, embodiment, example or claim disclosed herein. Such an apparatus may comprise one or more elements, for example one or more of receivers, transmitters, transceivers, processors, controllers, modules, units, and the like, each element configured to perform one or more corresponding processes, operations and / or method steps for implementing the techniques described herein. For example, an operation / function of X may be performed by a module configured to perform X (or an X-module). The one or more elements may be implemented in the form of hardware, software, or any combination of hardware and software.
[0755] It will be appreciated that examples of the present disclosure may be implemented in the form of hardware, software or any combination of hardware and software. Any such software may be stored in the form of volatile or non-volatile storage, for example a storage device like a ROM, whether erasable or rewritable or not, or in the form of memory such as, for example, RAM, memory chips, device or integrated circuits or on an optically or magnetically readable medium such as, for example, a CD, DVD, magnetic disk or magnetic tape or the like.
[0756] It will be appreciated that the storage devices and storage media are embodiments of machine-readable storage that are suitable for storing a program or programs comprising instructions that, when executed, implement certain examples of the present disclosure. Accordingly, certain examples provide a program comprising code for implementing a method, apparatus or system according to any example, embodiment, aspect and / or claim disclosed herein, and / or a machine-readable storage storing such a program. Still further, such programs may be conveyed electronically via any medium, for example a communication signal carried over a wired or wireless connection.
[0757] While the disclosure has been shown and described with reference to certain examples, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the scope of the disclosure.
[0758] Definitions / Acronyms:
[0759] In the present disclosure, the following acronyms / definitions may be used. Other acronyms may be defined in accordance with the 3GPP 5G and other relevant standards at the filing date of the present application.
[0760] 3GPP 3rdGeneration Partnership Project
[0761] 5G 5thGeneration
[0762] 5GC 5G Core
[0763] 5QI 5G QoS Identifier
[0764] 5GS 5G System
[0765] 5GSM 5G System Session Management
[0766] 5GMM 5G System Mobility Management
[0767] 6G 6thGeneration
[0768] ADRF Analytics Data Repository Function
[0769] AF Application Function
[0770] AI Artificial Intelligence
[0771] AIML Artificial Intelligence / Machine Learning
[0772] AM Acknowledged Mode
[0773] AMF Access and Mobility Management Function
[0774] AnLF Analytics Logical Function
[0775] AS Application Server
[0776] ASP Application Service Provider
[0777] ATG Air-To-Ground
[0778] AUSF Authentication Server Function
[0779] CDN Content Delivery Network
[0780] CFL Compressed Federated Learning
[0781] CN Core Network
[0782] DCAF Data Collection Application Function
[0783] DNAI Data Network Access Identifier
[0784] DNN Data Network Name
[0785] DNS Domain Name Server
[0786] DRB Data Radio Bearer
[0787] EC Energy Consumption
[0788] eNB Base Station
[0789] FL Federated Learning
[0790] FQDN Fully Qualified Domain Name
[0791] GBR Guaranteed Bit Rate
[0792] gNB 5G Base Station
[0793] GW Gateway
[0794] HSS Home Subscriber Service
[0795] HFL Horizontal Federated Learning
[0796] IAB Integrated Access and Backhaul
[0797] ID Identity / Identification
[0798] IIoT Industrial Internet of Things
[0799] IoT Internet of Things
[0800] IMEI International Mobile Equipment Identities
[0801] IP Internet Protocol
[0802] I-SMF Intermediate SMF
[0803] LADN Local Area Data Network
[0804] LL SSM Lower Layer SSM
[0805] MBMS Multimedia Broadcast / Multicast Service
[0806] MBS Multicast / Broadcast Service
[0807] MBSF Multicast / Broadcast Service Function
[0808] MBSTF Multicast / Broadcast Service Transport Function
[0809] MB-SMF Multicast / Broadcast Session Management Function
[0810] MB-UPF Multicast / Broadcast User Plane Function
[0811] KI Key Issue
[0812] ML Machine Learning
[0813] MME Mobility Management Entity
[0814] MN Master Node
[0815] MNF Monitoring Network Function
[0816] MNO Mobile Network Operator
[0817] MT Mobile Termination
[0818] MTLF Model Training Logical Function
[0819] NAS Non-Access Stratum
[0820] NB Base Station
[0821] NEF Network Exposure Function
[0822] NF Network Function
[0823] NR New Radio
[0824] NRF Network Repository Function
[0825] NG-RAN Next Generation Radio Access Network
[0826] NG-eNB Next Generation eNB
[0827] NSA Non-Standalone
[0828] NSSF Network Slice Selection Function
[0829] NTN Non-Terrestrial Networks
[0830] NW Network
[0831] NWDAF Network Data Analytics Function
[0832] OS Operating System
[0833] OSAPP OS Application
[0834] PCF Policy Control Function
[0835] PCO Protocol Configuration Options
[0836] PDR Packet Detection Rule
[0837] PDU Protocol Data Unit
[0838] PTM Point To Multipoint
[0839] PTP Point to Point
[0840] QFI QoS Flow Identifier (ID)
[0841] QoS Quality of Service
[0842] RACH Random Access Channel
[0843] PLMN Public Land Mobile Network
[0844] RAN Radio Access Network
[0845] Rel Release
[0846] SMF Session Management Function
[0847] SN Secondary Node
[0848] S-NSSAI Single Network Slice Selection Assistance Information
[0849] SSB Synchronization Signal Block
[0850] SSM Source Specific IP Multicast address
[0851] SSC Session and Service Continuity
[0852] SRB Signaling Radio Bearer
[0853] SUPI Subscription Permanent Identifier
[0854] TA Tracking Area
[0855] TAI Tracking Area Identity
[0856] TE Terminal Equipment
[0857] TM Transparent Mode
[0858] TMGI Temporary Mobile Group Identity
[0859] TR Technical Report
[0860] TS Technical Specification
[0861] UAV Unmanned Aerial Vehicle
[0862] UDM Unified Data Manager
[0863] UDR Unified Data Repository
[0864] UE User Equipment
[0865] UL Uplink
[0866] UM Unacknowledged Mode
[0867] UP User Plane
[0868] UPF User Plane Function
[0869] URLLC Ultra-Reliable and Low-Latency Communication
[0870] URSP UE Route Selection Policy
[0871] VFL Vertical Federated Learning
Claims
1.A method of a vertical federated learning (VFL) server for a VFL inference procedure in a wireless communication network, the method comprising:transmitting, to one or more VFL clients, a VFL inference request message including a VFL correlation identity (ID);receiving, from the one or more VFL clients, a response message including an intermediate inference result; andgenerating, a VFL inference result by aggregating the intermediate inference result based on the VFL correlation ID.2.The method of claim 1, wherein the VFL inference request message is transmitted via a network exposure function (NEF) in case that the one or more VFL clients is an untrusted application function (AF).3.The method of claim 1, wherein the VFL inference request message is triggered by an analytics consumer network function (NF).4.The method of claim 3, wherein the VFL inference result is generated for an analytics ID received from the analytics consumer NF.5.The method of claim 4, further comprising:obtaining analytics based on the generated VFL inference result; andtransmitting the analytics to the analytics consumer NF.6.The method of claim 1, further comprising:registering, to a network repository function (NRF), information including at least one of NF profile, analytics ID(s), service area, VFL capability information, or time interval supporting VFL.7.The method of claim 1, further comprising:transmit a VFL preparation request to the one or more VFL clients, the VFL preparation request including at least one of analytics ID, machine learning (ML) model interoperability information, available data requirement, availability time requirement, or information for checking whether a ML model training requirement can be met.8.The method of claim 7, further comprising:transmitting dataset identifier to the one or more VFL clients.9.The method of claim 8, further comprising:receiving, from the one or more VFL clients, information including an indication of whether the one or more VFL clients participate; andselecting a VFL client from among the one or more VFL clients based on the received information including the indication.10.The method of claim 9,wherein the received information including the indication further includes a reason why the one or more VFL clients does not participate.11.The method of claim 1, further comprising:transmitting, to the one or more VFL clients, a request to perform ML model training;receiving, from the one or more VFL clients, an intermediate training result; andperforming VFL computation based on the received intermediate training result.12.A vertical federated learning (VFL) server for a VFL inference procedure in a wireless communication network, the VFL server comprising:a transceiver; anda processor configured to control the transceiver, wherein the processor is configured to:transmit, to one or more VFL clients, a VFL inference request message including a VFL correlation identity (ID);receive, from the one or more VFL clients, a response message including an intermediate inference result; andgenerate, a VFL inference result by aggregating the intermediate inference result based on the VFL correlation ID.13.The VFL server of claim 12, wherein the VFL server is configured to perform any one of the methods of claims 2 to 11.14.A method of a vertical federated learning (VFL) client for a VFL inference procedure in a wireless communication network, the method comprising:receiving, from a VFL server, a VFL inference request message including a VFL correlation identity (ID);generating an intermediate inference result based on the VFL correlation ID; andtransmitting, to the VFL server, a response message including the intermediate inference result.15.The method of claim 14, wherein the VFL inference request message is received via a network exposure function (NEF) in case that the VFL client is an untrusted application function (AF).
Citation Information
Patent Citations
Communication method and apparatus
US20230308930A1
Cited By
Vertical federated learning feature and sample alignment
GB2701859A