Zone gradient diffusion (ZGD) for zone-based federated learning

By considering machine learning updates of adjacent zones in zone-based federated learning and dynamically adjusting model weights, the problem of neglected updates caused by inflexible zone boundaries is solved, and the accuracy and flexibility of the model are improved.

CN120035832APending Publication Date: 2025-05-23QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072866.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-05
Filing Date
2023-09-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In district-based federated learning, inflexible district boundaries lead to neglect of relevant model updates from outside the district of interest.

Method used

By considering machine learning updates from neighboring zones as a supplement to the region in a centralized server, the model weights are dynamically updated, and the self-attention mechanism is used to adjust the degree of influence of neighboring zones according to similarity parameters.

Benefits of technology

Improve the accuracy and flexibility of the zone model, ensure that relevant model updates are effectively utilized, and improve the performance of federated learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035832A_ABST
    Figure CN120035832A_ABST
Patent Text Reader

Abstract

A processor-implemented method includes receiving a machine learning model update from a client in a federated learning system. The method also includes determining a fixed local area associated with each of the clients, the fixed local area having a first fixed boundary. The method includes updating model weights of a central machine learning model based on local machine learning updates for a local subset of the clients corresponding to the fixed local region. The method includes updating model weights of the central machine learning model based on neighboring machine learning updates for neighboring subsets of clients. The adjacent subset corresponds to a fixed adjacent region adjacent to the fixed local region and having a second fixed boundary. The neighboring machine learning updates have different weights than the local machine learning updates when updating the model weights. The values of the different weights correspond to similarity parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. patent application No. 18 / 461,410, filed on September 5, 2023, entitled “ZONE GRADIENT DIFFUSION (ZGD) FOR ZONE-BASED FEDERATED LEARNING,” which claims the benefit of U.S. Provisional Patent Application No. 63 / 418,454, filed on October 21, 2022, and entitled “ZONE GRADIENT DIFFUSION (ZGD) FOR ZONE BASED FEDERATED LEARNING,” the disclosures of which are expressly incorporated by reference in their entireties. Technical Field

[0003] The present disclosure relates generally to wireless communications, and more particularly to methods for a zone gradient diffusion (ZGD) technique for zone-based federated learning. Background Art

[0004] Federated learning is a machine learning technique that trains a federated learning model across multiple decentralized edge devices or servers that keep local data samples without sharing the data samples with a central server. Federated learning provides the benefits of privacy-preserving machine learning and continuous learning on the edge. However, the performance of federated learning is impaired when the data at the device is non-independent and identically distributed (non-IID). Data augmentation is a method to address non-IID data. Another method is zone-based federated learning.

[0005] Zone-based federated learning groups participating devices into zones that contribute to the distribution of non-IID data at the edge. However, data from devices outside the zone of interest may be relevant. If the zone boundaries are not flexible, those relevant model updates will be ignored. Techniques for capturing relevant neighboring device information would be desirable. Summary of the invention

[0006] In various aspects of the present disclosure, a method implemented by a processor includes receiving machine learning model updates from multiple clients in a federated learning system. The method also includes determining a fixed local region associated with each of the clients, the fixed local region having a first fixed boundary. The method also includes updating a model weight of a central machine learning model based on a local machine learning update for a local subset of the client. The local subset corresponds to the fixed local region. The method also includes updating the model weight of the central machine learning model based on a neighboring machine learning update for a neighboring subset of the client. The neighboring subset corresponds to a fixed neighboring region that is adjacent to the fixed local region. When updating the model weight, the neighboring machine learning update has a different weight than the local machine learning update. The values ​​of the different weights correspond to a similarity parameter, and the fixed neighboring region has a second fixed boundary.

[0007] Other aspects of the present disclosure relate to an apparatus. The apparatus has at least one memory and one or more processors coupled to the at least one memory. The processor is configured to receive machine learning model updates from multiple clients in a federated learning system. The processor is also configured to determine a fixed local region associated with each of the clients, the fixed local region having a first fixed boundary. The processor is further configured to update the model weights of the central machine learning model based on local machine learning updates for a local subset of the client. The local subset corresponds to the fixed local region. The processor is also configured to update the model weights of the central machine learning model based on neighboring machine learning updates for neighboring subsets of the client. The neighboring subset corresponds to a fixed neighboring region that is adjacent to the fixed local region. When updating the model weights, the neighboring machine learning updates have different weights from the local machine learning updates. The values ​​of the different weights correspond to similarity parameters, and the fixed neighboring regions have a second fixed boundary.

[0008] Other aspects of the present disclosure relate to an apparatus. The apparatus includes a component for receiving machine learning model updates from multiple clients in a federated learning system. The apparatus also includes a component for determining a fixed local region associated with each of the clients, the fixed local region having a first fixed boundary. The apparatus also includes a component for updating a model weight of a central machine learning model based on a local machine learning update for a local subset of the client. The local subset corresponds to the fixed local region. The apparatus also includes a component for updating the model weight of the central machine learning model based on a neighboring machine learning update for a neighboring subset of the client. The neighboring subset corresponds to a fixed neighboring region that is adjacent to the fixed local region. When updating the model weight, the neighboring machine learning update has a different weight than the local machine learning update. The values ​​of the different weights correspond to a similarity parameter, and the fixed neighboring region has a second fixed boundary.

[0009] In other aspects of the present disclosure, a non-transitory computer-readable medium having program code recorded thereon is disclosed. The program code is executed by a processor and includes program code for receiving machine learning model updates from clients in a federated learning system. The program code also includes a fixed region associated with each of the clients, the fixed local region having a first fixed boundary. The program code also includes program code for updating the model weights of the central machine learning model based on local machine learning updates for a local subset of the client. The local subset corresponds to the fixed local region. The program code also includes program code for updating the model weights of the central machine learning model based on neighboring machine learning updates for neighboring subsets of the client. The neighboring subset corresponds to a fixed neighboring region that is adjacent to the fixed local region. When updating the model weights, the neighboring machine learning updates have different weights from the local machine learning updates. The values ​​of the different weights correspond to similarity parameters, and the fixed neighboring regions have a second fixed boundary.

[0010] Aspects generally include methods, apparatus, systems, computer program products, non-transitory computer readable media, user equipment, base stations, wireless communication devices, and processing systems as substantially described with reference to and as illustrated in the accompanying drawings and description.

[0011] The features and technical advantages of examples according to the present disclosure have been outlined quite broadly above so that the following specific embodiments may be better understood. Additional features and advantages will be described. The disclosed concepts and specific examples may be readily utilized as a basis for modifying or designing other structures for achieving the same purpose of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the disclosed concepts, both in terms of their organization and method of operation, and the associated advantages will be better understood by considering the following description in conjunction with the accompanying drawings. Each of the figures in the accompanying drawings is provided for the purpose of illustration and description and not as a definition of limitations of the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order that the features of the present disclosure may be understood in detail, a more specific description may be made with reference to various aspects, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only certain aspects of the present disclosure and therefore should not be considered as limiting the scope thereof, as the description may allow for other equally effective aspects. The same reference numerals in different drawings may identify the same or similar elements.

[0013] Figure 1 is a block diagram conceptually illustrating an example of a wireless communication network in accordance with various aspects of the present disclosure.

[0014] Figure 2is a block diagram conceptually illustrating an example of a base station communicating with a user equipment (UE) in a wireless communication network according to various aspects of the present disclosure.

[0015] Figure 3 is a block diagram illustrating an example decomposed base station architecture in accordance with aspects of the present disclosure.

[0016] Figure 4 An example implementation of designing a neural network using a system on a chip (SOC) including a general-purpose processor in accordance with certain aspects of the present disclosure is illustrated.

[0017] Figure 5 is a block diagram illustrating an exemplary deep convolutional network (DCN) according to aspects of the present disclosure.

[0018] Figure 6 is a diagram illustrating an example of different zones in a federated learning system according to various aspects of the present disclosure.

[0019] Figure 7 is a diagram illustrating an example zone network topology for zone-based federated learning in accordance with various aspects of the present disclosure.

[0020] Figure 8 is a block diagram illustrating an example of participating devices including a federated learning (FL) manager according to various aspects of the present disclosure.

[0021] Fig. 9 is a timeline illustrating a zone membership check in accordance with various aspects of the present disclosure.

[0022] Fig.10 is a diagram illustrating exemplary pseudo code for implementing zone gradient diffusion according to various aspects of the present disclosure.

[0023] Fig.11 is a flow chart illustrating an example process performed, for example, by a federated learning device according to various aspects of the present disclosure. DETAILED DESCRIPTION

[0024] The following is a more comprehensive description of various aspects of the present disclosure with reference to the accompanying drawings. However, the present disclosure can be embodied in many different forms, and should not be interpreted as being limited to any specific structure or function presented throughout the present disclosure. Instead, these aspects are provided so that the present disclosure will be thorough and complete, and the scope of the present disclosure will be fully conveyed to those skilled in the art. Based on the teachings, those skilled in the art should recognize that the scope of the present disclosure is intended to cover any aspect of the present disclosure, whether the aspect is implemented independently of any other aspect of the present disclosure or implemented in combination with any other aspect. For example, a device or a method can be implemented using any number of aspects described. In addition, the scope of the present disclosure is intended to cover such devices or methods that are practiced using other structures, functionality, or structures and functionality as supplements to the various aspects of the present disclosure described or in addition. It should be understood that any aspect of the present disclosure disclosed can be embodied by one or more elements of the claims.

[0025] Several aspects of telecommunication systems will now be presented with reference to various devices and techniques. These devices and techniques will be described in the following detailed description and illustrated in the accompanying drawings by various blocks, modules, components, circuits, steps, processes, algorithms, etc. (collectively referred to as "elements"). These elements may be implemented using hardware, software, or a combination thereof. Whether such elements are implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system.

[0026] It should be noted that although various aspects may be described using terms commonly associated with 5G and later wireless technologies, various aspects of the present disclosure may be applied in communication systems based on other generations (such as and including 3G and / or 4G technologies).

[0027] Federated learning is a machine learning technique that trains machine learning models across multiple decentralized edge devices or servers that keep local data samples without sharing the data samples with a central server. Federated learning provides the benefits of privacy-preserving machine learning and continuous learning on the edge. However, the performance of federated learning is impaired when the data at the device is non-independent and identically distributed (non-IID). Data augmentation is a method to address non-IID data. Another method is zone-based federated learning.

[0028] Zone-based federated learning groups participating devices into zones that contribute to the distribution of non-IID data at the edge. However, data from devices outside the zone of interest may be relevant. If the zone boundaries are not flexible, those relevant model updates will be ignored. Various aspects of the present disclosure include a method for considering machine learning updates (or gradients) from devices in neighboring zones as a supplement to machine learning updates (or gradients) from local devices in the zone corresponding to the local machine learning model being updated. In these aspects, when updating the model weights of the local machine learning model for the zone of interest, the centralized server considers machine learning updates (or gradients) from neighboring zones as a supplement to the local machine learning updates (or gradients) from the zone of interest. The extent to which neighboring machine learning weights affect the model used for the zone of interest can be learned over time.

[0029] Certain aspects of the subject matter described in this disclosure may be implemented to achieve one or more of the following potential advantages. In some examples, the described techniques (such as updating model weights based on local machine learning updates and neighboring machine learning updates) can improve a zone model by aggregating contextual information derived from local gradients of neighboring zones.

[0030] Figure 1 1 is a diagram illustrating a network 100 in which various aspects of the present disclosure may be practiced. The network 100 may be a 5G or NR network, or some other wireless network (such as an LTE network). The wireless network 100 may include a plurality of BSs 110 (shown as BSs 110a, BSs 110b, BSs 110c, and BSs 110d) and other network entities. A BS is an entity that communicates with a user equipment (UE), and may also be referred to as a base station, an NR BS, a Node B, a gNB, a 5G Node B, an access point, a transmission and reception point (TRP), a network node, a network entity, etc. The BS may be implemented as an aggregated base station, a decomposed base station, an integrated access and backhaul (IAB) node, a relay node, a side link node, etc. The BS may be implemented in an aggregated or monolithic base station architecture, or alternatively in a decomposed base station architecture, and may include one or more of a central unit (CU), a distributed unit (DU), a radio unit (RU), a near real-time (near RT) RAN intelligent controller (RIC), or a non-real-time (non-RT) RIC. Each BS can provide communication coverage for a specific geographical area. In 3GPP, the term "cell" can refer to the coverage area of ​​a BS and / or a BS subsystem serving the coverage area, depending on the context in which the term is used.

[0031] A BS may provide communication coverage for a macro cell, a pico cell, a femto cell, and / or another type of cell. A macro cell may cover a relatively large geographic area (e.g., a radius of several kilometers) and may allow unrestricted access by UEs with service subscriptions. A pico cell may cover a relatively small geographic area and may allow unrestricted access by UEs with service subscriptions. A femto cell may cover a relatively small geographic area (e.g., a home) and may allow restricted access by UEs associated with the femto cell (e.g., UEs in a closed subscriber group (CSG)). A BS for a macro cell may be referred to as a macro BS. A BS for a pico cell may be referred to as a pico BS. A BS for a femto cell may be referred to as a femto BS or a home BS. In Figure 1 In the example shown in , BS 110a may be a macro BS for macro cell 102a, BS 110b may be a pico BS for pico cell 102b, and BS 110c may be a femto BS for femto cell 102c. A BS may support one or more (e.g., three) cells. The terms "eNB", "base station", "NR BS", "gNB", "AP", "Node B", "5G NB", "TRP", and "cell" may be used interchangeably.

[0032] In some aspects, the cell need not be stationary, and the geographic area of ​​the cell may move according to the location of the mobile BS. In some aspects, the BSs may be interconnected with each other and / or to one or more other BSs or network nodes (not shown) in the wireless network 100 through various types of backhaul interfaces (such as direct physical connections, virtual networks, etc.) using any suitable transport network.

[0033] The wireless network 100 may also include a relay station. A relay station is an entity that can receive transmissions of data from an upstream station (e.g., a BS or a UE) and transmit the transmissions of the data to a downstream station (e.g., a UE or a BS). A relay station may also be a UE that can relay transmissions for other UEs. Figure 1 In the example shown in FIG. 1 , a relay station 110 d may communicate with a macro BS 110 a and a UE 120 d to facilitate communication between the BS 110 a and the UE 120 d. A relay station may also be referred to as a relay BS, a relay base station, a relay, or the like.

[0034] The wireless network 100 may be a heterogeneous network including different types of BSs (e.g., macro BSs, pico BSs, femto BSs, relay BSs, etc.). These different types of BSs may have different transmit power levels, different coverage areas, and different effects on interference in the wireless network 100. For example, a macro BS may have a high transmit power level (e.g., 5 watts to 40 watts), while a pico BS, a femto BS, and a relay BS may have a lower transmit power level (e.g., 0.1 watt to 2 watts).

[0035] A network controller 130 may be coupled to a group of BSs and may provide coordination and control for these BSs. The network controller 130 may communicate with the BSs via a backhaul. The BSs may also communicate with each other (eg, directly or indirectly via a wireless or wired backhaul).

[0036] UE 120 (e.g., 120a, 120b, 120c) may be dispersed throughout the wireless network 100, and each UE may be stationary or mobile. UE may also be referred to as an access terminal, terminal, mobile station, subscriber unit, station, etc. UE may be a cellular phone (e.g., a smart phone), a personal digital assistant (PDA), a wireless modem, a wireless communication device, a handheld device, a laptop computer, a cordless phone, a wireless local loop (WLL) station, a tablet computer, a camera, a gaming device, a netbook, a smartbook, an ultrabook, a medical device or equipment, a biometric sensor / device, a wearable device (smart watch, smart clothing, smart glasses, smart wristband, smart jewelry (e.g., smart ring, smart bracelet)), an entertainment device (e.g., a music or video device or a satellite radio), a component or sensor of a vehicle, a smart meter / sensor, industrial manufacturing equipment, a global positioning system device, or any other suitable device configured to communicate via a wireless or wired medium.

[0037] Some UEs may be considered as machine type communication (MTC) or evolved or enhanced machine type communication (eMTC) UEs. For example, MTC and eMTC UEs include robots, drones, remote devices, sensors, meters, monitors, location tags, etc. that can communicate with a base station, another device (e.g., a remote device), or some other entity. A wireless node may provide connectivity to or to a network (e.g., a wide area network (such as the Internet) or a cellular network), for example, via a wired or wireless communication link. Some UEs may be considered as Internet of Things (IoT) devices and / or may be implemented as NB-IoT (narrowband Internet of Things) devices. Some UEs may be considered as user premises equipment (CPE). UE 120 may be included in a housing that houses components of UE 120, such as a processor component, a memory component, etc.

[0038] Generally speaking, any number of wireless networks can be deployed in a given geographic area. Each wireless network can support a specific radio access technology (RAT) and can operate on one or more frequencies. RAT can also be referred to as radio technology, air interface, etc. Frequency can also be referred to as carrier, frequency channel, etc. In a given geographic area, each frequency can support a single RAT to avoid interference between wireless networks of different RATs. In some cases, NR or 5G RAT networks can be deployed.

[0039] In some aspects, two or more UEs 120 (e.g., shown as UE 120a and UE 120e) may communicate directly (e.g., without using base station 110 as an intermediary to communicate with each other) using one or more sidelink channels. For example, UE 120 may communicate using peer-to-peer (P2P) communication, device-to-device (D2D) communication, vehicle-to-everything (V2X) protocols (e.g., which may include vehicle-to-vehicle (V2V) protocols, vehicle-to-infrastructure (V2I) protocols, etc.), mesh networks, etc. In this case, UE 120 may perform scheduling operations, resource selection operations, and / or other operations described elsewhere herein performed by base station 110. For example, base station 110 may configure UE 120 via downlink control information (DCI), radio resource control (RRC) signaling, medium access control-control element (MAC-CE), or via system information (e.g., system information block (SIB)).

[0040] The base station 110 may include a zone gradient diffusion module 140. The network controller 130 may also or alternatively include a zone gradient diffusion module 150. The zone gradient diffusion modules 140, 150 may perform various functions, such as referring to Fig.11 One or more of the elements of process 1100 are described.

[0041] As indicated above, Figure 1 This is provided as an example only. Other examples can be found in the reference Figure 1 The examples described are different.

[0042] Figure 2 A block diagram of a design 200 of a base station 110 and a UE 120 is shown, which may be Figure 1 A base station in the base station and Figure 1 Base station 110 may be equipped with T antennas 234a through 234t, and UE 120 may be equipped with R antennas 252a through 252r, where in general T≧1 and R≧1.

[0043] At the base station 110, the transmit processor 220 may receive data for one or more UEs from the data source 212, select one or more modulation and coding schemes (MCS) for each UE based at least in part on a channel quality indicator (CQI) received from the UE, process (e.g., encode and modulate) the data for the UE based at least in part on the MCS selected for each UE, and provide data symbols for all UEs. Reducing the MCS lowers throughput but increases the reliability of transmission. The transmit processor 220 may also process system information (e.g., for semi-static resource allocation information (SRPI), etc.) and control information (e.g., CQI requests, grants, upper layer signaling, etc.), and provide overhead symbols and control symbols. The transmit processor 220 may also generate reference symbols for reference signals (e.g., cell-specific reference signals (CRS)) and synchronization signals (e.g., primary synchronization signals (PSS) and secondary synchronization signals (SSS)). The transmit (TX) multiple-input multiple-output (MIMO) processor 230 may perform spatial processing (e.g., pre-decoding) on ​​data symbols, control symbols, overhead symbols, and / or reference symbols where applicable, and may provide T output symbol streams to T modulators (MOD) 232a to 232t. Each modulator 232 may process a corresponding output symbol stream (e.g., for orthogonal frequency division multiplexing (OFDM) or the like) to obtain an output sample stream. Each modulator 232 may further process (e.g., convert to analog, amplify, filter, and up-convert) the output sample stream to obtain a downlink signal. The T downlink signals from modulators 232a to 232t may be transmitted via T antennas 234a to 234t, respectively. According to various aspects described in more detail below, position coding may be used to generate synchronization signals to convey additional information.

[0044] At the UE 120, antennas 252a to 252r may receive downlink signals from the base station 110 and / or other base stations, and may provide received signals to demodulators (DEMODs) 254a to 254r, respectively. Each demodulator 254 may condition (e.g., filter, amplify, downconvert, and digitize) the received signal to obtain input samples. Each demodulator 254 may further process the input samples (e.g., for OFDM, etc.) to obtain received symbols. The MIMO detector 256 may obtain received symbols from all R demodulators 254a to 254r, perform MIMO detection on the received symbols (if applicable), and provide detected symbols. The receive processor 258 may process (e.g., demodulate and decode) the detected symbols, provide decoded data for the UE 120 to the data sink 260, and provide decoded control information and system information to the controller / processor 280. The channel processor may determine reference signal received power (RSRP), received signal strength indicator (RSSI), reference signal received quality (RSRQ), channel quality indicator (CQI), etc. In some aspects, one or more components of UE 120 may be included in a housing.

[0045] On the uplink, at the UE 120, a transmit processor 264 may receive data from a data source 262 and control information (e.g., for reports including RSRP, RSSI, RSRQ, CQI, etc.) from a controller / processor 280 and process the data and control information. The transmit processor 264 may also generate reference symbols for one or more reference signals. The symbols from the transmit processor 264 may be pre-decoded by a TX MIMO processor 266, if applicable, further processed by modulators 254a to 254r (e.g., for discrete Fourier transform spread OFDM (DFT-s-OFDM), CP-OFDM, etc.), and transmitted to the base station 110. At the base station 110, uplink signals from the UE 120 and other UEs may be received by the antenna 234, processed by the demodulator 254, detected by the MIMO detector 236 (if applicable), and further processed by the receive processor 238 to obtain decoded data and control information transmitted by the UE 120. The receive processor 238 may provide the decoded data to the data sink 239 and the decoded control information to the controller / processor 240. The base station 110 may include a communication unit 244 and communicate with the network controller 130 via the communication unit 244. The network controller 130 may include a communication unit 294, a controller / processor 290, and a memory 292.

[0046] The controller / processor 240 of the base station 110, the controller / processor 280 of the UE 120, and / or Figure 2 Any other components of the controller / processor 240 of the base station 110, the controller / processor 280 of the UE 120, and / or Figure 2 Any other component of may perform or direct e.g. Fig.11 The processes and / or operations of other processes as described. Memory 242 and memory 282 may store data and program codes for base station 110 and UE 120, respectively. Scheduler 246 may schedule UEs for data transmission on the downlink and / or uplink.

[0047] In some aspects, the UE 120 may include means for receiving, means for determining, means for updating, and means for learning. In some aspects, the base station 110 may include means for sending, means for determining, means for updating, and means for learning. Such means may include combining Figure 2 One or more components of a UE 120 or base station 110 are described.

[0048] As indicated above, Figure 2 This is provided as an example only. Other examples can be found in the reference Figure 2 The examples described are different.

[0049] In some cases, different types of devices supporting different types of applications and / or services may coexist in a cell. Examples of different types of devices include UE phones, user premises equipment (CPE), vehicles, Internet of Things (IoT) devices, etc. Examples of different types of applications include ultra-reliable low-latency communications (URLLC) applications, massive machine type communications (mMTC) applications, enhanced mobile broadband (eMBB) applications, vehicle-to-everything (V2X) applications, etc. In addition, in some cases, a single device may support different applications or services at the same time.

[0050] The deployment of a communication system such as a 5G New Radio (NR) system may be arranged with various components or constituent parts in a variety of ways. In a 5G NR system or network, a network node, a network entity, a mobility element of a network, a radio access network (RAN) node, a core network node, a network element or network equipment such as a base station (BS) or one or more units (or one or more components) performing base station functionality may be implemented in an aggregated or decomposed architecture. For example, a BS such as a Node B (NB), an evolved NB (eNB), an NR BS, a 5G NB, an access point (AP), a transmit and receive point (TRP) or a cell, etc. may be implemented as an aggregated base station (also referred to as an independent BS or a monolithic BS) or a decomposed base station.

[0051] A converged base station can be configured to utilize a radio protocol stack that is physically or logically integrated within a single RAN node. A split base station can be configured to utilize a protocol stack that is physically or logically distributed between two or more units, such as one or more central or centralized units (CUs), one or more distributed units (DUs), or one or more radio units (RUs). In some aspects, a CU can be implemented within a RAN node, and one or more DUs can be co-located with the CU, or alternatively, can be geographically or virtually distributed among one or more other RAN nodes. A DU can be implemented to communicate with one or more RUs. Each of the CU, DU, and RU can also be implemented as a virtual unit (e.g., a virtual central unit (VCU), a virtual distributed unit (VDU), or a virtual radio unit (VRU)).

[0052] Base station type operations or network designs can consider the aggregation characteristics of base station functionality. For example, split base stations can be utilized in an integrated access backhaul (IAB) network, an open radio access network (O-RAN, such as a network configuration initiated by the O-RAN Alliance), or a virtualized radio access network (vRAN, also known as a cloud radio access network (C-RAN)). Splitting can include distributing functionality across two or more units at various physical locations, as well as virtually distributing the functionality of at least one unit, which can enable flexibility in network design. The various units of a split base station or a split RAN architecture can be configured for wired or wireless communication with at least one other unit.

[0053] Figure 3 A diagram illustrating an example split base station 300 architecture is shown. The split base station 300 architecture can include one or more central units (CUs) 310, which can communicate directly with a core network 320 via a backhaul link, or indirectly with the core network 320 through one or more split base station units, such as a near real-time (near RT) RAN intelligent controller (RIC) 325 via an E2 link, or a non-real-time (non RT) RIC 315 associated with a service management and orchestration (SMO) framework 305, or both. The CU 310 can communicate with one or more distributed units (DUs) 330 via respective midhaul links, such as an F1 interface. The DU 330 can communicate with one or more radio units (RUs) 340 via respective fronthaul links. The RU 340 can communicate with a respective UE 120 via one or more radio frequency (RF) access links. In some embodiments, the UE 120 can be served simultaneously by multiple RUs 340.

[0054] Each of these units (e.g., CU 310, DU 330, RU 340, and near-RT RIC 325, non-RTRIC 315, and SMO framework 305) may include one or more interfaces, or may be coupled to one or more interfaces configured to receive or send signals, data, or information (collectively referred to as signals) via a wired or wireless transmission medium. Each of these units or an associated processor or controller that provides instructions to the communication interface of these units may be configured to communicate with one or more of the other units via a transmission medium. For example, these units may include a wired interface that is configured to receive or send signals to one or more of the other units via a wired transmission medium. In addition, the unit may include a wireless interface that may include a receiver, a transmitter, or a transceiver (such as a radio frequency (RF) transceiver) that is configured to receive or send signals, or both, to one or more of the other units via a wireless transmission medium.

[0055] In some aspects, CU 310 may host one or more higher layer control functions. Such control functions may include radio resource control (RRC), packet data convergence protocol (PDCP), or service data adaptation protocol (SDAP), etc. Each control function may be implemented using an interface configured to communicate signals with other control functions hosted by CU 310. CU 310 may be configured to handle user plane functions (e.g., central unit-user plane (CU-UP)), control plane functions (e.g., central unit-control plane (CU-CP)), or a combination thereof. In some specific implementations, CU 310 may be logically split into one or more CU-UP units and one or more CU-CP units. When implemented in an O-RAN configuration, the CU-UP unit may communicate bidirectionally with the CU-CP unit via an interface (such as an E1 interface). As needed, CU 310 may be implemented to communicate with DU 330 for network control and signaling.

[0056] DU 330 may correspond to a logical unit that includes one or more base station functions for controlling the operation of one or more RU 340. In some aspects, DU 330 may host one or more of a radio link control (RLC) layer, a medium access control (MAC) layer, and one or more high physical (PHY) layers (such as modules for forward error correction (FEC) encoding and decoding, scrambling, modulation and demodulation, etc.) depending at least in part on a functional partitioning such as that defined by the Third Generation Partnership Project (3GPP). In some aspects, DU 330 may also host one or more low PHY layers. Each layer (or module) may be implemented using an interface that is configured to communicate signals with other layers (and modules) hosted by DU 330 or with control functions hosted by CU 310.

[0057] The lower layer functionality may be implemented by one or more RUs 340. In some deployments, a RU 340 controlled by a DU 330 may correspond to a logical node that hosts RF processing functions or low PHY layer functions (such as performing fast Fourier transform (FFT), inverse FFT (iFFT), digital beamforming, or physical random access channel (PRACH) extraction and filtering, etc.), or both, based at least in part on a functional split (such as a lower layer functional split). In such an architecture, the RU 340 may be implemented to handle over-the-air (OTA) communications with one or more UEs 120. In some implementations, real-time and non-real-time aspects of control plane communications and user plane communications with the RU 340 may be controlled by the corresponding DU 330. In some scenarios, this configuration may enable the DU 330 and the CU 310 to be implemented in a cloud-based RAN architecture (such as a vRAN architecture).

[0058] The SMO framework 305 may be configured to support RAN deployment and provisioning of non-virtualized network elements and virtualized network elements. For non-virtualized network elements, the SMO framework 305 may be configured to support the deployment of dedicated physical resources for RAN coverage requirements, which may be managed via an operation and maintenance interface (such as an O1 interface). For virtualized network elements, the SMO framework 305 may be configured to interact with a cloud computing platform (such as an open cloud (O-Cloud) platform 390) to perform network element lifecycle management (such as instantiating virtualized network elements) via a cloud computing platform interface (such as an O2 interface). Such virtualized network elements may include, but are not limited to, CU 310, DU 330, RU 340, and near-RT RIC 325. In some specific implementations, the SMO framework 305 may communicate with hardware aspects of the 4G RAN (such as an open eNB (O-eNB) 311) via the O1 interface. Additionally, in some specific implementations, the SMO framework 305 may communicate directly with one or more RUs 340 via the O1 interface. The SMO framework 305 may also include a non-RT RIC 315 configured to support the functionality of the SMO framework 305 .

[0059] The non-RT RIC 315 may be configured to include logic functions that enable non-real-time control and optimization of RAN elements and resources, artificial intelligence / machine learning (AI / ML) workflows including model training and updating, or policy-based guidance of applications / features in the near-RT RIC 325. The non-RT RIC 315 may be coupled to or in communication with the near-RT RIC 325 (such as via an A1 interface). The near-RT RIC 325 may be configured to include logic functions that enable near-real-time control and optimization of RAN elements and resources via data collection and actions on an interface (such as via an E2 interface) that connects one or more CUs 310, one or more DUs 330, or both, and the O-eNB 311 with the near-RT RIC 325.

[0060] In some implementations, in order to generate an AI / ML model to be deployed in the near-RT RIC 325, the non-RT RIC 315 may receive parameters or external enrichment information from an external server. Such information may be utilized by the near-RT RIC 325 and may be received from a non-network data source or from a network function at the SMO framework 305 or the non-RT RIC 315. In some examples, the non-RT RIC 315 or the near-RT RIC 325 may be configured to tune RAN behavior or performance. For example, the non-RT RIC 315 may monitor long-term trends and patterns of performance and employ AI / ML models to perform corrective actions through the SMO framework 305 (such as via reconfiguration of O1) or via the creation of RAN management policies (such as A1 policies).

[0061] Figure 4 An example implementation of a system on chip (SOC) 400 that may include a central processing unit (CPU) 402 or a multi-core CPU configured for zone gradient diffusion according to certain aspects of the present disclosure is illustrated. The SOC 400 may be included in a base station 110 or a UE 120. Variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device (e.g., a neural network with weights), delays, frequency bin information, and task information may be stored in a memory block associated with a neural processing unit (NPU) 408, a memory block associated with the CPU 402, a memory block associated with a graphics processing unit (GPU) 404, a memory block associated with a digital signal processor (DSP) 406, a memory block 418, or may be distributed across multiple blocks. Instructions executed at the CPU 402 may be loaded from a program memory associated with the CPU 402, or may be loaded from a memory block 418.

[0062] The SOC 400 may also include additional processing blocks tailored for specific functions, such as a GPU 404, a DSP 406, a connectivity block 410 (which may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 412 that may, for example, detect and recognize gestures. In one specific implementation, the NPU is implemented in the CPU, DSP, and / or GPU. The SOC 400 may also include a sensor processor 414, an image signal processor (ISP) 416, and / or a navigation module 420, which may include a global positioning system.

[0063] SOC 400 may be based on the ARM instruction set. In various aspects of the present disclosure, the instructions loaded into the general processor 402 may include code for receiving machine learning model updates from multiple clients in a federated learning system. The instructions may also include code for determining a fixed local region associated with each of the clients, the fixed local region having a first fixed boundary. The instructions may also include code for updating the model weights of the central machine learning model based on local machine learning updates for a local subset of the client, the local subset corresponding to the fixed local region. The instructions may also include code for updating the model weights of the central machine learning model based on neighboring machine learning updates for neighboring subsets of the client. The neighboring subset corresponds to a fixed neighboring region adjacent to the fixed local region. When updating the model weights, the neighboring machine learning updates have different weights from the local machine learning updates. The values ​​of the different weights correspond to similarity parameters, and the fixed neighboring regions have a second fixed boundary.

[0064] Deep learning architectures can perform object recognition tasks by learning to represent inputs at successively higher levels of abstraction in each layer, thereby building useful feature representations of the input data. In this way, deep learning solves the main bottleneck of traditional machine learning. Before the advent of deep learning, machine learning methods for object recognition problems may rely heavily on features designed by humans, possibly combined with shallow classifiers. A shallow classifier can be a two-class linear classifier, for example, in which the weighted sum of the feature vector components can be compared to a threshold to predict which class the input belongs to. Features designed by humans can be templates or kernels customized for a specific problem domain by engineers with domain expertise. In contrast, although deep learning architectures can learn to represent features similar to those that human engineers may design, they require training. In addition, deep networks can learn to represent and recognize new types of features that humans may not have considered.

[0065] A deep learning architecture can learn a hierarchy of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power in specific frequencies. The second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes for visual data or combinations of sounds for auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.

[0066] Deep learning architectures can perform particularly well when applied to problems with a natural hierarchy. For example, the classification of motorized vehicles can benefit from first learning to identify wheels, windshields, and other features. These features can be combined in different ways at higher levels to identify cars, trucks, and airplanes.

[0067] Neural networks can be designed with various connection patterns. In a feedforward network, information passes from lower layers to higher layers, where each neuron in a given layer communicates with neurons in the higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one block of input data presented sequentially to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the identification of high-level concepts can assist in discerning specific low-level features of the input.

[0068] The connections between the layers of a neural network can be fully connected or locally connected. To adjust the weights, a learning algorithm can compute the gradient vector of the weights. The gradient can indicate the amount by which the error will increase or decrease when the weights are adjusted. At the top layer, the gradient can directly correspond to the value of the weight connecting the activated neuron in the penultimate layer and the neuron in the output layer. In lower layers, the gradient can depend on the value of the weights and the error gradient computed in higher layers. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be called "backpropagation" because it involves a "backward pass" through the neural network.

[0069] In practice, the error gradient of the weights can be computed over a small number of examples, such that the computed gradient is close to the true error gradient. This approximation method can be called stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level. After learning, new images can be presented to the DCN, and the forward pass through the network can produce an output that can be regarded as an inference or prediction of the DCN.

[0070] A Deep Belief Network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. A DBN can be used to extract a hierarchical representation of a training data set. A DBN can be obtained by stacking layers of Restricted Boltzmann Machines (RBMs). RBMs are a type of artificial neural network that can learn a probability distribution over a set of inputs. Because RBMs can learn probability distributions without information about the category to which each input should be classified, RBMs are often used for unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBM of a DBN can be trained in an unsupervised manner and can be used as a feature extractor, while the top RBM can be trained in a supervised manner (on the joint distribution of inputs from the previous layer and the target class) and can be used as a classifier.

[0071] A DCN is a network of convolutional networks, configured with additional pooling and normalization layers. DCN achieves state-of-the-art performance on many tasks. DCN can be trained using supervised learning, where both the input and output targets are known for many examples and are used to modify the weights of the network by using a gradient descent method.

[0072] The DCN may be a feed-forward network. In addition, as described above, the connections from a neuron in the first layer of the DCN to a set of neurons in the next higher layer are shared across the neurons in the first layer. The feed-forward and shared connections of the DCN may be used for fast processing. For example, the computational burden of the DCN may be much smaller than that of a similarly sized neural network that includes loops or feedback connections.

[0073] The processing of each layer of the convolutional network can be thought of as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then the convolutional network trained on this input can be thought of as three-dimensional, where two spatial dimensions are along the axes of the image and the third dimension captures color information. The output of the convolutional connection can be thought of as forming a feature map in the subsequent layer, where each element in the feature map receives input from a certain range of neurons in the previous layer and from each of the multiple channels. The values ​​in the feature map can be further processed with nonlinearities (such as correction, max(0,x)). The values ​​from neighboring neurons can be further pooled, which corresponds to downsampling and can provide additional local invariance and dimensionality reduction. Normalization corresponding to whitening can also be applied by lateral inhibition between neurons in the feature map.

[0074] Figure 5 is a block diagram illustrating a DCN 550. The DCN 550 may include multiple different types of layers based on connection and weight sharing. Figure 5As shown, DCN 550 includes convolution blocks 554A and 554B. Each of the convolution blocks 554A and 554B may be configured with a convolution layer (CONV) 556, a normalization layer (LNorm) 558, and a maximum pooling layer (MAX POOL) 560.

[0075] Although only two of the convolution blocks 554A, 554B are shown, the present disclosure is not limited thereto, and any number of convolution blocks 554A, 554B may be included in the DCN 550 according to design preferences.

[0076] The convolution layer 556 may include one or more convolution filters that may be applied to the input data to generate a feature map. The normalization layer 558 may normalize the output of the convolution filter. For example, the normalization layer 558 may provide whitening or lateral suppression. The maximum pooling layer 560 may provide spatial downsampling aggregation to achieve local invariance and dimensionality reduction.

[0077] For example, a parallel filter bank of a deep convolutional network may be loaded on SOC 400 (e.g., Figure 4 ) on the CPU 402 or GPU 404 of the SOC 400 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter bank can be loaded onto the DSP 406 or ISP 416 of the SOC 400. In addition, the DCN 550 can access other processing blocks that may exist on the SOC 400, such as the sensor processor 414 and the navigation module 420 dedicated to sensors and navigation, respectively.

[0078] The DCN 550 may also include one or more fully connected layers 562 (FC1 and FC2). The DCN 550 may also include a logistic regression (LR) layer 564. Between each layer 556, 558, 560, 562, 564 of the DCN 550 are weights (not shown) to be updated. The output of each layer in the layer (e.g., 556, 558, 560, 562, 564) can be used as an input to a subsequent layer in the layer (e.g., 556, 558, 560, 562, 564) in the DCN 550 to learn a hierarchical feature representation from the input data 552 (e.g., image, audio, video, sensor data, and / or other input data) supplied at the first convolution block in the convolution block 554A. The output of the DCN 550 is a classification score 566 for the input data 552. The classification score 566 can be a set of probabilities, where each probability is a probability of the input data, including a feature from a feature set.

[0079] Federated learning is a machine learning technique that trains a federated learning model across multiple decentralized edge devices or servers that keep local data samples without sharing the data samples with a central server. Federated learning provides the benefits of privacy-preserving machine learning and continuous learning on the edge. However, the performance of federated learning is impaired when the data at the device is non-independent and identically distributed (non-IID). Data augmentation is a method to address non-IID data. Another method is zone-based federated learning.

[0080] Zone-based federated learning groups participating devices into zones that contribute to the distribution of non-IID data at the edge. However, data from devices outside the zone of interest may be relevant. If the zone boundaries are not flexible, those relevant model updates will be ignored. Various aspects of the present disclosure introduce a method for considering machine learning updates (or gradients) from devices in neighboring zones as a supplement to machine learning updates (or gradients) from devices in the zone corresponding to the machine learning model being updated. In these aspects, when updating the model weights of the local machine learning model for the zone of interest, the centralized server considers machine learning updates (or gradients) from neighboring zones as a supplement to the local machine learning updates (or gradients). The extent to which neighboring machine learning weights affect the local model for the zone of interest can be learned over time.

[0081] Figure 6 600 is a diagram illustrating an example of different zones in a federated learning system according to various aspects of the present disclosure. Figure 6 In the example 600 of FIG. 600 , each UE 620 may be an example of a device participating in federated learning. Such a device may be referred to as a participating device. Additionally, each UE 620 may be as described in reference Figure 1 and Figure 2 An example of a UE 120 is described. In some implementations, Figure 6 As shown in example 600, each UE 620 can be placed in a group 610, 612, 614 based on one or more common attributes or settings. Each group 610, 612, 614 can correspond to a specific zone. Figure 6 As shown in FIG. 6 , the first group 610 corresponds to the first zone, the second group 612 corresponds to the second zone, and the third group 614 corresponds to the third zone. In some examples, the UE 620 may be placed in more than one group 610, 612, 614 ( Figure 6 Additionally or alternatively, two or more zones may overlap ( Figure 6614). As described, the properties and settings may include, but are not limited to, geographic location, default language, or user interface theme. As an example, each group 610, 612, 614 may be based on the geographic location of the UE. In this example, the UEs 620 in the first group 610 have a common geographic location, the UEs 620 in the second group 612 have a common geographic location, and the UEs 620 in the third group 614 have a common geographic location.

[0082] Additionally, if Figure 6 As shown in , each group 610, 612, 614 can be associated with a different zone server 654, 656, 658, where each zone server 654, 656, 658 stores a different zone model 604, 606, 608. The zone server may also be referred to as a zone device. The zone model 604, 606, 606 may be an example of a machine learning model, such as a deep neural network. Each zone server 654, 656, 658 may be a different network device, such as a federated learning (FL) server. In some examples, each zone server 654, 656, 658 may be associated with a base station (such as Figure 1 and Figure 2 Base station 110) or backend server equipment (such as Figure 1 and Figure 2 As described, each zone model 604, 606, 608 can be customized based on training performed at participating devices of the corresponding group 610, 612, 614. In addition, Figure 6 , each zone model 604, 606, 608 can be associated with a global model 602 stored in a global server 652, such as a network device (e.g., a server). The global server 652 can be located at the same location as one or more of the zone servers 654, 656, 658. Alternatively, each device 652, 654, 656, 658 can be located in a different geographic location. As described, each zone model 604, 606, 608 can be customized based on training performed at the UE 620 of the associated group.

[0083] Figure 7 is a diagram illustrating an example zone network topology for zone-based federated learning according to various aspects of the present disclosure. Figure 7, the example zone network topology 700 includes two zones, namely, zone 1 704a and zone 2 704b. For the sake of brevity and simplicity of illustration, only two zones are shown, however, the zone network topology may include more than two zones. Each of the zones 704a, 704b may include multiple participating devices 710a-710f. Each of the participating devices 710a-710f may be a mobile communication device, such as, for example, a smart phone or an electric vehicle, or an Internet of Things (IoT) device. Each of the participating devices may be included in a group corresponding to a zone (e.g., 704a or 704b) based on one or more common attributes or settings. In some examples, the participating devices (e.g., 710a-710f) may be in more than one group ( Figure 7 Additionally or alternatively, two or more zones may overlap ( Figure 7 As described, the properties and settings may include, but are not limited to, geographic location, default language, or user interface theme. As an example, each zone 704a or 704b may be based on the geographic location of participating devices 710a-710f.

[0084] Each of the participating devices 710a-710f can be docked and communicated with one or more communicator edge nodes (e.g., 706, 708a, and 708b). In some aspects, the communicator edge nodes (e.g., 706, 708a, and 708b) can also serve as aggregators for a given zone. The aggregator can be configured to perform zone-level federal averaging. That is, the aggregator can receive model updates calculated at each of the participating devices (e.g., 710a-710f) in the zone (e.g., 704a or 704b), and can calculate representative values ​​of the zone, such as average values. For example, the communicator edge node 708a can also serve as an aggregator for zone 1 704a. On the other hand, the communicator edge node 708b can also be used as an aggregator for zone 2 704b. In some aspects, the communicator edge nodes (e.g., 706, 708a, and 708b) and the aggregator node can be base stations (e.g., gNode B). For example, in 5G NR and later deployments, mobile edge computing (MEC) devices may be used as aggregators (e.g., 708a, 708b) or communicators (e.g., 706, 708a, and 708b).

[0085] Each zone (e.g., 704a, 704b) may include one or more communicator edge nodes (e.g., 706) and an aggregator (e.g., 708a, 708b) that also operates as a communicator edge node. The aggregator (e.g., 708a, 708b) may receive a global model from the cloud device 702. The aggregator (e.g., 708a, 708b) may distribute the global model to each of the participating devices (e.g., 710a-710f) in the zone. Each participating device in the participating devices (e.g., 710a-710f) may be trained with the global model to generate a local model. Since each device (e.g., 710a-710f) may collect data and operate a local model, each participating device in the participating devices may be retrained (e.g., according to a loss function) to generate a local model update. Each aggregator in the aggregator (e.g., 708a, 708b) may receive a local model update from the devices (e.g., 710a to 710f) in the corresponding zone. For example, aggregator 708a may receive local model updates from devices 710a and 710b. Aggregator 708a may aggregate local model updates and compute zone model updates, for example, using a federated averaging process. Aggregator (e.g., 708a) may then provide zone model updates to each of the participating devices (e.g., 710a-710b) in the zone. Additionally, aggregator (e.g., 708a, 708b) may provide zone model updates to cloud device 702 that manages the global model. Updates may include all model weights, changed model weights, incremental values ​​of model weights, or in some cases may include the entire model.

[0086] Figure 8 is a block diagram illustrating an example of a participating device 800 including a federated learning (FL) manager 802 according to various aspects of the present disclosure. Figure 8 In the example of FIG. 8 , the participating device 800 may be an example of a UE, such as referring to Figure 1 , Figure 2 , Figure 3 and Figure 7 UE 120, 710 described. Participating device 800 can communicate with at least one network device in zone 850, which includes a federated learning zone manager 852 (only one network device with a zone manager is shown). The network device in each zone 850 can be an example of a base station 110, such as reference 1 Figure 1 , Figure 2 , Figure 3 and Figure 7 The described base stations 110, 706, 708. The network device including the federated learning zone manager 852 in the zone 850 can communicate with the network device including the zone partition keeper 872 in the cloud 870. The network device having the zone partition keeper 872 in the cloud 870 can be as described in reference Figure 7The cloud device 702 described or as referenced Figure 1 and Figure 2 The examples of the network controller 130 are described, but the network device is not limited thereto.

[0087] like Figure 8 As shown in , the participating device 800 may include multiple components, such as a local weight storage area 804, a global weight storage area 814, a model trainer 806, a model runner 816, a processed data storage area 808, a data preprocessor 810, a raw data storage area 812, an inter-process communication component 818, a data collector 822, and a local privacy protection manager 824. The various storage components 804, 806, 808, 812, 814 can be the same storage device (such as reference Figure 2 In another example, the storage components 804, 806, 808, 812, 814 may be different storage devices. An inter-process communication component 818 (such as a bus or controller / processor) may facilitate communication between the different components 804, 806, 808, 810, 812, 814, 816. The inter-process communication component 818 may be as described with reference to Figure 2 An example of the controller / processor 280 is depicted. 'Application' 820 represents an interface (e.g., an application programming interface (API)) that can be used by an application (e.g., a third party application) to communicate with the FL phone manager 802 and related components 802, 804, 806, 808, 810, 812, 814, 816, 818, 822, 824 in order to participate in federated training or to run inference with a model managed by the FL phone manager 802.

[0088] In some examples, the federated learning (FL) manager 802 controls data collection using one or more data collectors 822. Each data collector 822 can collect data from sensors ( Figure 8 In some implementations, the data collector 822 may be embedded with another data collector 822 such that the two data collectors 822 collect different types of data simultaneously. Controlling data collection via the FL manager 802 may improve resource usage, such as battery usage and / or processor usage, because the FL manager 802 may prevent multiple data collectors 822 from collecting the same data. Additionally, sensor access control may be simplified based on the FL manager 802 controlling data collection. In some examples, the FL manager 802 may dynamically (e.g., on demand) configure one or more of the following: sensor type, sampling rate, and the parameters for transferring data from the memory ( Figure 8The FL manager 802 may be configured to perform a flush (not shown) to a storage area (such as the process data storage area 808) at a time. Each model may inform the FL manager 802 of the data type and specified sampling rate required for training of the model. Based on the information provided by each model, the FL manager 802 may identify the appropriate data collector 822 to invoke and the corresponding sampling rate. In some implementations, the FL manager 802 may use one or more strategies to balance sensing accuracy (e.g., sampling rate) with resource consumption (e.g., battery usage, process load, etc.).

[0089] exist Figure 8 In the example of Figure 8 812. In some examples, the data collector 822 may buffer a certain amount of sensed data in memory before submitting the sensed data to the raw data storage area 812. The FL manager 802 may dynamically reconfigure a data dump period that defines when data is written to the raw data storage area 812. In such examples, the data dump period may be initially set by the data collector 822.

[0090] In some examples, the model may use the raw data. In other examples, the model may specify additional processing for the raw data. The additional processing may be performed by data processor 810. Figure 8 800. In some examples, the FL manager 802 may determine when to call a model-specific data processor 810. Each data processor 810 may store data in a processed data storage area 808. Data may be stored at intervals or based on new data becoming available in a raw data storage area 812. In some examples, all data is pre-processed before initiating a new local model training operation.

[0091] In some examples, data processor 810 and data collector 822 may be implemented by third-party developers. In some such examples, FL manager 802 may use inter-process communication (IPC) component 818 functionality provided by the phone's operating system to interact with third-party components.

[0092] As described, the FL manager 802 can initiate a model trainer for a given model and determine the location of data in the processing data store 808 or the raw data store 812. After training is complete, the model trainer 806 can store the newly computed weights in the local weight store 804. Additionally, the FL manager 802 can determine when the stored weights can be uploaded to the network device.

[0093] In some examples, the FL manager 802 can receive multiple models from one or more zone managers 852. That is, multiple models (e.g., federated learning models or applications) can be provided to the participating devices 800. As an example, a first application can be a text prediction model, and a second application can be a location-based advertising model. In such examples, the FL manager 802 can determine the training time for each model. In some examples, the participating device 800 can be associated with two different zone managers, where each zone server is associated with a different zone. Each zone server can send a different model. As another example, a single zone server can send two or more different zone models.

[0094] The models can be stored in the model trainer 806. The local weights of each model can be stored in the local weight store 804, and the global weights can be stored in the global weight store 814. In some implementations, the FL manager 802 can work in conjunction with one or more components 804, 806, 808, 810, 812, 814, 816 of the participating device 800 to determine the training priorities of the various models stored in the model trainer 806. In some examples, the priority of a model can be determined based on various criteria, such as but not limited to one or more of the following: the number of samples available for training a given model, the current accuracy of the model, the estimated model training time based on previous training times, and whether training can be successfully completed based on current resource availability (e.g., battery level, current system load, etc.). Additionally, the FL manager 802 can manage the local training status of the various models stored in the model trainer 806. As an example, the FL manager 802 can stop training the first model and start training the second model. In such an example, the FL manager 802 can store the local weights of the first model in the local weight store 804 to maintain the training status of the first model so that training can be resumed at a later time.

[0095] In some implementations, the FL manager 802 can determine current device resources to evaluate whether one or more models can be trained locally (e.g., on the device). It may be desirable to train the model locally to protect data privacy. Since participating devices 800 such as UEs and edge devices may have a limited number of resources, local training may still be limited. In such implementations, if the current device resources meet the resource condition and the current connectivity state meets the connection condition, the FL manager 802 can use the local privacy protection manager 824.

[0096] As described, the amount of available resources (such as available memory or processor load) can prevent the participating device 800 from training the model locally. In this example, the resource condition can be met when the amount of available resources prevents local training. That is, the amount of available resources can be less than a threshold. In some examples, when the resource condition is met, the FL manager 802 can determine the current connectivity state. The connectivity state refers to the connection state between the participating device 800 and the network device through a communication channel (such as a Wi-Fi channel or a cellular channel). In this example, the connection condition can be met if the participating device can communicate with the network device (e.g., an inter-network or intra-network device) through the communication channel. In this example, the FL manager 802 can use the network device as a proxy for training the model.

[0097] In some specific implementations, the local privacy protection manager 824 can be controlled individually by each participating device 800 to increase the speed of training while still protecting privacy. The local privacy protection manager 824 can be a network device that can receive both the model and the training data. The network device can train the model and return the trained weights and biases to the participating devices. In some examples, the local privacy protection manager 824 can delete the data corresponding to the model, weights, and biases after the training session. In addition, in some examples, the local privacy protection manager 824 may not understand the overall context of the model. Instead, the local privacy protection manager 824 may only be responsible for training the model. In addition, the global server may not know the local privacy protection manager 824. Because of the decentralized nature of training, and because the local privacy protection manager 824 is unaware of the entire context, the privacy of participating devices can be protected.

[0098] The zone partition maintainer 872 is in communication with each federated learning zone manager 852. The zone partition maintainer 872 includes a zone partition assignment module 874 that maintains an overall zone topology map. The overall zone topology map identifies the zones and the zone manager 852 for each zone.

[0099] The federated learning region manager 852 is responsible for communicating with participating devices 800 such as smart phones and performing region-level aggregation. The federated learning region manager 852 also interacts with neighboring regions to perform merge or split operations and updates the latest region partition information to the region partition holder 872. These regions can be suitable for improving the overall model accuracy.

[0100] In some aspects of the present disclosure, when enough updates have been uploaded or when the training round timer expires, the FL zone manager 852 calls the model aggregator 854 for the model. The model aggregator 854 reads the updates from the zone local model weight storage area 856, calculates the aggregate weights, and stores them in the zone global model weight storage area 858. Intermediate training states are stored in the training state storage area 860 to provide lower input / output (I / O) latency compared to other types of cloud storage in the design. This is because the FL zone manager 852 needs to access data frequently during training. Next, the model aggregator 854 transmits a notification via the new model notification service 862 to let the participating devices 800 know that a new model version is available. The zone local model utility storage area 864 and the zone partition updater 866 are used for model verification and zone management.

[0101] When a device 800 registers to participate in federated learning, the device 800 is provided with the latest zone topology map and a "zone determination function". This function accepts a set of parameters from the device and returns the zone information to which the device belongs. These parameters may include, for example, global positioning system (GPS) coordinates. This function can run offline and can be local to each device 800.

[0102] The device 800 periodically checks its zone membership and tabulates its training data based on the zone membership. When the device 800 is ready to perform local training, the device 800 communicates with the federated learning zone manager 852 for the zone to which the device belongs. Whenever the zone topology or zone determination function changes, the zone partitioning maintainer 872 updates and notifies all participating devices 800.

[0103] In other aspects, the device stores training data and parameters used by the zone membership function. For example, if an attempt is made to train a human activity recognition (HAR) model using sensor data, and if the zone partition holder provides a zone determination function that accepts GPS coordinates as parameters, the device can store the raw sensor data along with the GPS information in a sequential or time-stamped manner. When the device is ready to perform local training, the device can use the GPS data included in the data sample to determine which zone the data will be used for. These aspects differ from existing methods in that, in existing methods, the device determines zone membership when collecting data, and then tabulates the data. The previously described method involves periodic lookups of zone membership. In a second method, the device collects training data and parameters for determining zone membership. The data is partitioned to match zones at a later point in time.

[0104] Zone membership checks may be performed periodically or in an event-driven manner. According to aspects of the present disclosure, the device stores locally generated training, testing, and validation data to reflect the zone in which the data was collected. The zone determination functionality may be used in an offline mode (e.g., not connected to a network). According to aspects of the present disclosure, the device 800 maintains storage even when not connected to a network (e.g., moving from zone one to zone two). When connectivity is restored, previously stored training weights may be uploaded to the zone one manager even when the device 800 has moved to a different zone (e.g., zone two).

[0105] Fig. 9 is a timeline illustrating a zone membership check in accordance with various aspects of the present disclosure. Fig. 9 In the example of , at time t1, participating device 800 performs a periodic zone membership check using the zone determination function. At time t1, device 800 determines that the device is a member of zone one. Therefore, device 800 stores any locally generated training, testing, and validation data to reflect the data collected in zone one. At times t2 and t3, participating device 800 again performs a periodic zone membership check. At times t2 and t3, device 800 is still a member of zone one. Therefore, device 800 stores the locally generated data collected at these times (e.g., t2, t3) to reflect the data collected in zone one.

[0106] At time t4, an event-driven zone check occurs. For example, a sensor-based handover or different types of activities can trigger this event-driven zone check. An accelerometer is an example of a type of sensor that can trigger an event-driven zone check. At time t4, the device 800 determines that the device is now a member of zone two. Therefore, the data collected at this time is stored with reference to zone two. At time t5, the device 800 performs another periodic zone check. At time t5, the device 800 is still in zone two and stores data accordingly. At time t6, the periodic zone check indicates that the device 800 is now in zone three. Therefore, the device 800 stores its data with reference to zone three. Once the device 800 is ready to perform local training, the device 800 communicates with the federated learning zone manager 852 for the zone for which it has collected local data. In other words, the device 800 obtains the latest federated learning model for the zone of which the device 800 is a member, and performs local training and uploads model updates.

[0107] When compared to the standard federated learning model, zone-based federated learning has advantages. In some examples, zone-based federated learning helps to handle non-independent and identically distributed (non-IID) or class-imbalanced data in practical applications. Additionally, zone-based federated learning helps to customize model evolution. Zone-based federated learning also exploits the specific behavior of a given region or group of participants.

[0108] Identifying optimal zone boundaries for zone-based federated learning can be challenging. Aspects of the present disclosure address the problem when zone boundaries are not properly identified and / or cannot be changed. In many deployments or applications, zones may be defined based on a priori knowledge (e.g., administrative districts, counties, geometric patterns, etc.). In some cases, a simple zone-based model evolution of zone boundaries may be suboptimal.

[0109] According to aspects of the present disclosure, when updating a given zone “Z i "When the model weight is from zone Z i At least some of the clients in the neighboring zones of are included in the training process. i The adjacent area affects the area Z i The degree of model evolution can be controlled by the self-attention parameter β. i The β value of each neighboring region of the interest region can be determined based on the similarity of the neighboring regions to the interest region. For example, the neighboring group can be compared with the group of the interest region. This comparison can help determine which neighboring regions are most similar.

[0110] Fig.10 is a diagram illustrating an exemplary pseudo code for implementing zone gradient diffusion according to various aspects of the present disclosure. In the pseudo code, at line 1, zone Zi At row 2, the process spans all adjacent zones Z n Iteration. Adjacent zone Z n This is the region Z i For each neighboring region, the relationship between the gradient of the neighboring region and the gradient of the local client is determined at row 3, such as the normalized inner product, to obtain the similarity between the gradients of the region e in The present disclosure is not limited to normalized inner products, as other relations may also be useful. The symbol σ denotes the sigmoid function in row 3, while the symbol θ i t Denotes the region Z used for training round t i Model parameters. i The gradient of is expressed as And the adjacent zone Z n The gradient of is expressed as

[0111] In line 4, based on Fig.10 The similarity parameter β is determined for all neighboring regions using the formula shown in in , such as the self-attention coefficient. The similarity parameter βin determines which neighboring regions should be most influential. In line 4, exp represents an exponential function. In line 4, the similarity parameter is calculated for all neighboring regions. By using the similarity e for all neighboring regions ij The denominator in line 4 is calculated by summing the counter variable j.

[0112] At row 5, the gradient from the neighboring region is aggregated with the local gradient based on the parameter calculated for the neighboring region in row 4. In some aspects, the parameter may indicate that the neighboring region is given less weight than the local update. In other aspects, the parameter indicates that more weight should be assigned to the local update.

[0113] Zone Gradient Diffusion (ZGD) improves the zone model by aggregating contextual information derived from local gradients of neighboring zones. In zone gradient diffusion, zones do not change, but instead capture user mobility behavior changes through the diffusion of information from neighboring zones. A self-attention mechanism is applied in zone gradient diffusion to dynamically quantify the influence of each zone on its neighboring zones.

[0114] Fig.11 1 is a flow chart illustrating an example process 1100 performed, for example, by a federated learning device according to various aspects of the present disclosure. The example process 1100 is an example of a zone gradient diffusion (ZGD) technique for zone-based federated learning. The operations of the process 1100 may be implemented by the network controller 130 and / or the base station 110.

[0115] At block 1102, a network controller and / or a base station receives machine learning model updates from a plurality of clients in a federated learning system. For example, a network controller or a base station (e.g., using controller / processor 290, communication unit 294, memory 292, antenna 234, MOD / DEMOD 232, MIMO detector 236, receive processor 238, controller / processor 240, memory 242, etc.) may receive machine learning model updates.

[0116] At block 1104, the network controller and / or the base station determines a fixed local region associated with each of the plurality of clients, the fixed local region having a first fixed boundary. For example, the network controller and / or the base station (e.g., using the controller / processor 290, the memory 292, the controller / processor 240, the memory 242, etc.) may determine the fixed local region.

[0117] At block 1106, the network controller and / or the base station updates the model weights of the central machine learning model based on the local machine learning updates for the local subset of clients. The local subset corresponds to a fixed local region. For example, the network controller and / or the base station (e.g., using controller / processor 290, memory 292, controller / processor 240, memory 242, etc.) may update the model weights.

[0118] At box 1108, the network controller and / or the base station updates the model weights of the central machine learning model based on the neighboring machine learning updates for the neighboring subset of the client. The neighboring subset corresponds to a fixed neighboring area adjacent to the fixed local area. The neighboring machine learning updates have different weights than the local machine learning updates when updating the model weights. The values ​​of the different weights correspond to the similarity parameters. The fixed neighboring area has a second fixed boundary. For example, the network controller and / or the base station (e.g., using the controller / processor 290, the memory 292, the controller / processor 240, the memory 242, etc.) can update the model weights. In some aspects, the parameter is learned. The parameter can be, for example, a self-attention coefficient that normalizes the relationship between the local machine learning update and the neighboring machine learning update. For example, the relationship can be an inner product.

[0119] Example aspects

[0120] Aspect 1: A processor-implemented method, comprising: receiving machine learning model updates from multiple clients in a federated learning system; determining a fixed local region associated with each of the multiple clients, the fixed local region having a first fixed boundary; updating a model weight of a central machine learning model based on local machine learning updates for a local subset of the multiple clients, the local subset corresponding to the fixed local region; and updating the model weight of the central machine learning model based on neighboring machine learning updates for a neighboring subset of the multiple clients, the neighboring subset corresponding to a fixed neighboring region adjacent to the fixed local region, wherein when updating the model weight, the neighboring machine learning update has a different weight from the local machine learning update, the value of the different weight corresponds to a similarity parameter, and the fixed neighboring region has a second fixed boundary.

[0121] Aspect 2: The processor-implemented method according to Aspect 1 further comprises using machine learning training to learn the similarity parameter.

[0122] Aspect 3: A processor-implemented method according to Aspect 1 or 2, wherein the similarity parameter comprises a self-attention coefficient.

[0123] Aspect 4: A processor-implemented method according to any one of the preceding aspects, wherein the self-attention coefficient normalizes the relationship between the local machine learning update of the local subset and the neighboring machine learning updates of the neighboring subset.

[0124] Aspect 5: The processor-implemented method according to any of the preceding aspects, wherein the relationship comprises an inner product.

[0125] Aspect 6: A device comprising: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory, the at least one processor being configured to: receive machine learning model updates from multiple clients in a federated learning system; determine a fixed local region associated with each of the multiple clients, the fixed local region having a first fixed boundary; update a model weight of a central machine learning model based on local machine learning updates for a local subset of the multiple clients, the local subset corresponding to the fixed local region; and update the model weight of the central machine learning model based on neighboring machine learning updates for a neighboring subset of the multiple clients, the neighboring subset corresponding to a fixed neighboring region adjacent to the first region, and when updating the model weight, the neighboring machine learning update has a different weight from the local machine learning update, the value of the different weight corresponds to a similarity parameter, and the fixed neighboring region has a second fixed boundary.

[0126] Aspect 7: The apparatus according to aspect 6, wherein the at least one processor is further configured to: learn the similarity parameter using machine learning training;

[0127] Aspect 8: The apparatus according to Aspect 6 or 7, wherein the similarity parameter comprises a self-attention coefficient.

[0128] Aspect 9: An apparatus according to any one of Aspects 6 to 8, wherein the self-attention coefficient normalizes the relationship between the local machine learning update of the local subset and the neighboring machine learning updates of the neighboring subset.

[0129] Aspect 10: The apparatus of any one of aspects 6 to 9, wherein the relationship comprises an inner product.

[0130] Aspect 11: An apparatus comprising: a component for receiving machine learning model updates from multiple clients in a federated learning system; a component for determining a fixed local region associated with each of the multiple clients, the fixed local region having a first fixed boundary; a component for updating a model weight of a central machine learning model based on local machine learning updates for a local subset of the multiple clients, the local subset corresponding to the fixed local region; and a component for updating the model weight of the central machine learning model based on neighboring machine learning updates for a neighboring subset of the multiple clients, the neighboring subset corresponding to a fixed neighboring region adjacent to the first region, the neighboring machine learning update having a different weight from the local machine learning update when updating the model weight, the value of the different weight corresponding to a similarity parameter, and the fixed neighboring region having a second fixed boundary.

[0131] Aspect 12: The apparatus according to Aspect 11 further comprises a component for learning the similarity parameter using machine learning training.

[0132] Aspect 13: The apparatus according to Aspect 11 or 12, wherein the similarity parameter comprises a self-attention coefficient.

[0133] Aspect 14: An apparatus according to any one of Aspects 11 to 13, wherein the self-attention coefficient normalizes the relationship between the local machine learning update of the local subset and the neighboring machine learning updates of the neighboring subset.

[0134] Aspect 15: An apparatus according to any one of Aspects 11 to 14, wherein the relationship comprises an inner product.

[0135] Aspect 16: A non-transitory computer-readable medium having program code recorded thereon, the program code being executed by a processor and comprising: program code for receiving machine learning model updates from multiple clients in a federated learning system; program code for determining a fixed local region associated with each of the multiple clients, the fixed local region having a first fixed boundary; program code for updating model weights of a central machine learning model based on local machine learning updates for a local subset of the multiple clients, the local subset corresponding to the fixed local region; and program code for updating the model weights of the central machine learning model based on neighboring machine learning updates for a neighboring subset of the multiple clients, the neighboring subset corresponding to a fixed neighboring region adjacent to the first region, the neighboring machine learning updates having different weights from the local machine learning updates when updating the model weights, the values ​​of the different weights corresponding to a similarity parameter, the fixed neighboring region having a second fixed boundary.

[0136] Aspect 17: The non-transitory computer readable medium according to aspect 16, wherein the program code further comprises program code for learning the similarity parameter using machine learning training;

[0137] Aspect 18: The non-transitory computer-readable medium of aspect 16 or 17, wherein the similarity parameter comprises a self-attention coefficient.

[0138] Aspect 19: A non-transitory computer-readable medium according to any one of Aspects 16 to 18, wherein the self-attention coefficient normalizes the relationship between the local machine learning update of the local subset and the neighboring machine learning updates of the neighboring subset.

[0139] Aspect 20: The non-transitory computer-readable medium of any one of Aspects 16 to 19, wherein the relationship comprises an inner product.

[0140] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the aspects to the precise forms disclosed. Modifications and variations may be made in light of the above disclosure or may be acquired from practice of the various aspects.

[0141] As used, the term "component" is intended to be broadly interpreted as hardware, firmware, and / or a combination of hardware and software. As used, a processor is implemented using hardware, firmware, and / or a combination of hardware and software.

[0142] Some aspects are described in conjunction with thresholds. As used, satisfying a threshold may refer to a value being greater than a threshold, greater than or equal to a threshold, less than a threshold, less than or equal to a threshold, equal to a threshold, not equal to a threshold, etc., depending on the context.

[0143] It will be apparent that the described systems and / or methods can be implemented in various forms of hardware, firmware, and / or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the aspects. Therefore, the operation and performance of these systems and / or methods are described without reference to specific software code, and it should be understood that software and hardware used to implement these systems and / or methods can be designed based at least in part on these descriptions.

[0144] Although the specific combination of features is stated in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of various aspects. In fact, many of these features can be combined in a manner that is not specifically set forth in the claims and / or is not disclosed in the specification. Although each dependent claim listed below may directly rely only on one claim, the disclosure of various aspects includes each dependent claim and each other claim combination in the claim set. The phrase "at least one" mentioned in the list of items refers to any combination of those items, including a single member. For example, "at least one of a, b or c" is intended to cover a, b, c, ab, ac, bc and abc, and any combination with multiple identical elements (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc and ccc, or any other ordering of a, b and c).

[0145] The elements, actions or instructions used should not be interpreted as critical or essential unless explicitly described as such. In addition, as used, the articles "a" and "an" are intended to include one or more items, and they can be used interchangeably with "one or more". In addition, as used, the terms "set" and "group" are intended to include one or more items (e.g., related items, unrelated items, combinations of related items and unrelated items, etc.), and can be used interchangeably with "one or more". If only one item is intended to be referred to, the phrase "only one" or similar terms will be used. In addition, as used, the terms "having" and the like are intended to be open terms. In addition, the phrase "based on" is intended to mean "based at least in part on", unless otherwise explicitly stated.

Claims

1. A processor-implemented method, include: Receive machine learning model updates from multiple clients in a federated learning system; determining a fixed local area associated with each of the plurality of clients, the fixed local area having a first fixed boundary; updating model weights of a central machine learning model based on local machine learning updates for a local subset of the plurality of clients, the local subset corresponding to the fixed local region; as well as The model weights of the central machine learning model are updated based on neighboring machine learning updates of neighboring subsets for the multiple clients, wherein the neighboring subsets correspond to fixed neighboring areas adjacent to the fixed local area, and when updating the model weights, the neighboring machine learning updates have different weights from the local machine learning updates, and the values ​​of the different weights correspond to similarity parameters, and the fixed neighboring areas have a second fixed boundary.

2. The processor-implemented method of claim 1 , further comprising utilizing machine learning training to learn the similarity parameter.

3. The processor-implemented method of claim 1 , wherein the similarity parameter comprises a self-attention coefficient.

4. The processor-implemented method of claim 3, wherein the self-attention coefficient normalizes the relationship between the local machine learning update of the local subset and the neighboring machine learning updates of the neighboring subset. The processor-implemented method of claim 4 , wherein the relation comprises an inner product.

6. A device, include: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory, the at least one processor configured to: Receive machine learning model updates from multiple clients in a federated learning system; determining a fixed local area associated with each of the plurality of clients, the fixed local area having a first fixed boundary; updating model weights of a central machine learning model based on local machine learning updates for a local subset of the plurality of clients, the local subset corresponding to the fixed local region; as well as The model weights of the central machine learning model are updated based on neighboring machine learning updates of neighboring subsets for the multiple clients, wherein the neighboring subsets correspond to fixed neighboring areas adjacent to the fixed local area, and when updating the model weights, the neighboring machine learning updates have different weights from the local machine learning updates, and the values ​​of the different weights correspond to similarity parameters, and the fixed neighboring areas have a second fixed boundary.

7. The apparatus of claim 6, wherein the at least one processor is further configured to learn the similarity parameter using machine learning training.

8. The apparatus of claim 6, wherein the similarity parameter comprises a self-attention coefficient.

9. An apparatus according to claim 8, wherein the self-attention coefficient normalizes the relationship between the local machine learning update of the local subset and the neighboring machine learning updates of the neighboring subset.

10. The apparatus of claim 9, wherein the relationship comprises an inner product.

11. A device, include: A component for receiving machine learning model updates from multiple clients in a federated learning system; means for determining a fixed local region associated with each of the plurality of clients, the fixed local region having a first fixed boundary; means for updating model weights of a central machine learning model based on local machine learning updates for a local subset of the plurality of clients, the local subset corresponding to the fixed local region; as well as A component for updating the model weights of the central machine learning model based on neighboring machine learning updates for neighboring subsets of the multiple clients, wherein the neighboring subsets correspond to fixed neighboring areas adjacent to the fixed local area, and when updating the model weights, the neighboring machine learning updates have different weights from the local machine learning updates, and the values ​​of the different weights correspond to similarity parameters, and the fixed neighboring areas have a second fixed boundary.

12. The apparatus of claim 11, further comprising means for learning the similarity parameter using machine learning training.

13. The apparatus of claim 11, wherein the similarity parameter comprises a self-attention coefficient.

14. An apparatus according to claim 13, wherein the self-attention coefficient normalizes the relationship between the local machine learning update of the local subset and the neighboring machine learning updates of the neighboring subset. The apparatus of claim 14 , wherein the relationship comprises an inner product.

16. A non-transitory computer readable medium having a program code recorded thereon, the program code being executed by a processor and include: Program code for receiving machine learning model updates from multiple clients in a federated learning system; program code for determining a fixed local region associated with each of the plurality of clients, the fixed local region having a first fixed boundary; program code for updating model weights of a central machine learning model based on local machine learning updates for a local subset of the plurality of clients, the local subset corresponding to the fixed local region; as well as Program code for updating the model weights of the central machine learning model based on neighboring machine learning updates for a neighboring subset of the plurality of clients, the neighboring subset corresponding to a fixed neighboring zone adjacent to the fixed local area, the neighboring machine learning updates having different weights from the local machine learning updates when updating the model weights, the values ​​of the different weights corresponding to a similarity parameter, the fixed neighboring zone having a second fixed boundary.

17. The non-transitory computer readable medium of claim 16, wherein the program code further comprises program code for learning the similarity parameter using machine learning training.

18. The non-transitory computer-readable medium of claim 16, wherein the similarity parameter comprises a self-attention coefficient.

19. The non-transitory computer-readable medium of claim 18, wherein the self-attention coefficient normalizes the relationship between the local machine learning update of the local subset and the neighboring machine learning updates of the neighboring subset.

20. The non-transitory computer readable medium of claim 19, wherein the relationship comprises an inner product.