Method and apparatus for learning personalized representations through layer decoupling-based scheduling in federated learning

The personalized representation learning method through layer decoupling-based scheduling in federated learning addresses the challenges of data heterogeneity and privacy by using a shared base layer and local head layer, improving model accuracy and learning efficiency while reducing communication costs.

WO2025135736A1PCT designated stage expired Publication Date: 2025-06-26FOUND OF SOONGSIL UNIV IND COOP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/020513
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-12-17
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing personalized federated learning methods face challenges in efficiently handling data heterogeneity and protecting client privacy, particularly when applying Curriculum Learning in heterogeneous environments.

Method used

A personalized representation learning method and device that employs layer decoupling-based scheduling in federated learning, where a shared base layer captures common features across clients, and a local head layer maintains client-specific information, with gradient freezing and unfreezing strategies applied across hidden layers.

Benefits of technology

This approach improves model accuracy in heterogeneous environments, reduces communication costs, and enhances learning efficiency by simplifying the federated learning process while protecting client privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024020513_26062025_PF_FP_ABST
    Figure KR2024020513_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are a method and an apparatus for learning representations through layer decoupling-based scheduling in federated learning. According to the present invention, an apparatus for learning representations through layer decoupling-based scheduling in federated learning is provided, the apparatus comprising a processor and a memory connected to the processor, wherein the memory stores program instructions executed on the processor to: between a base layer and a head layer, share the base layer for representation learning with a plurality of clients, wherein the base layer includes a plurality of hidden layers for processing input data to determine representations; and during a process of generating a global model by receiving variables of local models from the plurality of clients, perform representation learning for every n global rounds by sequentially releasing gradient freeze among the plurality of hidden layers from the rearmost hidden layer to the foremost hidden layer, or from the foremost hidden layer to the rearmost hidden layer.
Need to check novelty before this filing date? Find Prior Art

Description

A personalized representation learning method and device using layer decoupling-based scheduling in federated learning.

[0001] The present invention relates to a personalized representation learning method and device through layer decoupling-based scheduling in federated learning.

[0002] Federated learning refers to a method of training an artificial intelligence model together with a set of other local models by sharing only the information (local variables) of the local model with the server without sharing the client's local data.

[0003] Federated learning can successfully train local models in situations where there is insufficient local data for model training, and has advantages in terms of data security because local data is not shared with the server.

[0004] Federated learning requires an appropriate personalization method due to the heterogeneity of data stored on clients.

[0005] There have been attempts to apply the Curriculum Learning method to existing personalized federated learning, and the Curriculum Learning method uses a data-based difficulty approach.

[0006] It relies on expert models to assess the difficulty of data points, which limits the classification of difficulty of client data points in a federated learning environment.

[0007] Additionally, this approach adds complexity to the federated learning process and has the disadvantage of making it difficult to evaluate difficulty in highly heterogeneous environments.

[0008] In order to solve the problems of the above-mentioned prior art, the present invention proposes a personalized representation learning method and device through layer decoupling-based scheduling in federated learning, which can protect personal information and improve learning efficiency.

[0009] In order to achieve the above object, according to one embodiment of the present invention, a personalized representation learning device through layer decoupling-based scheduling in federated learning is provided, which is connected to a plurality of clients through a network, comprising: a processor; and a memory connected to the processor, wherein the memory shares a base layer for representation learning among a base layer and a head layer with the plurality of clients, the base layer includes a plurality of hidden layers for processing input data to determine a representation, and in a process of receiving variables of local models from the plurality of clients and generating a global model, the representation learning is performed in a manner of sequentially releasing gradient freeze from a hidden layer arranged at a rear end to a hidden layer arranged at a front end among the plurality of hidden layers or from a hidden layer arranged at a front end to a hidden layer arranged at a rear end in every n global rounds.

[0010] The above program commands may proceed with representation learning in a state where, in n global rounds at a first time point, the gradients of at least some of the first hidden layers arranged in the front end among the plurality of hidden layers are frozen and the gradients of the second hidden layers arranged in the rear end are not frozen, and in n global rounds at a second time point after the first time point, the gradients of at least some of the first hidden layers are unfrozen and then representation learning is performed.

[0011] The above program commands may proceed with representation learning in a state where, in n global rounds at a first time point, the gradients of at least some of the first hidden layers arranged at the rear end among the plurality of hidden layers are frozen and the gradients of the second hidden layers arranged at the front end are not frozen, and in n global rounds at a second time point after the first time point, the gradients of at least some of the first hidden layers are unfrozen and then representation learning is performed.

[0012] The above multiple hidden layers include a first convolution layer, a second convolution layer, and a fully connected layer from the front end to the rear end, and the program instructions can perform representation learning in a state where the gradients of the first convolution layer and the second convolution layer are frozen in n global rounds at the first time point, and can perform representation learning in a state where the gradients of the first convolution layer are frozen in n global rounds at the second time point.

[0013] After all global rounds are completed, the variables of the global model are transmitted to the plurality of clients, and the plurality of clients can perform fine tuning using the head layer and the base layer according to the variables of the received global model.

[0014] The hidden layer placed in the above-mentioned front end can be defined as a shallow layer that extracts low-level features of input data, and the hidden layer placed in the rear end can be defined as a deep layer that extracts high-level features of input data.

[0015] The data stored in each of the above multiple clients may have data heterogeneity that does not satisfy the IID (Independent Identically Distributed) condition and exhibits unbalanced distribution characteristics (Non-IID).

[0016] According to another aspect of the present invention, there is provided a personalized representation learning method through layer decoupling-based scheduling in federated learning performed on a server connected to a plurality of clients through a network, the method comprising: a step of sharing a base layer for representation learning among a base layer and a head layer with the plurality of clients, the base layer including a plurality of hidden layers for processing input data to determine a representation; and a step of performing representation learning by sequentially unfreezing gradients from a hidden layer arranged at a rear end to a hidden layer arranged at a front end among the plurality of hidden layers or from a hidden layer arranged at a front end to a hidden layer arranged at a rear end in every n global rounds.

[0017] According to another aspect of the present invention, there is provided a personalized representation learning method through layer decoupling-based scheduling in federated learning on a device via a server and a network, the method comprising: sharing a base layer for representation learning among a base layer and a head layer with the plurality of clients; receiving variables of a global model generated after all global rounds have ended from the server; and performing fine tuning using the head layer and the base layer according to the variables of the received global model, wherein the base layer includes a plurality of hidden layers for processing input data to determine a representation, and the server, in the process of receiving variables of local models from the plurality of clients and generating a global model, sequentially releases gradient freezing from a hidden layer positioned at a rear end to a hidden layer positioned at a front end among the plurality of hidden layers, or from a hidden layer positioned at a front end to a hidden layer positioned at a rear end, for every n global rounds.

[0018] According to the present invention, the base portion is shared with the server to capture the characteristics of various clients, and there is an advantage in that the accuracy of the model is improved even in an environment with high data heterogeneity.

[0019] In addition, according to the present invention, there is an advantage in that the head portion remains local, thereby reducing communication costs and protecting the client's unique information.

[0020] Furthermore, according to the present invention, the efficiency of learning is increased through representation learning through layer decoupling-based scheduling in federated learning, and the process of federated learning is simplified compared to curriculum learning in existing federated learning.

[0021] FIG. 1 is a diagram illustrating the configuration of a personalized expression learning system through layer decoupling-based scheduling in federated learning according to the present embodiment.

[0022] Figures 2 and 3 are drawings for explaining the expression learning process according to the present embodiment.

[0023] Figure 4 is a drawing showing a detailed configuration of a server according to this embodiment.

[0024] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention.

[0025] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, it should be understood that the terms "comprises" or "has" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0026] In addition, it is to be understood that the components of the embodiments described with reference to each drawing are not limited to the specific embodiments, but may be implemented to be included in other embodiments within the scope in which the technical idea of ​​the present invention is maintained, and that multiple embodiments may be re-implemented as a single integrated embodiment even if a separate description is omitted.

[0027] In addition, when describing with reference to the attached drawings, identical components will be assigned identical or related reference numerals regardless of the drawing reference numbers, and redundant descriptions thereof will be omitted. When describing the present invention, if a detailed description of a related known technology is judged to unnecessarily obscure the gist of the present invention, the detailed description thereof will be omitted.

[0028]

[0029] FIG. 1 is a diagram illustrating the configuration of a personalized expression learning system through layer decoupling-based scheduling in federated learning according to the present embodiment.

[0030] As illustrated in FIG. 1, the federated learning system according to the present embodiment may include a plurality of clients (100) and a server (102) connected to the plurality of clients (100) via a network.

[0031] Here, the multiple clients may be conventional smartphones, but may also include various IoT devices without limitation.

[0032] The data stored in each of the multiple clients is data collected from the devices, dynamic environments, and spatiotemporal data used by multiple users participating in federated learning, and is assumed to have data heterogeneity that does not satisfy the IID (Independent Identically Distributed) condition and exhibits unbalanced distribution characteristics (Non-IID).

[0033] Additionally, the network may include wired and wireless Internet networks and mobile communication networks.

[0034] In this embodiment, federated learning is performed using a personalized representation learning method through layer decoupling-based scheduling.

[0035] Representation learning divides the model into a base layer and a head layer to perform model learning.

[0036] Here, the base layer may include multiple hidden layers for processing input data to determine a representation, wherein the hidden layers may include convolutional layers, pooling layers, fully connected layers, etc.

[0037] According to this embodiment, the base layer is shared by multiple clients (100) and servers (102), and the head layer is not shared by the servers and remains with multiple clients (100).

[0038] In the process of generating a global model by receiving variables of a local model from a plurality of clients (100), the server (102) according to the present embodiment performs representation learning by sequentially releasing gradient freeze from a hidden layer positioned at the rear end to a hidden layer positioned at the front end among the plurality of hidden layers, or from a hidden layer positioned at the front end to a hidden layer positioned at the rear end, for every n global rounds.

[0039] The advantage of this approach is that the base layer is shared by the server, so the characteristics of all clients can be well captured, improving model accuracy in environments with high data heterogeneity.

[0040] And since the head layer is not shared with the server but remains only for each of the multiple clients (100), communication costs are reduced and privacy is improved by not revealing client-specific information.

[0041] This embodiment proposes a novel layer-based difficulty approach that leverages the unique properties of representation learning.

[0042] For reference, curriculum learning is a method of learning from easy examples to progressively more difficult examples, similar to the way humans learn.

[0043] Traditional curriculum learning focuses on generating the difficulty of a dataset and using expert models pre-trained on the entire dataset to evaluate the difficulty of each data point.

[0044] Next, the expert model assigns a difficulty score based on the client's loss and proceeds with learning based on the data and the client's difficulty.

[0045] A drawback of curriculum learning is that it relies on expert models to assess the difficulty of data points, and in a real-world federated learning environment, the ability to divide the difficulty of client data points may be limited.

[0046] Additionally, assigning difficulty scores and implementing curricula based on client losses adds complexity to the federated learning process and makes difficulty assessment difficult in highly heterogeneous environments.

[0047] Accordingly, in this embodiment, we propose a personalized representation learning method through scheduling based on layer decoupling in federated learning.

[0048]

[0049] According to this embodiment, learning is performed by adding hidden layers as global rounds progress after decoupling of the hidden layers included in the base layer.

[0050] In addition, the layer decoupling-based representation learning scheduling according to the present embodiment is designed by considering the characteristics of the shallow hidden layer in the early stage that captures low-level features (edges, colors, textures, etc.) of input data well in a general deep learning model, and the characteristics of the deep layer in the later stage that extracts complex and abstract features and high-dimensional patterns.

[0051] This approach simplifies the federated learning process compared to traditional curriculum learning, as there is no need to evaluate the difficulty of data points, thus eliminating the need to train expert models or assign difficulty scores.

[0052] Figures 2 and 3 are drawings for explaining the expression learning process according to the present embodiment.

[0053] Figure 2 assumes a network structure in which the base layer includes multiple hidden layers for processing input data to determine an expression, and includes a first convolutional layer (Conv1), a second convolutional layer (Conv2), and a first fully connected layer (FC1) from the front to the back, and a second fully connected layer (FC2) in the head layer.

[0054] Here, the head layer freezes the gradient at all points in time, and when the final global round T ends, the gradient of the head layer is unfrozen and used for fine-tuning.

[0055] Referring to Fig. 2, the gradient of the first convolutional layer of the shear is not frozen in the 0 to 100 global rounds of the first time point, and the global model is trained while the gradient of the second convolutional layer and the first fully connected layer are frozen.

[0056] Afterwards, in the 101st to 200th global round of the second time point, the gradients of the first and second convolutional layers are not frozen, and the global model is trained while the gradients of the first fully connected layer are frozen.

[0057] Finally, the global model is trained in the 201st to final global round T of the third time point with all base layers including the first and second convolutional layers and the first fully connected layer unfrozen.

[0058] FIG. 3 is a representation learning process according to another embodiment of the present invention. Referring to FIG. 3, in the process of generating a global model by receiving variables of local models from multiple clients (100), representation learning is performed by sequentially releasing gradient freeze from hidden layers arranged at the rear end to hidden layers arranged at the front end among the multiple hidden layers for every n global rounds.

[0059] More specifically, the gradient of the first fully connected layer at the rear end is not frozen in the 0 to 100 global rounds of the first time point, and the global model is trained while the gradients of the first convolutional layer and the second convolutional layer are frozen.

[0060] Thereafter, in the 101st to 200th global round of the second time point, the gradients of the second convolutional layer and the first fully connected layer are not frozen, and the global model is trained while the gradient of the first convolutional layer is frozen.

[0061] Finally, the global model is trained in the 201st to final global round T of the third time point with all base layers including the first and second convolutional layers and the first fully connected layer unfrozen.

[0062] After all global rounds are completed, the server (102) transmits the variables of the global model to multiple clients (100), and the multiple clients (100) perform fine tuning using the head layer and the base layer according to the variables of the received global model.

[0063] When representation learning is performed through scheduling based on layer decoupling in this way, the learning efficiency is increased because there is no need to consider the difficulty of the data set.

[0064] Figure 4 is a drawing showing a detailed configuration of a server according to this embodiment.

[0065] Referring to FIG. 4, the server (102) according to the present embodiment may include a processor (400) and a memory (402).

[0066] Here, the processor (400) may include a central processing unit (CPU) capable of executing a computer program or a virtual machine, etc.

[0067] Memory (402) may include a non-volatile storage device such as a fixed hard drive or a removable storage device. Removable storage devices may include a compact flash unit, a USB memory stick, etc. Memory (402) may also include volatile memory such as various random access memories, and may be defined as a computer-readable recording medium.

[0068] In the memory (402) according to the present embodiment, program commands for performing representation learning are stored in a manner of sequentially releasing gradient freeze from a hidden layer positioned at the rear end to a hidden layer positioned at the front end among the plurality of hidden layers or from a hidden layer positioned at the front end to a hidden layer positioned at the rear end in the process of generating a global model by receiving variables of a local model from a plurality of clients for every n global rounds.

[0069] In addition, a configuration such as FIG. 3 can be configured with multiple clients (100), and through this, the base layer for representation learning among the base layer and the head layer can be shared with the multiple clients, while receiving the variables of the global model generated after all global rounds are completed from the server (102), and performing fine tuning using the head layer and the base layer according to the variables of the received global model.

[0070] The above-described embodiments of the present invention are disclosed for the purpose of illustration, and those skilled in the art with common knowledge of the present invention will be able to make various modifications, changes, and additions within the spirit and scope of the present invention, and such modifications, changes, and additions should be considered to fall within the scope of the following patent claims.

Claims

1. As a representation learning device through scheduling based on layer decoupling in federated learning connected to multiple clients and networks, processor; and Including a memory connected to the above processor, The above memory is, Among the Base layer and Head layer, the Base layer for representation learning is shared with the above multiple clients. The above base layer includes multiple hidden layers for processing input data to determine the representation, In the process of generating a global model by receiving variables of a local model from the above multiple clients, representation learning is performed by sequentially releasing gradient freeze from a hidden layer positioned at the rear to a hidden layer positioned at the front or from a hidden layer positioned at the front to a hidden layer positioned at the rear for every n global rounds. A representation learning device using layer decoupling-based scheduling in federated learning that stores program instructions executed on the above processor.

2. In paragraph 1, The above program commands are: In the nth global round at the first time point, the gradients of at least some of the first hidden layers arranged in the front end among the plurality of hidden layers are frozen, and the gradients of the remaining second hidden layers arranged in the back end are not frozen, and representation learning is performed. A representation learning device using scheduling based on layer decoupling in federated learning, which performs representation learning after releasing the gradient freeze of at least some of the hidden layers among the first hidden layers in n global rounds of the second time point after the first time point.

3. In paragraph 2, The above multiple hidden layers are, From the front end to the back end, it includes a first convolutional layer, a second convolutional layer, and a first fully connected layer, The above program commands are: In the n global rounds of the first time point, representation learning is performed while freezing the gradients of the first convolutional layer and the second convolutional layer. A representation learning device using scheduling based on layer decoupling in federated learning that performs representation learning while freezing the gradient of the first convolutional layer in n global rounds of the second time point.

4. In paragraph 1, The above program commands are: In the nth global round at the first time point, the gradients of at least some of the first hidden layers arranged in the rear end among the plurality of hidden layers are frozen, and the gradients of the remaining second hidden layers arranged in the front end are not frozen, and representation learning is performed. A representation learning device using scheduling based on layer decoupling in federated learning, which performs representation learning after releasing the gradient freeze of at least some of the hidden layers among the first hidden layers in n global rounds of the second time point after the first time point.

5. In paragraph 1, After all global rounds are completed, the variables of the global model are sent to the above multiple clients. A representation learning device using scheduling based on layer decoupling in federated learning in which the plurality of clients perform fine tuning using the head layer and the base layer according to the variables of the received global model.

6. In paragraph 1, The hidden layer placed in the above shear is defined as a shallow layer that extracts low-level features of the input data. A representation learning device through scheduling based on layer decoupling in federated learning, where the hidden layer positioned at the rear end is defined as a deep layer that extracts high-level features of input data.

7. In paragraph 1, A personalized representation learning device in federated learning, in which data stored in each of the above multiple clients does not satisfy the IID (Independent Identically Distributed) condition and has data heterogeneity exhibiting unbalanced distribution characteristics (Non-IID).

8. A personalized representation learning method in federated learning performed on a server connected to multiple clients and a network. A step of sharing a base layer for representation learning among the base layer and the head layer with the plurality of clients, wherein the base layer includes a plurality of hidden layers for processing input data to determine a representation; and A method for representation learning through scheduling based on layer decoupling in federated learning, comprising: a step of sequentially releasing gradient freeze from a hidden layer positioned at the rear to a hidden layer positioned at the front or from a hidden layer positioned at the front to a hidden layer positioned at the rear among the plurality of hidden layers for every n global rounds in the process of generating a global model by receiving variables of local models from the plurality of clients; 9. In paragraph 8, The steps for learning the above expressions are: In the nth global round of the first time point, a step of performing representation learning while freezing the gradients of at least some of the first hidden layers arranged in the front end among the plurality of hidden layers and not freezing the gradients of the second hidden layers arranged in the back end; and A method for representation learning through scheduling based on layer decoupling in federated learning, comprising a step of performing representation learning after releasing gradient freezing of at least some hidden layers among the first hidden layers in n global rounds of the second time point after the first time point.

10. In paragraph 9, The above multiple hidden layers are, From the front end to the back end, it includes a first convolutional layer, a second convolutional layer, and a fully connected layer, The steps for learning the above expressions are: A step of performing representation learning while freezing the gradients of the first convolutional layer and the second convolutional layer in the n global rounds of the first time point; and A method for representation learning through scheduling based on layer decoupling in federated learning, comprising a step of performing representation learning while freezing the gradient of the first convolutional layer in n global rounds of the second time point.

11. In paragraph 8, The steps for learning the above expressions are: In the nth global round of the first time point, a step of performing representation learning while freezing the gradients of at least some of the first hidden layers arranged in the front end among the plurality of hidden layers and not freezing the gradients of the second hidden layers arranged in the back end; and A method for representation learning through scheduling based on layer decoupling in federated learning, comprising a step of performing representation learning after releasing gradient freezing of at least some hidden layers among the first hidden layers in n global rounds of the second time point after the first time point.

12. A representation learning method through layer decoupling-based scheduling in federated learning on devices via servers and networks, A step of sharing the base layer for representation learning among the base layer and the head layer with the plurality of clients; A step of receiving variables of a global model generated after all global rounds have ended from the above server; and Including a step of performing fine tuning using the head layer and the base layer according to the variables of the received global model, The above base layer includes multiple hidden layers for processing input data to determine the representation, The above server, in the process of generating a global model by receiving variables of local models from multiple clients, sequentially releases gradient freeze from a hidden layer arranged in the rear to a hidden layer arranged in the front or from a hidden layer arranged in the front to a hidden layer arranged in the rear among the multiple hidden layers for every n global rounds, a representation learning method through scheduling based on layer decoupling in federated learning.

Citation Information

Patent Citations

  • Low-power-consumption federated learning method based on tensor neural network

    CN117077803A

  • Personalized federal learning identification method and system based on differential privacy

    CN117196012A

  • Automatic terminology recommendation device and method for big data standardization

    KR102046640B1

  • Layer-by-layer training for federated learning

    US20230316062A1