Federated learning with heterogeneous model types and architectures

By supporting the selection of heterogeneous model types and architectures in joint learning, users can independently select models based on local data, solving the problem that local models may be overfitted or underfitted, and improving the robustness and accuracy of the global model.

CN114514519BActive Publication Date: 2025-05-16TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201980101110.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-10-04
Publication Date
2025-05-16
Estimated Expiration
2039-10-04

AI Technical Summary

Technical Problem

Existing joint learning methods restrict users from choosing model types and architectures, resulting in local models that may overfit or underfit and cannot effectively handle different data distributions between users.

Method used

Allows users to use heterogeneous model types and architectures in federated learning, select the best filters for each layer by cascading and train the fully connected layer locally, combining different local models to build a global model.

Benefits of technology

Unlocks the user's freedom to choose a model type and architecture suitable for local data, reduces overfitting and underfitting problems, and can handle different data distributions, improving the robustness and accuracy of the global model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114514519B_ABST
    Figure CN114514519B_ABST
Patent Text Reader

Abstract

A method on a central node or server is provided. The method includes: receiving a first model from a first user device and receiving a second model from a second user device, wherein the first model is a neural network model type and has a first layer set, and the second model is a neural network model type and has a second layer set different from the first layer set; for each layer in the first layer set, selecting a first subset of filters from the layer in the first layer set; for each layer in the second layer set, selecting a second subset of filters from the layer in the second layer set; constructing a global model by forming a global layer set based on the first layer set and the second layer set, so that for each layer in the global layer set, the layer includes filters based on the corresponding first subset of filters and / or the corresponding second subset of filters; and forming a fully connected layer of the global model, wherein the fully connected layer is the final layer of the global layer set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments related to federated learning using heterogeneous model types and architectures are disclosed. Background Art

[0002] Over the past few years, machine learning has enabled major breakthroughs in a variety of fields, including those related to task automation and digitization (e.g., natural language processing, computer vision, speech recognition, Internet of Things (IoT)). Much of this success is based on collecting and processing large amounts of data (so-called “big data”) in the right environment. For certain applications of machine learning, this need to collect data can be incredibly invasive of privacy.

[0003] For example, as an example of such privacy-invading data collection, consider a model for speech recognition and language translation or for predicting the next word that may be typed on a mobile phone to help people type faster. In both cases, it is beneficial to train the model directly based on user data (e.g., what a specific user is saying or typing) rather than using data from other (non-personalized) sources. Doing so will allow the model to be trained based on the data distribution that is also used to make predictions. However, for various reasons, and especially because such data may be very private, there are problems with directly collecting such data. Users are not interested in sending everything they type to a server that they cannot control. Other examples of data that may be particularly sensitive to users include financial data (e.g., credit card transactions) or business or property data. For example, a telecommunications operator collects data about alarms triggered by nodes operated by the telecommunications operator (e.g., for determining false alarms and true alarms), but such telecommunications operators do not usually wish to share this data (including customer data) with others.

[0004] A recent solution to this problem is the introduction of federated learning, a new approach to machine learning where the training data never leaves the user's computer at all. Instead of sharing their data, individual users use locally available data to compute weight updates themselves. This is a way to train a model without directly checking the user's data on a centralized server. Federated learning is a collaborative form of machine learning where the training process is distributed among many users. A server coordinates everything, but the bulk of the work is not performed by a central entity, but by a federation of users.

[0005] In federated learning, after the model is initialized, some number of users may be randomly selected to improve the model. Each randomly selected user receives the current (or global) model from the server and computes a model update using the data available to them locally. All of these updates are sent back to the server, where they are averaged and weighted by the number of training examples used by the client. The server then applies the update to the model, typically by using some form of gradient descent.

[0006] Current machine learning methods require large datasets to be available. These are often created by collecting massive amounts of data from users. Federated learning is a more flexible technique that allows models to be trained without directly viewing the data. Despite using learning algorithms in a distributed manner, federated learning is very different from the way machine learning is used in data centers. Many guarantees cannot be made about the statistical distribution, and communication with users is often slow and unstable. In order to be able to perform federated learning effectively, appropriate optimization algorithms can be adapted within each user device. Summary of the invention

[0007] Federated learning is based on the following approach: building a machine learning model based on a data set distributed on multiple devices while preventing data from leaking from the multiple devices. In existing federated learning implementations, it is assumed that users try to train or update the same model type and model architecture. That is, for example, each user is training the same type of convolutional neural network (CNN) model, which has the same layers and each layer has the same filters. In this existing implementation, users do not have the freedom to choose their own personal architecture and model type. This may also lead to problems such as local model overfitting or local model underfitting, and if the model type or architecture is not suitable for some users, it may cause the global model to be suboptimal. Therefore, it is necessary to improve the existing federated learning implementation to solve these and other problems. This improvement should allow users to run their own model types and model architectures, while being able to use centralized resources (e.g., nodes or servers) to handle these different model architectures and model types, for example, by intelligently combining individual local models to form a global model.

[0008] Embodiments disclosed herein allow for the use of heterogeneous model types and architectures between users of joint learning. For example, users can choose different model types and model architectures for their own data and fit the data to these models. For example, by cascading selected filters corresponding to each layer, each user's local best working filter can be used to build a global model. The global model can also include a fully connected layer at the output of the layer built from the local model. The fully connected layer can be sent back to the individual user with the initial layer fixed, where only the fully connected layer is then trained locally for the user. The learned weights for each individual user can then be combined (e.g., averaged) to construct the fully connected layer weights of the global model.

[0009] Embodiments provided herein enable users to build their own models while still adopting a federated learning approach, which allows users to make local decisions about which model type and architecture will work best for the user's local data while benefiting from other users' input by conducting federated learning in a privacy-preserving manner. Embodiments may also reduce the overfitting and underfitting issues that may result when using a federated learning approach as discussed previously. In addition, embodiments may handle different data distributions between users, which is something that current federated learning techniques cannot do.

[0010] According to a first aspect, a method on a central node or server is provided. The method includes receiving a first model from a first user device and receiving a second model from a second user device, wherein the first model is a neural network model type and has a first layer set, and the second model is the neural network model type and has a second layer set different from the first layer set. The method also includes: for each layer in the first layer set, selecting a first subset of filters from the layer in the first layer set; and for each layer in the second layer set, selecting a second subset of filters from the layer in the second layer set. The method also includes constructing a global model by forming a global layer set based on the first layer set and the second layer set, so that for each layer in the global layer set, the layer includes filters based on the corresponding first subset of filters and / or the corresponding second subset of filters; and forming a fully connected layer of the global model, wherein the fully connected layer is the final layer of the global layer set.

[0011] In some embodiments, the method also includes sending information about the fully connected layer of the global model to one or more user devices including the first user device and the second user device; receiving one or more coefficient sets from the one or more user devices, wherein the one or more coefficient sets correspond to: the result of each of the one or more user devices using information related to the fully connected layer of the global model to train a device-specific local model; and updating the global model by averaging the one or more coefficient sets to generate a new coefficient set for the fully connected layer.

[0012] In some embodiments, selecting a first subset of filters from a layer of the first set of layers comprises: determining k best filters from the layer, wherein the first subset comprises the determined k best filters. In some embodiments, selecting a second subset of filters from a layer of the second set of layers comprises: determining k best filters from the layer, wherein the second subset comprises the determined k best filters. In some embodiments, forming a global set of layers based on the first set of layers and the second set of layers comprises: for each layer common to the first set of layers and the second set of layers, generating a corresponding layer in the global model by cascading a corresponding first subset of filters and a corresponding second subset of filters; for each layer unique to the first set of layers, generating a corresponding layer in the global model by using a corresponding first subset of filters; and for each layer unique to the second set of layers, generating a corresponding layer in the global model by using a corresponding second subset of filters.

[0013] In some embodiments, the method further includes: instructing one or more of the first user device and the second user device to extract their corresponding local models as the neural network model type.

[0014] According to a second aspect, a method for utilizing federated learning on a user device is provided, the federated learning having heterogeneous model types and / or architectures. The method comprises: extracting a local model as a first extracted model, wherein the local model is a first model type, and the first extracted model is a second model type different from the first model type; sending the first extracted model to a server; receiving a global model from the server, wherein the global model is the second model type; and updating the local model based on the global model.

[0015] In some embodiments, the method further comprises updating the local model based on new data received at the user device; extracting the updated local model as a second extraction model, wherein the second extraction model is of the second model type; and sending a weighted average of the second extraction model and the first extraction model to the server. In some embodiments, the weighted average of the second extraction model and the first extraction model is given by W1+αW2, where W1 represents the first extraction model, W2 represents the second extraction model, and 0<α<1.

[0016] In some embodiments, the method further comprises determining coefficients of a final layer of the global model based on the local data; and sending the coefficients to a central node or server.

[0017] According to a third aspect, a central node or server is provided. The central node or server comprises: a memory; and a processor coupled to the memory. The processor is configured to: receive a first model from a first user device and receive a second model from a second user device, wherein the first model is a neural network model type and has a first layer set, and the second model is the neural network model type and has a second layer set different from the first layer set; for each layer in the first layer set, select a first subset of filters from the layer in the first layer set; for each layer in the second layer set, select a second subset of filters from the layer in the second layer set; construct a global model by forming a global layer set based on the first layer set and the second layer set, so that for each layer in the global layer set, the layer includes filters based on the corresponding first subset of filters and / or the corresponding second subset of filters; and form a fully connected layer of the global model, wherein the fully connected layer is the final layer of the global layer set.

[0018] According to a fourth aspect, a user device is provided. The user device includes a memory; and a processor coupled to the memory. The processor is configured to: extract a local model as a first extraction model, wherein the local model is a first model type, and the first extraction model is a second model type different from the first model type; send the first extraction model to a server; receive a global model from the server, wherein the global model is the second model type; and update the local model based on the global model.

[0019] According to a fifth aspect, there is provided a computer program comprising instructions which, when executed by a processing circuit, cause the processing circuit to perform the method of any one of the embodiments of the first aspect or the second aspect.

[0020] According to a sixth aspect, there is provided a carrier comprising the computer program of the fifth aspect, wherein the carrier is one of an electric signal, an optical signal, a radio signal and a computer-readable storage medium. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate various embodiments.

[0022] Figure 1 A federated learning system according to an embodiment is shown.

[0023] Figure 2 A model according to an embodiment is shown.

[0024] Figure 3 A message diagram is shown according to an embodiment.

[0025] Figure 4 An extraction according to an embodiment is shown.

[0026] Figure 5 A message diagram is shown according to an embodiment.

[0027] Figure 6 is a flow chart according to an embodiment.

[0028] Figure 7 is a flow chart according to an embodiment.

[0029] Figure 8 is a block diagram of an apparatus according to an embodiment.

[0030] Fig. 9 is a block diagram of an apparatus according to an embodiment. DETAILED DESCRIPTION

[0031] Figure 1 A joint learning system 100 according to an embodiment is shown. As shown, a central node or server 102 communicates with one or more users 104. Optionally, users 104 can communicate with each other using any of a variety of network topologies and / or network communication systems. For example, users 104 may include user devices such as smart phones, tablet computers, laptop computers, personal computers, and may also be communicatively coupled via a public network such as the Internet (e.g., via WiFi) or a communication network (e.g., LTE or 5G). Although a central node or server 102 is shown, the functions of the central node or server 102 may be distributed across multiple nodes and / or servers and may be shared between one or more users 104.

[0032] As described in the embodiments of this article, joint learning can involve one or more rounds, wherein the global model is iteratively trained in each round.Users 104 can register with a central node or server to indicate that they are willing to participate in the joint learning of the global model, and can do so continuously or rollingly.At the time of registration (and possibly at any time thereafter), the central node or server 102 can select a model type and / or a model architecture for local user training.Alternatively or additionally, the central node or server 102 can allow each user 104 to select a model type and / or a model architecture for himself.The central node or server 102 can send an initial model to the user 104.For example, the central node or server 102 can send a global model (e.g., newly initialized or partially trained by the first few rounds of joint learning) to the user.Users 104 can train their personal models locally with their own data.The results of this local training can then be reported back to the central node or server 102, which can collect the results and update the global model.This process can be repeated iteratively. Furthermore, in each round of training of the global model, the central node or server 102 may select a subset (eg, a random subset) of all registered users 104 to participate in the training round.

[0033] Embodiments provide a new architectural framework in which users 104 can choose their own architectural models when training their systems. Typically, architectural frameworks establish common practices for creating, interpreting, analyzing, and using architectural descriptions within an application domain or stakeholder community. In a typical federated learning system, each user 104 has the same model type and architecture, so it is relatively simple to combine the model inputs from each user 104 to form a global model. However, allowing users 104 to have heterogeneous model types and architectures raises the question of how the central node or server 102 that maintains the global model can resolve this heterogeneity.

[0034] In some embodiments, each individual user 104 may have a specific type of neural network (e.g., CNN) as a local model. The specific model architecture of the neural network is not constrained, and different users 104 may have different model architectures. For example, the neural network architecture may refer to the layered arrangement of neurons and the connection pattern, activation function, and learning method between the layers. With specific reference to CNN, the model architecture may refer to the specific layers of the CNN and the specific filters associated with each layer. In other words, in some embodiments, different users 104 may each train a local CNN type model, but between different users 104, the local CNN model may have different layers and / or filters. Typical joint learning systems cannot handle this situation. Therefore, some modifications to joint learning are required. Specifically, in some embodiments, the central node or server 102 generates a global model by intelligently combining different local models. By adopting this process, the central node or server 102 is able to adopt joint learning on different model architectures. Allowing the model architecture to be free from the constraints of a fixed model type can be referred to as a "same model type, different model architecture" approach.

[0035] In some embodiments, each individual user 104 may have any type of model as a local model and any architecture of the model type selected by the user 104. That is, the model type is not limited to neural networks, but may also include random forest type models, decision trees, etc. The user 104 may train the local model in a manner suitable for the specific model. Before sharing the model update with the central node or server 102, as part of the federated learning method, the user 104 converts the local model to a common model type and (in some embodiments) a common architecture. The conversion process may take the form of model extraction, as disclosed herein for some embodiments. If the conversion is to a common model type and model architecture, the central node or server 102 may essentially apply typical federated learning. If the conversion is to a common model type (e.g., a neural network type model), but not to a common model architecture, the central node or server 102 may adopt the "same model type, different model architecture" approach described for certain embodiments. Allowing both the model type and the model architecture to be unconstrained may be referred to as a "different model type, different model architecture" approach.

[0036] “Same model type, different model architecture”

[0037] As described herein, different users 104 may have local models with different model architectures between local models but sharing a common model type. Specifically, it is assumed herein that the shared model type is a neural network model type. An example of this is a CNN model type. In this case, the goal is to combine different models (e.g., different CNN models) to intelligently form a global model. Different local CNN models may have different filter sizes and different numbers of layers. More generally (e.g., when other types of neural network architectures are used), different layers may include considerations of the neuronal structure of the layers, such as different layers may have neurons with different weights, unlike users having different layers or having layers with different filters (as discussed with respect to CNNs).

[0038] Figure 2 Models according to an embodiment are shown. As shown, local models 202, 204, and 206 are each CNN model types, but have different architectures. For example, CNN model 202 includes a first layer 210 with a filter set 211. CNN model 204 includes a first layer 220 with a filter set 221 and a second layer 222 with a filter set 223. CNN model 206 includes a first layer 230 with a filter set 231, a second layer 232 with a filter set 233, and a third layer 234 with a filter set 235. Different local models 202, 204, and 206 can be combined to form a global model 208. Global CNN model 208 includes a first layer 240 with a filter set 241, a second layer 242 with a filter set 243, and a third layer 244 with a filter set 245.

[0039] In some embodiments, some aspects of the model architecture can be shared between users 104 (e.g., using the same first layer, or using common filter types). It is also possible that two or more users 104 can adopt the same architecture as a whole. However, in general, it is expected that different users 104 may choose different model architectures to optimize local performance. Therefore, although each of the models 202, 204, 206 has a first layer L1, the first layer L1 of each of the models 202, 204, 206 can be constructed differently, for example by having different filter sets 211, 221, 231.

[0040] Users 104 employing each of the local models 202, 204, and 206 may train their personal models locally, for example, using local datasets (e.g., D1, D2, D3). Typically, the datasets will contain similar types of data, for example, for training a classifier, each dataset may include the same categories, although the representation of each category may differ between datasets.

[0041] Then, build (or update) the global model based on the different local models. The central node or server 102 can be responsible for some or all of the functions associated with building the global model. Individual users 104 (e.g., user devices) or other entities can also perform certain steps and report the results of these steps to the central node or server 102.

[0042] Typically, a global model can be constructed by cascading the filters in each layer of each local model. In some embodiments, a subset of filters in each layer can be used instead, for example by selecting the k best filters for each layer. The value of k (e.g., k=2) can vary from one local model to another, and can vary from one layer to another in the local model. In some embodiments, the central node or server 102 can signal the value of k that each user 104 should use. In some embodiments, two best filters (k=2) can be selected from each layer of each local model, while in other embodiments, different values ​​of k (e.g., k=1 or k>2) can be selected. In some embodiments, k can be selected to reduce the total number of filters in the layer by a relative amount (e.g., selecting the first third of the filters). The selection of the best filter can use any suitable technique to determine the best working filter. For example, the PCT application entitled “Understanding Deep Learning Models” with application number PCT / IN2019 / 050455 describes some such techniques that can be used. Selecting a subset of filters in this way can help reduce the computational load while maintaining a high accuracy. In some embodiments, the central node or server 102 may perform the selection; in some embodiments, the user 104 or other entity may perform the selection and report the results to the central node or server 102.

[0043] The global model 208 will be used to illustrate the process. Each of the local models 202, 204, and 206 includes a first layer L1. Therefore, the global model 208 also includes a first layer L1, and the filter 241 of L1 of the global model 208 includes the filters 211, 221, 231 (or a subset of filters) of each model in the local models 202, 204, and 206, which are cascaded together. Only the local models 204 and 206 include a second layer L2. Therefore, the global model 208 also includes a second layer L2, and the filter 242 of L2 of the global model 208 includes the filters 222, 232 (or a subset of filters) of each model in the local models 204 and 206, which are cascaded together. Only the local model 206 includes a third layer L3. Therefore, the global model 208 also includes a third layer L3, and the filters 245 of L3 of the global model 208 include the filters 235 (or a subset of filters) of the local model 206.

[0044] In other words, if N(M i ) represents the local model M i The number of layers, then the global model will be constructed to have at least max(N(M i )) layers, where the max operator is used on all local models M from which the global model is being built (or updated). i For a given layer L of the global model j , layer L j Include Filters where the range of index i covers different local models with layer j, and Fi refers to a specific local model M i The filter (or filter subset) of the jth layer. indicates cascade, and Where the set I = {i}.

[0045] After cascading the local models, the global model can be further constructed by adding a dense layer (e.g., a fully connected layer) to the model as the final layer.

[0046] Once the global model is thus constructed (or updated), equations for training the model can be generated. These equations can be sent to different users 104 who can each train the final dense layer (e.g., by keeping other local filters unchanged). Users 104 who have trained the final dense layer locally can then report the model coefficients of their local dense layer to the central node or server 102. Ultimately, the global model can combine such coefficients from different users 104 who reported model coefficients to form a global model. For example, combining the model coefficients can include averaging the coefficients, including by using a weighted average, such as weighted by the amount of local data trained by each user 104.

[0047] In an embodiment, a global model built in this way will be robust and contain features learned from different local models. Such a global model can work well as a classifier, for example. An advantage of this embodiment is that the global model can be updated based on only a single user 104 (in addition to updating based on input from multiple users 104). In this single-user update case, only the weights of the last layer can be adjusted by keeping everything else fixed.

[0048] Figure 3 A message diagram according to an embodiment is shown. As shown, users 104 (e.g., first user 302 and second user 304) work with a central node or server 102 to update a global model. The first user 302 and the second user 304 each train their respective local models at 310 and 314, and each report their local models to the central node or server 102 at 312 and 316. The training and reporting of the model can be performed simultaneously, or can be staggered to a certain extent. Before continuing, the central node or server 102 can wait until it receives a model report from each user 104 that it expects to report, or it can wait until a threshold number of such model reports are received, or it can wait for a certain time period, or any combination. After having received the model report, the central node or server 102 can build or update the global model (e.g., as described above, such as by connecting filters or filter subsets of different local models at each layer, and adding a dense fully connected layer as the final layer), and form the equations required for the dense layer of the training global model. The central node or server 102 then reports the dense layer equation to the first user 302 and the second user 304 at 320 and 322. Next, the first user 302 and the second user 304 use their local models to train the dense layer at 324 and 328, and report the coefficients of the dense layer equation they have trained at 326 and 330 back to the central node or server 102. With this information, the central node or server 102 can then update the global model by updating the dense layer based on the coefficients from the local users 104.

[0049] “Different model types, different model architectures”

[0050] As described herein, different users can have local models with different model types and different model architectures. The problem that this approach addresses is that the unconstrained nature of both model type and model architecture between different local models makes combining different local models difficult to solve, as there can be significant differences between the available model types, such that training applied to one model type may not make sense for training applied to another model type. For example, a user may fit different models such as a random forest type model, a decision tree, etc.

[0051] To address this problem, embodiments convert local models to common model types, and in some embodiments also to common model architectures. One way to convert models is to use a model extraction method. Model extraction can convert any model (e.g., a complex model trained on a large amount of data) into a smaller, simpler model. The idea is to train a simpler model on the output of a complex model rather than the original output. This can convert features learned on a complex model to a simpler model. In this way, any complex model can be converted to a simpler model by retaining features.

[0052] Figure 4 Extraction according to an embodiment is shown. There are two models for extraction, a local model 402 (also referred to as a "teacher" model) and an extraction model 404 (also referred to as a "student" model). Typically, the teacher model is complex and trained using a GPU or another device with similar processing resources, while the student model is trained on a device with less powerful computing resources. This is not required, however, because the "student" model is easier to train than the original "teacher" model, it can be trained using fewer processing resources. In order to maintain the knowledge of the "teacher" model, the "student" model is trained on the probabilities of the "teacher" model's predictions. The local model 402 and the extraction model 404 can be different model types and / or model architectures.

[0053] In some embodiments, one or more individual users 104 with their own personal models, which may have different model types and model architectures, may convert their local models (e.g., by extraction) to an extracted model with a specified model type and model architecture. For example, the central node or server 102 may indicate to each user what model type and model architecture the user 104 should extract the model to. The model type will be common to each user 104, but the model architecture may be different in some embodiments.

[0054] The extracted local models may then be sent to a central node or server 102, where they are combined to build (or update) a global model. The central node or server 102 may then send the global model to one or more of the users 104. In response, the users 104 who receive the updated global model may update their own personal local models based on the global model.

[0055] In some embodiments, the extraction model sent to the central node or server 102 can be based on a previous extraction model. Assume that the user 104 has previously sent (e.g., in a recent round of federated learning) a first extraction model that represents an extraction of the local model of the user 104. The user 104 can then update the local model based on the new data received at the user 104, and can extract a second extraction model based on the updated local model. The user 104 can then take a weighted average of the first extraction model and the second extraction model (e.g., W1+αW2, where W1 represents the first extraction model, W2 represents the second extraction model, and 0<α<1), and send the weighted average of the first extraction model and the second extraction model to the central node or server 102. The central node or server 102 can then use the weighted average to update the global model.

[0056] Figure 5 A message diagram according to an embodiment is shown. As shown, users 104 (e.g., first user 302 and second user 304) work with a central node or server 102 to update a global model. The first user 302 and the second user 304 each extract their respective local models at 510 and 514, and each report their extracted models to the central node or server 102 at 512 and 516. The training and reporting of the model can be performed simultaneously, or can be staggered to a certain extent. Before continuing, the central node or server 102 can wait until it receives a model report from each user 104 that it expects to report, or it can wait until such model reports of a threshold number are received, or it can wait for a certain time period, or any combination. After having received the model report, the central node or server 102 can build or update a global model 318 (e.g., as described in the disclosed embodiment). Then the central node or server 102 reports the global model to the first user 302 and the second user 304 at 520 and 522. Next, the first user 302 and the second user 304 then update their respective local models based on the global model at 524 and 526 (eg, as described in the disclosed embodiments).

[0057] Returning to the example where each user 102 has a different model architecture for the same CNN model type, a mathematical formula related to the proposed embodiment is provided. For a given CNN, the output of each filter can be expressed as

[0058]

[0059] This is valid for N filters, where the size of the input data (in[k]) is M, the size of the filter (c) is P, and the stride is 1. That is, in[k] represents the kth element of the input (of size M) to the filter, and c[j] is the jth element of the filter (of size P). Moreover, for the purpose of illustration, only one layer is considered in this CNN model. The above representation ensures the dot product between the input data and the filter coefficients. From this representation, the filter coefficients c can be learned by using backpropagation. Typically, among these filters, only a few (e.g., two or three) filters will work well. Therefore, the above equation can be simplified to only a subset N of the filters that work well S (N S ≤ N). As mentioned above, these filters can be obtained (ie, work well compared to other filters) by various methods.

[0060] As discussed in this article, a global model can then be constructed that takes the filters of each of the models of different users and concatenates them for each layer. The global model also includes a fully connected dense layer as the final layer. For a fully connected layer with L nodes (or neurons), the mathematical formula for this layer can be expressed as:

[0061]

[0062] Where cm represents one of the filters in the subset of the best working filters, W is the set of weights of the final layer, b is the bias, and g(.) is the activation function of the final layer. The input to the fully connected layer will be flattened before being passed to this layer. This equation is sent to each of the users to calculate the weights using conventional back-propagation techniques. Assume that the weights learned by different users are W1, W2, ..., W U , where U is the number of users in the federated learning method, and the final layer weights of the global model can be determined by averaging, such as

[0063]

[0064] Example:

[0065] The following examples are prepared to evaluate the performance of the embodiment. An alarm dataset corresponding to three telecom operators is collected. The three telecom operators correspond to three different users. The alarm datasets have the same features and have different patterns. The goal is to classify alarms into true alarms and false alarms based on the features.

[0066] Users can choose their own models. In this example, each user can choose a specific architecture for the CNN model type. That is, each user can choose a different number of layers and different filters in each layer than other users.

[0067] For this example, operator 1 (the first user) chooses to fit a three-layer CNN with 32 filters in the first layer, 64 filters in the second layer, and 32 filters in the final layer. Similarly, operator 2 (the second user) chooses to fit a two-layer CNN with 32 filters in the first layer and 16 filters in the second layer. Finally, operator 3 (the third user) chooses to fit a five-layer CNN with 32 filters in each of the first four layers and 8 filters in the fifth layer. These models are selected based on the nature of the data available to each operator, and the models can be selected based on the current round of joint learning.

[0068] The global model is constructed as follows. The number of layers in the global model includes the maximum number of layers that different local models have, which is 5 layers here. The first two filters in each layer of each local model are identified, and the global model is constructed with two filters in each layer of each local model. Specifically, the first layer of the global model contains 6 filters (from the first layer of each local model), the second layer contains 6 filters (from the second layer of each local model), the third layer contains two filters from the first model and two filters from the third model, the fourth layer contains two filters from the fourth layer of the third model, and the fifth layer contains two filters from the fifth layer of the third model. Next, a dense fully connected layer is constructed as the final layer of the global model. The dense layer has 10 nodes (neurons). Once constructed, the global model is sent to the user to train the last layer, and the training results (coefficients) of each local model are collected. These coefficients are then averaged to obtain the last layer of the global model.

[0069] Applying this to the three data sets of telecom operators, the accuracy obtained for the local model is 82%, 88%, and 75%. Once the global model is built, the accuracy obtained at the local model is improved to 86%, 94%, and 80%. It can be seen from the examples that the joint learning model of the disclosed embodiment is good and can produce a better model when compared with the local model.

[0070] Figure 6A flow chart according to an embodiment is shown. Process 600 is a method performed by a central node or a server. Process 600 may start from step s602.

[0071] Step s602 includes: receiving a first model from a first user device and receiving a second model from a second user device, wherein the first model is a neural network model type and has a first layer set, and the second model is the neural network model type and has a second layer set different from the first layer set.

[0072] Step s604 comprises: for each layer in the first set of layers, selecting a first subset of filters from that layer in the first set of layers.

[0073] Step s606 comprises: for each layer in the second set of layers, selecting a second subset of filters from that layer in the second set of layers.

[0074] Step s608 comprises constructing a global model by forming a global layer set based on the first layer set and the second layer set, such that for each layer in the global layer set, the layer comprises filters based on the corresponding first subset of filters and / or the corresponding second subset of filters.

[0075] Step s610 includes forming a fully connected layer of the global model, wherein the fully connected layer is the final layer of the global layer set.

[0076] In some embodiments, the method may also include: sending information about a fully connected layer of a global model to one or more user devices including a first user device and a second user device; receiving one or more coefficient sets from the one or more user devices, wherein the one or more coefficient sets correspond to: the result of each of the one or more user devices using information related to the fully connected layer of the global model to train a device-specific local model; and updating the global model by averaging the one or more coefficient sets to generate a new coefficient set for the fully connected layer.

[0077] In some embodiments, selecting a first subset of filters from a layer of a first set of layers comprises: determining k best filters from the layer, wherein the first subset comprises the determined k best filters. In some embodiments, selecting a second subset of filters from a layer of a second set of layers comprises: determining k best filters from the layer, wherein the second subset comprises the determined k best filters. In some embodiments, forming a global set of layers based on the first set of layers and the second set of layers comprises: for each layer common to the first set of layers and the second set of layers, generating a corresponding layer in a global model by cascading a corresponding first subset of filters and a corresponding second subset of filters; for each layer unique to the first set of layers, generating a corresponding layer in the global model by using a corresponding first subset of filters; and for each layer unique to the second set of layers, generating a corresponding layer in the global model by using a corresponding second subset of filters.

[0078] In some embodiments, the method may further include: instructing one or more of the first user device and the second user device to extract their corresponding local models as the neural network model type.

[0079] Figure 7 A flow chart according to an embodiment is shown. Process 700 is a method performed by a user 104 (eg, a user device). Process 700 may start at step s702.

[0080] Step s702 comprises extracting the local model into a first extracted model, wherein the local model is of a first model type, and the first extracted model is of a second model type different from the first model type.

[0081] Step s704 includes sending the first extraction model to the server.

[0082] Step s706 comprises receiving a global model from the server, wherein the global model is of the second model type.

[0083] Step s708 comprises updating the local model based on the global model.

[0084] In some embodiments, the method may further include updating the local model based on new data received at the user device; extracting the updated local model as a second extraction model, wherein the second extraction model is of the second model type; and sending a weighted average of the second extraction model and the first extraction model to the server. In some embodiments, the weighted average of the second extraction model and the first extraction model is given by W1+αW2, wherein W1 represents the first extraction model, W2 represents the second extraction model, and 0<α<1.

[0085] In some embodiments, the method may further include determining coefficients of a final layer of the global model based on the local data; and sending the coefficients to a central node or server.

[0086] Figure 8 8 is a block diagram of an apparatus 800 (eg, a user 102 and / or a central node or server 104) according to some embodiments. Figure 8 As shown, the apparatus may include: a processing circuit (PC) 802, which may include one or more processors (P) 855 (e.g., a general-purpose microprocessor and / or one or more other processors, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc.); a network interface 848 including a transmitter (Tx) 845 and a receiver (Rx) 847, for enabling the apparatus to send data to and receive data from other nodes connected to a network 810 (e.g., an Internet Protocol (IP) network), to which the network interface 848 is connected; and a local storage unit (also known as a "data storage system") 808, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In an embodiment where the PC 802 includes a programmable processor, a computer program product (CPP) 841 may be provided. The CPP 841 includes a computer-readable medium (CRM) 842 storing a computer program (CP) 843, which includes a computer-readable instruction (CRI) 844. The CRM 842 can be a non-transitory computer-readable medium, such as a magnetic medium (e.g., a hard disk), an optical medium, a memory device (e.g., a random access memory, a flash memory), etc. In some embodiments, the CRI 844 of the computer program 843 is configured so that when executed by the PC 802, the CRI causes the device to perform the steps described herein (e.g., the steps described herein with reference to the flowchart). In other embodiments, the device can be configured to perform the steps described herein without the need for code. That is, for example, the PC 802 can consist of only one or more ASICs. Therefore, the features of the embodiments described herein can be implemented in hardware and / or software.

[0087] Fig. 9 800 according to some other embodiments. The device 800 includes one or more modules 900, each of which is implemented in software. The module 900 provides the functions of the device 800 described herein (e.g., Figure 6 to Figure 7 steps).

[0088] Although various embodiments of the present disclosure are described herein, it should be understood that they are presented only by way of example and not limitation. Therefore, the breadth and scope of the present disclosure should not be limited by any of the above exemplary embodiments. In addition, any combination of the above elements with all possible variations thereof is included in the present disclosure, unless otherwise indicated or otherwise clearly conflicting with the context.

[0089] Additionally, although the processes described above and shown in the accompanying drawings are shown as a series of steps, this is for illustrative purposes only. Therefore, it is contemplated that some steps may be added, some steps may be omitted, the order of steps may be rearranged, and some steps may be performed in parallel.

Claims

1. A method performed on a central node or server in a communication network, the method comprising: Receiving a first model from a first user device and receiving a second model from a second user device, wherein the first model is a neural network model type and has a first set of layers, and the second model is a neural network model type and has a second set of layers different from the first set of layers; for each layer in the first set of layers, selecting a first subset of filters from that layer in the first set of layers; for each layer in the second set of layers, selecting a second subset of filters from that layer in the second set of layers; constructing a global model by forming a global layer set based on the first layer set and the second layer set, such that for each layer in the global layer set, the layer includes a filter based on the corresponding first subset of filters and / or the corresponding second subset of filters; forming a fully connected layer of the global model, wherein the fully connected layer is the final layer of the global layer set; as well as The global model is sent to one or more user equipment including the first user equipment and the second user equipment.

2. The method according to claim 1, wherein: Sending the global model to one or more user devices including the first user device and the second user device includes sending information about the fully connected layer of the global model to the one or more user devices, the method further comprising: receiving one or more coefficient sets from the one or more user devices, wherein the one or more coefficient sets correspond to: results of each of the one or more user devices training a device-specific local model using information related to the fully connected layer of the global model; as well as The global model is updated by averaging the one or more coefficient sets to produce a new set of coefficients for the fully connected layer.

3. The method according to any one of claims 1 to 2, wherein: Selecting a first subset of filters from a layer of the first set of layers includes determining k best filters from the layer, wherein the first subset includes the determined k best filters.

4. The method according to any one of claims 1 to 2, wherein: Selecting a second subset of filters from a layer of the second set of layers includes determining k best filters from the layer, wherein the second subset includes the determined k best filters.

5. The method according to any one of claims 1 to 2, wherein: Forming a global layer set based on the first layer set and the second layer set includes: For each layer common to the first set of layers and the second set of layers, generating a corresponding layer in the global model by cascading a corresponding first subset of filters and a corresponding second subset of filters; For each layer unique to the first set of layers, generating a corresponding layer in the global model by using a corresponding first subset of filters; and For each layer unique to the second set of layers, a corresponding layer is generated in the global model by using a corresponding second subset of filters.

6. The method according to any one of claims 1 to 2, further comprising: Instruct one or more of the first user device and the second user device to extract their corresponding local models as the neural network model type.

7. A method for utilizing federated learning performed on a user device in a communication network, the federated learning having heterogeneous model types and / or architectures, the method comprising: extracting a local model as a first extracted model, wherein the local model is of a first model type and the first extracted model is of a second model type different from the first model type; Sending the first extraction model to a server; receiving a global model from the server, wherein the global model is constructed based on at least the first extraction model, and the global model is of the second model type; as well as The local model is updated based on the global model.

8. The method according to claim 7, further comprising: updating the local model based on new data received at the user device; extracting the updated local model as a second extracted model, wherein the second extracted model is of the second model type; A weighted average of the second extraction model and the first extraction model is sent to the server.

9. The method according to claim 8, wherein: The weighted average of the second extraction model and the first extraction model is given by W1+αW2, where W1 represents the first extraction model, W2 represents the second extraction model, and 0<α<1.

10. The method according to any one of claims 7 to 9, further comprising: determining coefficients of a final layer of the global model based on the local data; as well as The coefficients are sent to a central node or server.

11. A central node or server in a communication network, comprising: Memory; as well as a processor, coupled to the memory, wherein the processor is configured to: Receiving a first model from a first user device and receiving a second model from a second user device, wherein the first model is a neural network model type and has a first set of layers, and the second model is a neural network model type and has a second set of layers different from the first set of layers; for each layer in the first set of layers, selecting a first subset of filters from that layer in the first set of layers; for each layer in the second set of layers, selecting a second subset of filters from that layer in the second set of layers; constructing a global model by forming a global layer set based on the first layer set and the second layer set, such that for each layer in the global layer set, the layer includes a filter based on the corresponding first subset of filters and / or the corresponding second subset of filters; forming a fully connected layer of the global model, wherein the fully connected layer is the final layer of the global layer set; as well as The global model is sent to one or more user equipment including the first user equipment and the second user equipment.

12. The central node or server according to claim 11, wherein: The processor is further configured to: Sending the global model to the one or more user devices including the first user device and the second user device by sending information about the fully connected layer of the global model to the one or more user devices; receiving one or more coefficient sets from the one or more user devices, wherein the one or more coefficient sets correspond to: results of each of the one or more user devices training a device-specific local model using information related to the fully connected layer of the global model; as well as The global model is updated by averaging the one or more coefficient sets to produce a new set of coefficients for the fully connected layer.

13. A central node or server according to any one of claims 11 to 12, wherein: Selecting a first subset of filters from a layer of the first set of layers includes determining k best filters from the layer, wherein the first subset includes the determined k best filters.

14. The central node or server according to any one of claims 11 to 12, wherein: Selecting a second subset of filters from a layer of the second set of layers includes determining k best filters from the layer, wherein the second subset includes the determined k best filters.

15. The central node or server according to any one of claims 11 to 12, wherein: Forming a global layer set based on the first layer set and the second layer set includes: For each layer common to the first set of layers and the second set of layers, generating a corresponding layer in the global model by cascading a corresponding first subset of filters and a corresponding second subset of filters; For each layer unique to the first set of layers, generating a corresponding layer in the global model by using a corresponding first subset of filters; and For each layer unique to the second set of layers, a corresponding layer is generated in the global model by using a corresponding second subset of filters.

16. The central node or server according to any one of claims 11 to 12, wherein: The processor is also configured to instruct one or more of the first user device and the second user device to extract their corresponding local models as the neural network model type.

17. A user equipment in a communication network, comprising: Memory; a processor, coupled to the memory, wherein the processor is configured to: extracting a local model as a first extracted model, wherein the local model is of a first model type and the first extracted model is of a second model type different from the first model type; Sending the first extraction model to a server; receiving a global model from the server, wherein the global model is constructed based on at least the first extraction model, and the global model is of the second model type; as well as The local model is updated based on the global model.

18. The user equipment according to claim 17, wherein: The processor is further configured to: updating the local model based on new data received at the user device; extracting the updated local model as a second extracted model, wherein the second extracted model is of the second model type; A weighted average of the second extraction model and the first extraction model is sent to the server.

19. The user equipment according to claim 18, wherein: The weighted average of the second extraction model and the first extraction model is given by W1+αW2, where W1 represents the first extraction model, W2 represents the second extraction model, and 0<α<1.

20. The user equipment according to any one of claims 17 to 19, wherein: The processor is further configured to: determining coefficients of a final layer of the global model based on the local data; and The coefficients are sent to a central node or server.

21. A computer program product comprising instructions which, when executed by a processing circuit, cause the processing circuit to perform the method according to any one of claims 1 to 10.

22. A computer-readable storage medium storing instructions which, when executed by a processing circuit, cause the processing circuit to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Secure federated neural networks

    US20190012592A1

  • Methods and apparatus for federated training of a neural network using trusted edge devices

    US20190042937A1

  • Detection of malicious network activity

    US20190166144A1