A low-uplink-load federated learning method and device for a binary neural network
Patent Information
- Application Number
- CN202410021369.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-01-05
AI Technical Summary
[0003]过去,许多节省上传成本的方法被提取,如通过对边缘节点参数变化量进行压缩上传,然而这些方法受到参数的变化量大小限制,压缩倍数有限
[0052]根据本发明实施例的一种用于二值神经网络的低上行负载联邦学习装置,通过边缘节点上传的二值参数估计其对应的实值参数的变化量从而更新中心节点维护的全局模型,从而减少上传数据量,提高边缘节点模型推理速度。
Smart Images

Figure CN117852625B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, and in particular to a low uplink load federated learning method for binary neural networks and a low uplink load federated learning apparatus for binary neural networks. Background Technology
[0002] Federated learning (FL) has been proposed as a new distributed computing paradigm. It can train a global model without accessing the datasets of edge nodes, and has the advantages of protecting data privacy and breaking down data barriers. However, due to the different network environments of edge devices, how to reduce the cost of uploading data is an important research direction in federated learning.
[0003] In the past, many methods to reduce upload costs have been extracted, such as compressing the changes in edge node parameters before uploading. However, these methods are limited by the magnitude of the parameter changes, resulting in a limited compression ratio. Additionally, some methods use fewer bits to represent network parameters and upload the low-bit representation of the network parameters to reduce the number of bits used for uploading a single parameter; however, these methods suffer from model performance degradation. In summary, existing federated learning methods for reducing the amount of uploaded data still face the problem of large uploaded data volumes or compromised model performance. Summary of the Invention
[0004] This invention aims to at least partially solve one of the technical problems in the aforementioned technologies. To this end, one objective of this invention is to propose a low-upload federated learning method for binary neural networks. This method updates the global model maintained by the central node by estimating the changes in the corresponding real-valued parameters based on the binary parameters uploaded by the edge nodes, thereby reducing the amount of uploaded data and improving the inference speed of the edge node model.
[0005] A second objective of this invention is to provide a low uplink load federated learning device for binary neural networks.
[0006] To achieve the above objectives, a first aspect of the present invention proposes a low-uplink-load federated learning method for binary neural networks, comprising the following steps: each edge node instantiates a local binary neural network model according to the model definition of the binary neural network sent by the central node; in each training round, the central node sends global model parameters to each edge node selected in the current round, so that each edge node selected in the current round loads the global model parameters into the corresponding local binary neural network model and performs training in conjunction with the local dataset to obtain the corresponding updated binary parameters, real-valued parameters, and auxiliary parameters, wherein the auxiliary parameters include the updated... The mean and standard deviation of the updated binary parameters, the slope and intercept of the linear mapping parameters of the network layer that performs linear mapping on the updated binary parameters and then binarizes them, and the size of the training dataset; the central node receives the updated binary parameters, real-valued parameters, and auxiliary parameters sent by each edge node selected in the current round, so as to estimate the change in the real-valued parameters of each edge node selected in the current round, so as to obtain the corresponding estimated value of the change in the real-valued parameters; the central node aggregates the estimated values of the change in the real-valued parameters corresponding to each edge node selected in the current round according to the size of the training dataset and updates the global model.
[0007] According to the low uplink load federated learning method for binary neural networks of the present invention, firstly, each edge node instantiates a local binary neural network model based on the model definition of the binary neural network sent by the central node; then, in each training round, the central node sends the global model parameters to each edge node selected in the current round, so that each edge node selected in the current round loads the global model parameters into the corresponding local binary neural network model and trains it with the local dataset to obtain the corresponding updated binary parameters, real-valued parameters, and auxiliary parameters, wherein the auxiliary parameters include the mean square of the updated binary parameters. The central node receives the updated binary parameters, real-valued parameters, and auxiliary parameters from each edge node selected in the current round. This allows it to estimate the changes in the real-valued parameters of each edge node selected in the current round, thus obtaining the corresponding estimated values of the real-valued parameter changes. Finally, the central node aggregates the estimated values of the real-valued parameter changes corresponding to each edge node selected in the current round based on the size of the training dataset and updates the global model.
[0008] In addition, the low uplink load federated learning method for binary neural networks proposed in the above embodiments of the present invention may also have the following additional technical features:
[0009] Optionally, the inference method of the binary neural network model is as follows:
[0010]
[0011] The gradient backpropagation method for binary parameters is as follows:
[0012]
[0013] Where, x, Let represent the real-valued input and real-valued inference result of the binary neural network, respectively, and θ represent all real-valued parameters of the binary neural network. Let θ represent the binary parameters of the l-th layer of the binary neural network. l This represents the real-valued parameters of the l-th layer of a binary neural network. This indicates that the function Binarize(·), f w The composite function formed by (·), Binarize(·) denotes the function that binarizes real values. This represents a function that linearly maps real-valued parameters, where... and Represents θ l For real-valued data of the same size, * indicates that the data is multiplied by its element position. This indicates that the output of the hidden layer of the binary neural network is generated by Binarize(·). Functions that perform linear mappings The composite function formed, where, and Indicates a l Real-valued data of the same size, N θ (·) represents the network inference process based on parameter θ. Let represent the cost function for training a model based on a training subset D of size N and model parameters θ, where y represents the training label of the model, Loss(·) represents the loss function used for network training, and L represents the total number of layers in the binary neural network.
[0014] Optionally, after the edge nodes k selected in the current round are updated, each element's binary parameter is represented by 1 bit. k θ b [t+1] uses 32 bits to represent the real-valued parameter of each element. k θ[t+1], the mean μ of the real-valued parameter represented using 32 bits. k [t+1]=mean( k θ[t+1]) and standard deviation σ k [t+1] = std( kθ[t+1]), the slope parameter included in the network layer l of the binary neural network after linear mapping and binarization of the parameters. and intercept parameter Dataset D for edge node training k Size | D k |, where mean(·) represents the function to calculate the mean of the input data, std(·) represents the function to calculate the standard deviation of the input data, and |·| represents the size of a set.
[0015] Optionally, the changes in the real-valued parameters of each edge node selected in the current round are estimated to obtain corresponding estimated values of the changes in the real-valued parameters, including:
[0016] A. Assume the following:
[0017] (1) Before training at any time t (t>0), the parameters of the model instance of the edge node numbered k. in, This represents the model parameters aggregated by the central node in round t. This indicates that the model parameters are randomly generated by the center node, meaning that before training begins, the edge nodes will download the parameters from the center node.
[0018] (2) At any time t, the j-th parameter of the i-th layer of the model instance of the edge node numbered k. k θ i,j All parameters of model instances that follow a normal distribution and have the same edge node are independently and identically distributed, i.e.
[0019] (3) It is a composite function consisting of the sign function and the linear function f(x) = qx + p (q > 0) with the range V ∈ (a, b), i.e. Where a,b∈R∪{+∞,-∞};
[0020] (4) The mean μ of the real-valued parameters updated after t+1 rounds of training uploaded by any edge node k. k [t+1], standard deviation σ k [t+1], binary parameter k θ b [t+1] and the set of parameters related to the linear mapping of parameters, i.e., the slope set. and intercept set Among them, L b N represents the number of binary network layers included in the model. i This represents the number of parameters contained in the i-th network layer, where, when When no data is uploaded, the central node's default value is 0. When no data is uploaded, the central node defaults to a value of 1, and retains the model parameters obtained from the aggregation at the previous time step t.
[0021] B. Using the assumptions in A, we obtain the change in the real-valued parameter of the k-th edge node in round t+1. The derivation process is as follows:
[0022] To simplify the symbolic representation and reasoning process, we will use f to represent f below. w Let p represent q represents That is, use f( k θ i,j )=q k θ i,j +p indicates The binary function Binarize(x) takes
[0023] The real-valued change Δ of the j-th parameter in the i-th layer of the k-th edge node k θ i,j [t+1] can be expressed by the formula:
[0024]
[0025] Assuming the range of function f is V∈(a,b), since the domain of f in a real network may belong to some interval in the field of all real numbers R, the algorithm defines a,b∈R∪{+∞,-∞}; when it is known At that time, the following relationship exists:
[0026]
[0027] Therefore, we can deduce Δ k θ i,j The range of values for [t+1]:
[0028]
[0029]
[0030] Because the algorithm estimates Δ from a probabilistic perspective. k θ i,j [t+1], so to avoid symbol confusion, we will use [t+1] below. k Y i,j The quantity to be estimated is Δ k θ i,j [t+1], k X i,j This indicates that N(f(μ) satisfies a normal distribution.k [t+1],(qσ k [t+1]) 2 ) the unknown variable f( k θ i,j [t+1]), k c i,j represents a known constant that is According to the superposition property of normal distribution, the variable to be estimated k Y i,j follows a normal distribution
[0031] The algorithm takes the expectation E of the random variable k Y i,j when the value of the conditional variable is known as the estimate of Δ k θ i,j [t+1] Combined with the value range of Δ k Z i,j [t+1] under the condition of known conditional variable k θ i,j [t+1], the following can be obtained:
[0032]
[0033]
[0034] Due to and there exists the 3-sigma rule for normal distribution, the following can be obtained thus the algorithm simplifies the above formula to
[0035]
[0036] wherein, the function g(x) represents the probability density function of the random variable X~N(μ,σ) under the condition t1<X<t2, that is
[0037]
[0038] Through mathematical derivation, the expectation of the random variable X under the condition t1<X<t2 is:
[0039]
[0040] wherein, the function ξ represents the probability density function of standard normal distribution, and the function ζ represents the cumulative distribution function of standard normal distribution;
[0041] Finally can be calculated by the following formula:
[0042]
[0043] Using the algorithm described above, the central node can estimate the changes in the real-valued parameters of the binary network layer at each edge node k at time t+1.
[0044] For real-valued parameters, the central node calculates the change in k of each edge node at time t+1 using the following formula.
[0045]
[0046] Among them, L r This indicates the number of network layers that submit real-valued parameters;
[0047] The change in real-valued parameters Δ of the binary neural network at the k-th edge node in the (t+1)-th round. k θ[t+1] can be expressed as:
[0048] Optionally, the estimated values of the changes in real-valued parameters corresponding to each edge node selected in the current round are aggregated and the global model is updated according to the following formula:
[0049]
[0050] Where, .+ means adding the elements of the set according to their corresponding positions, c.*S means multiplying each element of set S by a constant c, and β represents the update rate of the global model, which is usually taken as 1.
[0051] To achieve the above objectives, a second aspect of the present invention proposes a low-uplink-load federated learning device for binary neural networks, comprising: a model building module, configured to instantiate a local binary neural network model by each edge node according to the model definition of the binary neural network sent by the central node; and a local training module, configured to, in each training round, have the central node send global model parameters to each edge node selected in the current round, so that each edge node selected in the current round loads the global model parameters into the corresponding local binary neural network model and trains it in conjunction with a local dataset to obtain corresponding updated binary parameters, real-valued parameters, and auxiliary parameters, wherein the auxiliary parameters include the updated... The network includes the mean and standard deviation of the updated binary parameters, the slope and intercept of the linear mapping parameters of the network layer after linear mapping of the updated binary parameters and then binarization, and the size of the training dataset; an estimation module, used by the central node to receive the updated binary parameters, real-valued parameters and auxiliary parameters sent by each edge node selected in the current round, so as to estimate the change in the real-valued parameters of each edge node selected in the current round, so as to obtain the corresponding real-valued parameter change estimate; and a model update module, used by the central node to aggregate the real-valued parameter change estimates corresponding to each edge node selected in the current round according to the size of the training dataset and update the global model.
[0052] According to an embodiment of the present invention, a low uplink load federated learning device for binary neural networks estimates the change in the corresponding real-valued parameters by means of binary parameters uploaded by edge nodes, thereby updating the global model maintained by the central node, thereby reducing the amount of uploaded data and improving the inference speed of the edge node model. Attached Figure Description
[0053] Figure 1 This is a flowchart illustrating a low-upload federated learning method for binary neural networks according to an embodiment of the present invention.
[0054] Figure 2 This is a schematic diagram of the structure of a binary neural network model according to an embodiment of the present invention;
[0055] Figure 3 This invention relates to performance metrics for different numbers of edge nodes W, different single-round edge node selection ratios C, different models, and different aggregation methods on the MNIST dataset, according to an embodiment of the present invention.
[0056] Figure 4 This invention provides performance metrics for different edge node numbers W, different single-round edge node selection ratios C, different models, and different aggregation methods on the Fashion-MNIST dataset, according to an embodiment of the present invention.
[0057] Figure 5 This is a schematic diagram of the Euclidean distance between the real-valued parameter changes estimated by the P-aggregation method and the B-aggregation method during the training process according to an embodiment of the present invention and the actual real-valued parameter changes calculated by the F-aggregation method.
[0058] Figure 6 This is a block diagram of a low-upload federated learning device for a binary neural network according to an embodiment of the present invention. Detailed Implementation
[0059] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0060] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the invention to those skilled in the art.
[0061] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0062] Figure 1 This is a flowchart illustrating a low-upload federated learning method for binary neural networks according to an embodiment of the present invention, as shown below. Figure 1 As shown, this low-upload federated learning method for binary neural networks includes the following steps:
[0063] S101, each edge node instantiates a local binary neural network model according to the model definition of the binary neural network sent by the central node.
[0064] In other words, when the federated learning task is started, each edge node instantiates a local binary neural network model based on the model definition of the binary neural network broadcast by the central node.
[0065] As an example, the inference method of the above binary neural network model is as follows:
[0066] Furthermore, the gradient backpropagation method for binary parameters is as follows:
[0067]
[0068] Where, x, Let represent the real-valued input and real-valued inference result of the binary neural network, respectively, and θ represent all real-valued parameters of the binary neural network. Let θ represent the binary parameters of the l-th layer of the binary neural network. l This represents the real-valued parameters of the l-th layer of a binary neural network. This indicates that the function Binarize(·), f w The composite function formed by (·), Binarize(·) denotes the function that binarizes real values. This represents a function that linearly maps real-valued parameters, where... and Represents θ l For real-valued data of the same size, * indicates the dot product of the elements by their positions, φ(·) = Binarize(f a (a l )) indicates that the output of the hidden layer of the binary neural network is generated by Binarize(·). Functions that perform linear mappings The composite function formed, where, and Indicates a l Real-valued data of the same size, N θ (·) represents the network inference process based on parameter θ. Let represent the cost function for training a model based on a training subset D of size N and model parameters θ, where y represents the training label of the model, Loss(·) represents the loss function used for network training, and L represents the total number of layers in the binary neural network.
[0069] S102, in each training round, the central node sends the global model parameters to each edge node selected in the current round, so that each edge node selected in the current round can load the global model parameters into the corresponding local binary neural network model and train it in combination with the local dataset to obtain the corresponding updated binary parameters, real-valued parameters and auxiliary parameters. The auxiliary parameters include the mean and standard deviation of the updated binary parameters, the slope and intercept of the linear mapping parameters of the network layer that performs linear mapping on the updated binary parameters and then binarizes them, and the size of the dataset used for training.
[0070] As an example, after the edge nodes k selected in the current round are updated, each element's binary parameter is represented by 1 bit. k θ b [t+1] uses 32 bits to represent the real-valued parameter of each element. k θ[t+1], the mean μ of the real-valued parameter represented using 32 bits. k[t+1]=mean( k θ[t+1]) and standard deviation σ k [t+1] = std( k θ[t+1]), the slope parameter included in the network layer l of the binary neural network after linear mapping and binarization of the parameters. and intercept parameter Dataset D for edge node training k Size | D k |, where mean(·) represents the function to calculate the mean of the input data, std(·) represents the function to calculate the standard deviation of the input data, and |·| represents the size of a set.
[0071] S103, the central node receives the updated binary parameters, real-value parameters and auxiliary parameters sent by each edge node selected in the current round, so as to estimate the change in the real-value parameters of each edge node selected in the current round, so as to obtain the corresponding estimated value of the change in the real-value parameters.
[0072] As an example, the change in the real-valued parameters of each edge node selected in the current round is estimated to obtain the corresponding estimated value of the change in real-valued parameters, including:
[0073] A. Assume the following:
[0074] (1) Before training at any time t (t>0), the parameters of the model instance of the edge node numbered k. in, This represents the model parameters aggregated by the central node in round t. This indicates that the model parameters are randomly generated by the center node, meaning that before training begins, the edge nodes will download the parameters from the center node.
[0075] (2) At any time t, the j-th parameter of the i-th layer of the model instance of the edge node numbered k. k θ i,j All parameters of model instances that follow a normal distribution and have the same edge node are independently and identically distributed, i.e.
[0076] (3) It is a composite function consisting of the sign function and a linear function f(x) = qx + p with range V ∈ (a, b), where q > 0. Where a,b∈R∪{+∞,-∞};
[0077] (4) The mean μ of the real-valued parameters updated after t+1 rounds of training uploaded by any edge node k. k [t+1], standard deviation σ k [t+1], binary parameterk θ b [t+1] and the set of parameters related to the linear mapping of parameters, i.e., the slope set. and intercept set Among them, L b N represents the number of binary network layers included in the model. i This represents the number of parameters contained in the i-th network layer, where, when When no data is uploaded, the central node's default value is 0. When no data is uploaded, the central node defaults to a value of 1, and retains the model parameters obtained from the aggregation at the previous time step t.
[0078] B. Using the assumptions in A, we obtain the change in the real-valued parameter of the k-th edge node in round t+1. The derivation process is as follows:
[0079] To simplify the symbolic representation and reasoning process, we will use f to represent f below. w Let p represent q represents That is, use f( k θ i,j )=q k θ i,j +p indicates The binary function Binarize(x) takes
[0080] The real-valued change Δ of the j-th parameter in the i-th layer of the k-th edge node k θ i,j [t+1] can be expressed by the formula:
[0081]
[0082] Assuming the range of function f is V∈(a,b), since the domain of f in a real network may belong to some interval in the field of all real numbers R, the algorithm defines a,b∈R∪{+∞,-∞}; when it is known At that time, the following relationship exists:
[0083]
[0084] Therefore, we can deduce Δ k θ i,j The range of values for [t+1]:
[0085]
[0086]
[0087] Since the algorithm estimates Δ from the perspective of probability theory k θ i,j [t+1], in order to avoid symbol confusion, the following uses k Y i,j to represent the quantity to be estimated Δ k θ i,j [t+1], k X i,j represents the unknown variable f( that satisfies the normal distribution N(f(μ k [t+1],(qσ k [t+1]) 2 ) k θ i,j [t+1]), k c i,j represents a known constant that is According to the superposition property of normal distribution, the quantity to be estimated k Y i,j obeys the normal distribution
[0088]
[0089] The algorithm takes the expectation E of the random variable k Y i,j when the value of the known conditional variable is fixed as the estimate of Δ k θ i,j [t+1] Combined with the value range of Δ k Z i,j under the condition of the known conditional variable k θ i,j [t+1], the following can be obtained:
[0090]
[0091]
[0092] Since and the normal distribution has the 3-sigma rule, it can be obtained that thus the algorithm simplifies the above formula to
[0093]
[0094] wherein, the function g(x) represents the probability density function of the random variable X~N(μ,σ) under the condition that t1<X<t2, that is
[0095]
[0096] Through mathematical derivation, the expectation of the random variable X under the condition t1<X<t2 is:
[0097]
[0098] Where ξ represents the probability density function of the standard normal distribution, and ζ represents the cumulative distribution function of the standard normal distribution;
[0099] final It can be calculated using the following formula:
[0100]
[0101] Using the algorithm described above, the central node can estimate the changes in the real-valued parameters of the binary network layer at each edge node k at time t+1.
[0102] For real-valued parameters, the central node calculates the change in k of each edge node at time t+1 using the following formula.
[0103]
[0104] Among them, L r This indicates the number of network layers that submit real-valued parameters;
[0105] The change in real-valued parameters Δ of the binary neural network at the k-th edge node in the (t+1)-th round. k θ[t+1] can be expressed as:
[0106] S104, the central node aggregates the estimated values of the changes in real-valued parameters corresponding to each edge node selected in the current round based on the size of the dataset involved in training and updates the global model.
[0107] As an example, the estimated values of the changes in real-valued parameters corresponding to each edge node selected in the current round are aggregated and the global model is updated according to the following formula:
[0108]
[0109] Where, .+ means adding the elements of the set according to their corresponding positions, c.*S means multiplying each element of set S by a constant c, and β represents the update rate of the global model, which is usually taken as 1.
[0110] Furthermore, this application will be further described through the following specific embodiments and in conjunction with the accompanying drawings.
[0111] This application utilizes the MNIST dataset from the National Institute of Standards and Technology (NIST) and the Fashion-MNIST dataset, which contains images in 10 categories: t-shirt, trousers, pullover, dress, coat, sandal, shirt, sneaker, bag, and ankle boot. The training set of the MNIST dataset consists of handwritten digits from 250 different people, 50% of whom are high school students and 50% are from Census Bureau staff. The test set also contains the same proportion of handwritten digit data. The Fashion-MNIST dataset consists of a training set containing 60,000 instances and a test set containing 10,000 instances.
[0112] In terms of experimental setup, this embodiment divides the training set of the above dataset into 20, 40, and 60 edge nodes, sets the proportion of edge nodes selected in each round of training to 0.5 and 1 respectively, constructs a total of 6 different experiments, and uses the accuracy of the global model on the test set and the number of parameters (in megabits (Mb)) that need to be uploaded by edge nodes in a single round as evaluation indicators.
[0113] Regarding model selection, this application selects examples such as... Figure 2 The network structure shown, consisting of two convolutional layers, two pooling layers, and one fully connected layer, serves as a use case model. To demonstrate that this application can be applied to binary neural network types satisfying the above-mentioned requirements, the examples in this application employ the following two types of binary neural network layers satisfying the binary neural network type:
[0114] (1) BNN, characterized by the use of binary functions BNN can be used as a convolutional layer (hereinafter referred to as BNN-CNN) and a fully connected layer (hereinafter referred to as BNN-FC).
[0115] (2) IR-CNN, characterized by the use of binary functions Auxiliary parameters
[0116] Ultimately, this application uses the following two binary network models:
[0117] (1) BNN network: where the convolutional layer and the fully connected layer are BNN-CNN and BNN-FC, respectively.
[0118] (2) IR network: where the convolutional and fully connected layers are IR-CNN and BNN-FC, respectively.
[0119] To demonstrate the performance of the aggregation method proposed in this application, the examples in this application employ the following three different parameter aggregation methods.
[0120] (1) Real-valued parameter aggregation (hereinafter referred to as F-aggregation) is characterized by the fact that for the (k+1)th round of training, the parameter change of the j-th parameter of the i-th layer of any edge node k is:
[0121] (2) Binary parameter aggregation (hereinafter referred to as B aggregation) is characterized by the fact that for the (k+1)th round of training, the parameter change of the j-th parameter of the i-th layer of any edge node k is as follows: in
[0122] (3) The binary parameter aggregation method proposed in this application (hereinafter referred to as P aggregation).
[0123] Specific embodiments are given below.
[0124] The embodiments of this application include the following steps:
[0125] Step 1: The central node broadcasts the definitions of the BNN network and the IR network to each edge node, and the edge nodes instantiate the model according to the network definitions.
[0126] Step 2: The central node broadcasts the current global model parameters to the selected edge nodes.
[0127] Step 3: The selected edge nodes load the global model parameters into the locally instantiated model for training.
[0128] Step 4: In round t, for aggregation F, edge node k uploads the real-valued parameters of the local model. k θ[t+1]; For B aggregation, edge nodes upload the binary parameters of the local model. k θ b [t+1]; For P aggregation and the model is a BNN network, the edge nodes upload the binary parameters of the local model. k θ b [t+1], the mean μ of the real-valued parameters of the local model. k [t+1] and the standard deviation σ of the real parameter k [t+1]. For P-aggregation and an IR network model, edge nodes upload the binary parameters of their local models. k θ b [t+1], the mean μ of the real-valued parameters of the local model. k [t+1], the variance σ of the real-valued parameter k [t+1], and the slope of the linear mapping in the binary function of IR-CNN. and intercept
[0129] Step 5: The central node calculates the change in the real-valued parameter k of the edge node in round t+1. k θ[t+1], the calculation methods for the three aggregation methods are as described in the above introduction to aggregation methods.
[0130] Step 6: The central node updates the global model parameters.
[0131] Where .+ indicates that the elements of the set are added at their corresponding positions, c.*S indicates that each element of set S is multiplied by the constant c, and β represents the update rate of the global model, which is 1.
[0132] Step 7: The central node calculates the performance of the global model parameters on the test set.
[0133] Step 8: If the termination condition is not met, repeat steps 2 through 8.
[0134] In this embodiment, the evaluation index results obtained by different aggregation methods are as follows: Figure 3 , Figure 4 As shown, the Euclidean distance between the real-valued parameter changes estimated by the P-aggregation method and the B-aggregation method proposed in this application during training and the actual real-valued parameter changes calculated by the F-aggregation method is as follows: Figure 5 As shown. Compared with the prior art, this application better estimates the change of real-valued parameters corresponding to binary parameters based on the properties of neural network parameter distribution, thereby achieving excellent training results for the global model while reducing the amount of data uploaded by edge nodes in a single round by tens of times. At the same time, the edge model greatly improves the inference speed by running a binary neural network.
[0135] To achieve the above embodiments, this invention also proposes a low-upload federated learning device for binary neural networks, such as... Figure 6 As shown, the low uplink load federated learning device for binary neural networks includes: a model building module 10, a local training module 20, an estimation module 30, and a model update module 40.
[0136] The model building module is used by each edge node to instantiate a local binary neural network model based on the model definition of the binary neural network sent by the central node. The local training module is used in each training round where the central node sends global model parameters to each edge node selected in the current round. This allows each edge node to load the global model parameters into its corresponding local binary neural network model and train it using a local dataset to obtain updated binary parameters, real-valued parameters, and auxiliary parameters. The auxiliary parameters include the mean and standard deviation of the updated binary parameters, and the values of the real-valued parameters relative to the updated binary parameters. The network layer, after linear mapping of the value parameters and then binarization, includes the linear mapping parameters slope and intercept, and the size of the training dataset; the estimation module is used by the central node to receive the updated binary parameters, real-value parameters, and auxiliary parameters sent by each edge node selected in the current round, so as to estimate the change in the real-value parameters of each edge node selected in the current round, so as to obtain the corresponding estimated value of the change in the real-value parameters; the model update module is used by the central node to aggregate the estimated values of the change in the real-value parameters corresponding to each edge node selected in the current round according to the size of the training dataset and update the global model.
[0137] It should be noted that the above description of the low uplink load federated learning method for binary neural networks also applies to the low uplink load federated learning device for binary neural networks, and will not be repeated here.
[0138] In summary, the low uplink load federated learning device for binary neural networks according to an embodiment of the present invention estimates the change in the corresponding real-valued parameters by using the binary parameters uploaded by edge nodes to update the global model maintained by the central node, thereby reducing the amount of uploaded data by tens of times. At the same time, running the binary neural network on the edge nodes can increase the inference speed of the model by tens of times. Thus, the amount of uploaded data is small, the training quality of the global network is high, the inference speed of the edge node model is fast, and it is compatible with various binary neural networks that perform linear mapping of real-valued parameters before parameter binarization.
[0139] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0140] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0141] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0142] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0143] It should be noted that any reference signs placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0144] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0145] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0146] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0147] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0148] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.
[0149] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms should not be construed as necessarily referring to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0150] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A low-upload federated learning method for binary neural networks, characterized in that, Includes the following steps: Each edge node instantiates a local binary neural network model based on the model definition of the binary neural network sent by the central node; In each training round, the central node sends the global model parameters to each edge node selected in the current round, so that each edge node selected in the current round loads the global model parameters into the corresponding local binary neural network model and trains it in combination with the local dataset to obtain the corresponding updated binary parameters, real-valued parameters and auxiliary parameters. The auxiliary parameters include the mean and standard deviation of the updated binary parameters, the slope and intercept of the linear mapping parameters of the network layer that performs linear mapping on the updated binary parameters and then binarizes them, and the size of the dataset used for training. The dataset contains the MNIST dataset and images of 10 categories. The central node receives updated binary parameters, real-value parameters, and auxiliary parameters from each edge node selected in the current round, so as to estimate the change in the real-value parameters of each edge node selected in the current round and obtain the corresponding estimated value of the change in the real-value parameters. The central node aggregates the estimated values of the changes in real-valued parameters corresponding to each edge node selected in the current round based on the size of the dataset participating in the training and updates the global model. Among them, each edge node selected in the current round After the update, each element's binary parameter is represented by 1 bit. The real-valued parameter of each element is represented using 32 bits. The mean of real-valued parameters represented using 32 bits and standard deviation In a binary neural network, the network layer that performs linear mapping of parameters before binarization... The included slope parameters and intercept parameter Dataset used for edge node training Size ,in, A function that calculates the mean of the input data. This function represents the calculation of the standard deviation of the input data. It represents the size of a set.
2. The low uplink load federated learning method for binary neural networks as described in claim 1, characterized in that, The reasoning method of the binary neural network model is as follows: The gradient backpropagation method for binary parameters is as follows: in, These represent the real-valued input and real-valued inference result of the binary neural network, respectively. Represents all real-valued parameters of a binary neural network. Represents the binary neural network's... Binary parameters of the layer, Represents the binary neural network's... Real-valued parameters of the layer, Indicates by function The composite function formed This represents a function that binarizes real values. This represents a function that linearly maps real-valued parameters, where... and express Real-valued data of the same size This indicates that the data is multiplied by its element position. , indicating by and the output of the hidden layer of the binary neural network Functions that perform linear mappings The composite function formed, where, and express Real-valued data of the same size Indicates based on parameters The network reasoning process, Indicates based on size training subset and model parameters The cost function for model training, where, The training labels of the model. This represents the loss function used for network training. This represents the total number of layers in a binary neural network.
3. The low uplink load federated learning method for binary neural networks as described in claim 2, characterized in that, Estimate the changes in the real-valued parameters of each edge node selected in the current round to obtain the corresponding estimated values of the changes in the real-valued parameters, including: A. Assume the following: (1) In any Before the training session, numbered Parameters of the model instance of the edge node ,in, Indicates the first The model parameters aggregated by the round center node This indicates that the model parameters are randomly generated by the center node, meaning that before training begins, the edge nodes will download the parameters from the center node. (2) at any time Number The first model instance of the edge node The first layer Parameters All parameters of model instances that follow a normal distribution and have the same edge node are independently and identically distributed, i.e. ; (3) yes Functions and range linear functions The composite function formed by q>0, i.e. ,in, ; (4) Arbitrary edge nodes Upload process Mean of real-valued parameters updated after each training round Standard deviation binary parameters And the set of parameters related to the linear mapping of parameters, i.e., the set of slopes. and intercept set ,in, This indicates the number of binary network layers included in the model. Indicates the first The number of parameters contained in each network layer, where, when When no data is uploaded, the central node's default value is 0. When no data is uploaded, the central node defaults to a value of 1, and the central node retains the value from the previous moment. Model parameters obtained by aggregation ; B. Using the assumptions in A, we obtain the... In the round The derivation of the changes in the real-valued parameters of each edge node is as follows: To simplify symbolic representation and reasoning processes, the following uses... express ,use express , express , that is, use express binary function Pick ; For the The edge node of the first The first layer The real value change of each parameter This can be expressed by the formula: hypothesis function range Because in actual networks The domain may belong to the field of all real numbers. The algorithm is defined within a certain interval. When it is known At that time, the following relationship exists: Therefore, we can deduce that... The range of values for: Because the algorithm estimates from a probabilistic perspective... Therefore, to avoid symbol confusion, we will use [symbols] below. Indicates the quantity to be estimated , This indicates that it follows a normal distribution. unknown variables , Represent a known constant ,Right now Based on the superposition property of the normal distribution, the quantity to be estimated Follows a normal distribution ; The algorithm will use random variables Given condition variables Expected value when taking the value As Estimate Combined with known condition variables In this case The range of values for can be obtained as follows: because Furthermore, the normal distribution follows a 3-standard-deviation rule, which allows us to obtain... Thus, the algorithm simplifies the above expression to, Among them, the function Indicates a given Under the condition, random variable The probability density function, i.e. ; Through mathematical derivation, random variables exist The expected value under the given conditions is: Among them, the function The probability density function of the standard normal distribution is denoted by . The cumulative distribution function represents the standard normal distribution; final It can be calculated using the following formula: Using the algorithm described above, the central node can be estimated. Each edge node at any time The change in real-valued parameters of a binary network layer ; For real-valued parameters, the center node is calculated using the following formula. Each edge node at any time Change in, This indicates the number of network layers that submit real-valued parameters; Final Rounds The change in real-valued parameters of a binary neural network with edge nodes It can be represented as: .
4. The low uplink load federated learning method for binary neural networks as described in claim 3, characterized in that, The estimated values of the changes in real-valued parameters corresponding to each edge node selected in the current round are aggregated and the global model is updated according to the following formula: in, This indicates that the elements of the set are added together at their corresponding positions. Indicates the set Each element multiplied by a constant , This represents the update rate of the global model, which is usually set to 1.
5. A low-upload federated learning device for binary neural networks, characterized in that, include: The model building module is used by each edge node to instantiate a local binary neural network model based on the model definition of the binary neural network sent by the central node. A local training module is used in each training round, whereby the central node sends global model parameters to each edge node selected in the current round, so that each edge node selected in the current round loads the global model parameters into the corresponding local binary neural network model and trains it in conjunction with the local dataset to obtain the corresponding updated binary parameters, real-valued parameters, and auxiliary parameters. The auxiliary parameters include the mean and standard deviation of the updated binary parameters, the slope and intercept of the linear mapping parameters of the network layer that performs linear mapping on the updated binary parameters and then binarizes them, and the size of the dataset used for training. The dataset contains the MNIST dataset and images of 10 categories. The estimation module is used by the central node to receive the updated binary parameters, real-value parameters and auxiliary parameters sent by each edge node selected in the current round, so as to estimate the change in the real-value parameters of each edge node selected in the current round, so as to obtain the corresponding estimated value of the change in the real-value parameters. The model update module is used by the central node to aggregate the estimated values of the changes in real-valued parameters corresponding to each edge node selected in the current round based on the size of the dataset participating in training, and then update the global model. Among them, each edge node selected in the current round After the update, each element's binary parameter is represented by 1 bit. The real-valued parameter of each element is represented using 32 bits. The mean of real-valued parameters represented using 32 bits and standard deviation In a binary neural network, the network layer that performs linear mapping of parameters before binarization... The included slope parameters and intercept parameter Dataset used for edge node training Size ,in, A function that calculates the mean of the input data. This function represents the calculation of the standard deviation of the input data. It represents the size of a set.
6. The low uplink load federated learning device for binary neural networks as described in claim 5, characterized in that, The reasoning method of the binary neural network model is as follows: The gradient backpropagation method for binary parameters is as follows: in, These represent the real-valued input and real-valued inference result of the binary neural network, respectively. Represents all real-valued parameters of a binary neural network. Represents the binary neural network's... Binary parameters of the layer, Represents the binary neural network's... Real-valued parameters of the layer, Indicates by function The composite function formed This represents a function that binarizes real values. This represents a function that linearly maps real-valued parameters, where... and express Real-valued data of the same size This indicates that the data is multiplied by its element position. , indicating by and the output of the hidden layer of the binary neural network Functions that perform linear mappings The composite function formed, where, and express Real-valued data of the same size Indicates based on parameters The network reasoning process, Indicates based on size training subset and model parameters The cost function for model training, where, The training labels of the model. This represents the loss function used for network training. This represents the total number of layers in a binary neural network.
7. The low uplink load federated learning device for binary neural networks as described in claim 6, characterized in that, Estimate the changes in the real-valued parameters of each edge node selected in the current round to obtain the corresponding estimated values of the changes in the real-valued parameters, including: A. Assume the following: (1) In any Before the training session, numbered Parameters of the model instance of the edge node ,in, Indicates the first The model parameters aggregated by the round center node This indicates that the model parameters are randomly generated by the center node, meaning that before training begins, the edge nodes will download the parameters from the center node. (2) at any time Number The first model instance of the edge node The first layer Parameters All parameters of model instances that follow a normal distribution and have the same edge node are independently and identically distributed, i.e. ; (3) yes Functions and range linear functions The composite function formed by q>0, i.e. ,in, ; (4) Arbitrary edge nodes Upload process Mean of real-valued parameters updated after each training round Standard deviation binary parameters And the set of parameters related to the linear mapping of parameters, i.e., the set of slopes. and intercept set ,in, This indicates the number of binary network layers included in the model. Indicates the first The number of parameters contained in each network layer, where, when When no data is uploaded, the central node's default value is 0. When no data is uploaded, the central node defaults to a value of 1, and the central node retains the value from the previous moment. Model parameters obtained by aggregation ; B. Using the assumptions in A, we obtain the... In the round The derivation of the changes in the real-valued parameters of each edge node is as follows: To simplify symbolic representation and reasoning processes, the following uses... express ,use express , express , that is, use express binary function Pick ; For the The edge node of the first The first layer The real value change of each parameter This can be expressed by the formula: hypothesis function range Because in actual networks The domain may belong to the field of all real numbers. The algorithm is defined within a certain interval. When it is known At that time, the following relationship exists: Therefore, we can deduce that... The range of values for: Because the algorithm estimates from a probabilistic perspective... Therefore, to avoid symbol confusion, we will use [symbols] below. Indicates the quantity to be estimated , This indicates that it follows a normal distribution. unknown variables , Represent a known constant ,Right now Based on the superposition property of the normal distribution, the quantity to be estimated Follows a normal distribution ; The algorithm will use random variables Given condition variables Expected value when taking the value As Estimate Combined with known condition variables In this case The range of values for can be obtained as follows: because Furthermore, the normal distribution follows a 3-standard-deviation rule, which allows us to obtain... Thus, the algorithm simplifies the above expression to, Among them, the function Indicates a given Under the condition, random variable The probability density function, i.e. ; Through mathematical derivation, random variables exist The expected value under the given conditions is: Among them, the function The probability density function of the standard normal distribution is denoted by . The cumulative distribution function represents the standard normal distribution; final It can be calculated using the following formula: Using the algorithm described above, the central node can be estimated. Each edge node at any time The change in real-valued parameters of a binary network layer ; For real-valued parameters, the center node is calculated using the following formula. Each edge node at any time Change in, This indicates the number of network layers that submit real-valued parameters; Final Rounds The change in real-valued parameters of a binary neural network with edge nodes It can be represented as: .
8. The low uplink load federated learning device for binary neural networks as described in claim 7, characterized in that, The estimated values of the changes in real-valued parameters corresponding to each edge node selected in the current round are aggregated and the global model is updated according to the following formula: in, This indicates that the elements of the set are added together at their corresponding positions. Indicates the set Each element multiplied by a constant , This represents the update rate of the global model, which is usually set to 1.