Control device, model learning device, secret federated learning device, their methods, and programs
The control device in the federated learning system enhances security and efficiency by managing asynchronous or synchronous processing based on device speeds, using secure computation to prevent data inference and optimize model aggregation.
Patent Information
- Application Number
- JP2024511019
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2042-03-31
AI Technical Summary
Conventional federated learning systems face security issues due to plaintext worker models being transmitted, allowing the federated learning device to infer the learning data tendencies, and efficiency is poor due to unconsidered processing speeds among devices.
A control device manages a federated learning system with model learning devices and secret federated learning devices, performing asynchronous or synchronous control based on processing times to enhance security and efficiency by using secure computation methods like multi-party computation or homomorphic encryption, and executing local and secret aggregation processes without directly obtaining worker models.
This approach improves the security and efficiency of federated learning by preventing inference of learning data tendencies and optimizing processing speeds, ensuring secure and efficient model aggregation.
Smart Images

Figure 0007709657000001 
Figure 0007709657000002 
Figure 0007709657000003
Abstract
Description
Technical Field
[0001] The present invention relates to machine learning technology, and more particularly to federated learning technology.
Background Art
[0002] Federated learning, which performs machine learning in a distributed state without aggregating learning data, is known (see, for example, Non-Patent Document 1, etc.). In federated learning, a plurality of model learning devices perform machine learning using the learning data they hold to generate worker models (local models), and transmit the generated worker models to a federated learning device. The federated learning device generates an aggregated model (global model) by aggregating the worker models sent from the plurality of model learning devices, and transmits the generated aggregated model to the plurality of model learning devices. The plurality of model learning devices that have received the aggregated model update the aggregated model by machine learning using the learning data they hold to generate new worker models, and transmit the generated worker models to the federated learning device. By repeating such processing, each model learning device can obtain an aggregated model in which the learning data held by the plurality of model learning devices is reflected in machine learning without passing the learning data it holds to the outside.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in conventional federated learning, the federated learning device receives the plaintext worker models from each model learning device. Therefore, the federated learning device can know the tendency of the learning data held by each model learning device based on the difference between the transmitted aggregated model and the received worker model. Also, in conventional federated learning, since the processing speed of each model learning device is not considered, the efficiency is poor.
[0005] The present invention has been made in view of such points, and an object thereof is to improve the security and efficiency of federated learning.
Means for Solving the Problems
[0006] A control device controls a federated learning system including a plurality of model learning devices and one or more secret federated learning devices. Here, the model learning device executes local processing that updates an aggregated model by machine learning using local learning data to obtain information for specifying a worker model, and provides secret information of the information for specifying the worker model to the secret federated learning device. Also, the secret federated learning device obtains secret information of information for specifying a new aggregated model obtained by aggregating a plurality of worker models without obtaining the worker models, by performing secret calculation using the secret information of the information for specifying the obtained worker models, and executes secret aggregation processing for providing the information for specifying the new aggregated model or the secret information of the information for specifying the new aggregated model to the plurality of model learning devices. The control device compares the local processing time corresponding to the local processing with the aggregation processing time corresponding to the secret aggregation processing, and performs asynchronous control to execute the local processing of the plurality of model learning devices asynchronously with each other when the local processing time is longer than the aggregation processing time, or when the local processing time is equal to or longer than the aggregation processing time, and performs synchronous control to synchronize the local processing of the plurality of model learning devices with each other when asynchronous control is not performed.
Effects of the Invention
[0007] This can improve the safety and efficiency of federated learning.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. 「First Embodiment」 In the present embodiment, a form is exemplified in which a federated learning system including a plurality of model learning devices and one or a plurality of secure federated learning devices is controlled by a control device provided separately from the secure federated learning device.
[0010] <Configuration> As illustrated in FIG. 1, the federated learning system 1 of this embodiment includes N model learning devices 11-1, …, 11-N that perform model learning, M secret federated learning devices 12-1, …, 12-M that perform federated learning by means of secure computation, and a control device 13 that controls the federated learning system 1. There is no limitation to the secure computation method. For example, this secure computation method may be a multi-party computation method that performs secure computation using secretly-shared shares, or may be a homomorphic encryption method that performs secure computation using homomorphic encryption. N is an integer of 2 or more. M is an integer of 1 or more. For example, M is an integer of 2 or more. However, when the secure computation method is a multi-party computation method, M is an integer of 2 or more. When the secure computation method is a homomorphic encryption method, M is an integer of 1 or more. For example, M = 1.
[0011] As illustrated in FIG. 2, the model learning device 11-n of this embodiment includes a storage unit 111-n, an acquisition unit 112-n, a learning unit 113-n, a concealment unit 114-n, a provision unit 115-n, a control unit 116-n, and a determination unit 117-n. The model learning device 11-n executes each process based on the control of the control unit 116-n. The input information and the information obtained in each process are stored in the storage unit 111-n and read out and used as needed. Here, n is a positive integer, and n = 1, …, N. Unless otherwise specified, the configurations and processes related to n are the same for all n = 1, …, N. However, the content of the data (information) to be processed may vary depending on the value of n.
[0012] As illustrated in FIG. 3, the secret federated learning device 12-m of this embodiment includes an acquisition unit 121-m, a secret aggregation processing unit 122-m, a provision unit 123-m, a control unit 126-m, a storage unit 127-m, and a determination unit 128-m. The secret federated learning device 12-m executes each process based on the control of the control unit 126-m. The input information and the information obtained in each process are stored in the storage unit 127-m and read out and used as needed. Here, m is a positive integer, and m = 1, …, M. Unless otherwise specified, the configurations and processes related to m are the same for all m = 1, …, M. However, the content of the data (information) to be processed may vary depending on the value of m.
[0013] As illustrated in FIG. 4, the control device 13 of the present embodiment includes a measurement unit 131, a comparison unit 132, and a control unit 133.
[0014] <Preprocessing> In the storage unit 111-n of the model learning device 11-n, the local learning data D-n of each model learning device 11-n is stored. The local learning data D-n is learning data for machine learning, and may be learning data for supervised learning or unsupervised learning. Also, the local learning data D-n may be updated. Further, when the secret calculation method used is a homomorphic encryption method, an encryption key and a decryption key are stored in the storage unit 111-n of the model learning device 11-n.
[0015] <Learning Process> The learning process of the present embodiment is illustrated with reference to FIG. 5. Each model learning device 11-n (where n = 1,..., N) (FIG. 1) updates the aggregated model by machine learning using the local learning data D-n to obtain information WM-n (for example, a group of model parameters) that identifies the worker model, and executes a local process of providing the secret information [WM-n] of the information WM-n that identifies the worker model to the secret federated learning device 12-m (where m = 1,..., M).
[0016] Each secret federated learning device 12-m (where m = 1,..., M) uses the secret information [WM-n] of the information WM-n that identifies the acquired worker model to identify the secret information [GM] of the information GM that aggregates a plurality of worker models without obtaining the worker models through secret calculation. m to obtain the secret information [GM] of the information GM that identifies the new aggregated model m and executes a secret aggregation process of providing the secret information [GM] of the information GM that identifies the new aggregated model to a plurality of model learning devices 11-n.
[0017] These local processing and secret aggregation processing are alternately repeated until a predetermined end condition (for example, the number of updates, update amount, update time, etc. of the aggregation model reach a specified value, etc.) is satisfied. At this time, the anonymized information [GM] obtained by the secret aggregation processing m The new aggregation model corresponding to is used as the aggregation model in the local processing to be executed next. Here, since each secret federated learning device 12-m cannot obtain the worker model itself, it is also impossible to know the tendency of the local learning data D-n held by each model learning device 11-n based on the difference between the worker model and the aggregation model. Thereby, the security of federated learning can be improved.
[0018] The measurement unit 131 (Fig. 4) of the control device 13 measures the local processing time T1 corresponding to the local processing and the aggregation processing time T2 corresponding to the secret aggregation processing in the federated learning system 1. For example, the local processing time T1 is the time required for one local processing, and the aggregation processing time T2 is the time required for one secret aggregation processing. Alternatively, the local processing time T1 may be the total time during which local processing was performed out of the time required for B local processings, and the aggregation processing time T2 may be the total time during which secret aggregation processing was performed out of the time required for B secret aggregation processings. However, B is an integer of 2 or more. Alternatively, the local processing time T1 may be the average time required for one local processing, and the aggregation processing time T2 may be the average time required for one secret aggregation processing. For example, one local processing means that all N model learning devices 11-1, …, 11-N execute local processing, and one secret aggregation processing means that all M secret federated learning devices 12-1, …, 12-M perform secret aggregation processing. Alternatively, for example, one local processing means that C% or more of the N model learning devices 11-1, …, 11-N execute local processing, and one secret aggregation processing means that C% or more of the M secret federated learning devices 12-1, …, 12-M perform secret aggregation processing. However, C is a real number satisfying 0 < C ≤ 100. Alternatively, for example, one local processing means that one model learning device 11-n executes local processing, and one secret aggregation processing means that one secret federated learning device 12-m executes local processing. The measured local processing time T1 and aggregation processing time T2 are sent to the comparison unit 132 (step S131).
[0019] The comparison unit 132 compares the sent local processing time T1 with the aggregated processing time T2 (step S132). Here, when the local processing time T1 is longer than the aggregated processing time T2 (T1>T2), or when the local processing time T1 is greater than or equal to the aggregated processing time T2 (T1≧T2), the control unit 133 performs asynchronous control to execute the local processing of the plurality of model learning devices 11-1,…,11-N asynchronously with each other. Then, it returns to step S131. When T1>T2 or T1≧T2, there is a time margin on the side of the secret federated learning devices 12-1,…,12-M compared to the side of the model learning devices 11-1,…,11-N. In such a case, it is often more efficient to execute the local processing of the model learning devices 11-1,…,11-N asynchronously with each other and execute the secret aggregation processing of the secret federated learning devices 12-1,…,12-M before the local processing of all the model learning devices 11-1,…,11-N is completed (step S1331). On the other hand, when asynchronous control is not performed (that is, when the local processing time T1 is less than or equal to the aggregated processing time T2, or when the local processing time T1 is shorter than the aggregated processing time T2), the control unit 133 performs synchronous control to synchronize the local processing of the plurality of model learning devices 11-1,…,11-N with each other. Then, it returns to step S131. When T1≦T2 or T1<T2, there is a time margin on the side of the model learning devices 11-1,…,11-N compared to the side of the secret federated learning devices 12-1,…,12-M. In such a case, it is often more efficient to synchronize the local processing of the model learning devices 11-1,…,11-N with each other and execute the secret aggregation processing of the secret federated learning devices 12-1,…,12-M based on their processing results (step S1332).
[0020] [An example of synchronous control (step S1332)] The processing of the federated learning system 1 by synchronous control is exemplified. Note that the following synchronous control is an example and does not limit the present invention. In synchronous control, each model learning device 11-n executes local processing synchronized with other model learning devices 11-n' (where n, n' ∈ {1, …, N}) included in the federated learning system 1. That is, the plurality of model learning devices 11-1, …, 11-N execute local processing synchronized with each other. When synchronous control is performed, the control unit 133 instructs each model learning device 11-n (Fig. 2) to execute local processing by synchronous control. This instruction is acquired by the acquisition unit 112-n and sent to the control unit 116-n. The control unit 116-n executes synchronous control. The learning unit 113-n reads the local learning data D-n stored in the storage unit 111-n, updates the latest aggregated model by machine learning using the local learning data D-n to obtain a worker model, and outputs information WM-n (for example, a model parameter group) specifying the worker model. When the model learning device 11-n has not yet obtained an aggregated model, the initially set machine learning model is the "latest aggregated model". The machine learning model initially set by the control device 13 may be provided. The initially set model is, for example, a machine learning model with an initial model parameter group set. When the model learning device 11-n has obtained information GM for specifying an aggregated model as described later, among the aggregated models specified by the information GM, the latest one is the "latest aggregated model". In the latter case, the learning unit 113-n specifies the latest aggregated model based on the information GM read from the storage unit 111-n. Note that the aggregated model and the worker model are well-known machine learning models. There is no limitation to the aggregated model and the worker model. For example, it may be a model based on a deep learning method, a model based on a hidden Markov model method, a model based on a support vector machine method, or a model based on linear prediction. However, all the aggregated models and worker models handled in the federated learning system 1 are models based on the same method. The information WM-n for specifying the worker model is sent to the secrecy unit 114-n (step S113-n).
[0021] The concealment unit 114-n receives input of information WM-n that identifies a worker model. The concealment unit 114-n conceals the information WM-n that identifies a worker model using a method that allows the above-mentioned secure computation, obtains and outputs the concealed information [WM-n] of the information WM-n that identifies the worker model. For example, when the above-mentioned secure computation method is a multi-party computation method, the concealment unit 114-n secretly distributes the information WM-n into M pieces to obtain M shares [WM-n]1, ..., [WM-n] M For example, if the above-mentioned secure computation method is a homomorphic encryption method, the concealment unit 114-n uses the encryption key read from the storage unit 111-n to encrypt the information WM-n in accordance with the homomorphic encryption method, and outputs M (for example, 1) ciphertexts [WM-n]1, ..., [WM-n] M The ciphertext [WM-n]1, ..., [WM-n] M is output as secret information [WM-n]. Secret information [WM-n] = {[WM-n]1, ..., [WM-n] M} is sent to providing unit 115-n (step S114-n).
[0022] The provider 115-n stores confidential information [WM-n]={[WM-n]1, ..., [WM-n] of information WM-n for identifying a worker model. M The providing unit 115-n receives the confidential information [WM-n] of the information WM-n that identifies the worker model. m to the secret associative learning device 12-m (FIG. 3) (where m=1, . . . , M). Furthermore, the providing unit 115-n transmits (provides) the secret information [WM-n] to the secret associative learning device 12-m (FIG. 3) (where m=1, . . . , M). m The model learning device 11-n has completed learning of the worker model, and the confidential information [WM-n] of the worker model WM-n has been transmitted. m The secret association learning device 12-m transmits the synchronization information to the control device 13 (step S115-n).
[0023] The acquisition unit 121-m of the secret federated learning device 12-m (Fig. 3) receives the confidential information [WM-n] of the information WM-n that identifies the worker model sent from the model learning device 11-n m and stores the confidential information [WM-n] m in the storage unit 127-m. That is, the acquisition unit 121-m obtains the confidential information [WM-n] of the information WM-n that identifies a plurality of worker models from a plurality of model learning devices 11-n m and stores it in the storage unit 127-m (step S121-m).
[0024] Based on the synchronization information, the control unit 133 of the control device 13 determines whether all the model learning devices 11-1, …, 11-N have sent the confidential information [WM-n]1, …, [WM-n] to all the secret federated learning devices 12-1, …, 12-M M (step S1332a). Here, if it is determined that not all the model learning devices 11-n (where n = 1, …, N) have sent the confidential information [WM-n]1, …, [WM-n] to all the secret federated learning devices 12-1, …, 12-M M and it is determined that a predetermined time has not elapsed from the reference time point (no timeout), the control unit 133 makes the determination in step S1332a at regular intervals. On the other hand, if it is determined that all the model learning devices 11-n (where n = 1, …, N) have sent the confidential information [WM-n]1, …, [WM-n] to all the secret federated learning devices 12-1, …, 12-M M or if it is determined that a predetermined time has elapsed from the reference time point (timeout), the control unit 133 sends an instruction to start the secret aggregation process to the secret federated learning devices 12-1, …, 12-M. Note that the reference time point for the above-mentioned timeout can be any one, for example, it may be based on the start or end time of the previous secret aggregation process, or if the secret aggregation process has not been executed yet, it may be based on the start time of the learning process (S1332b).
[0025] An instruction to start the secret aggregation process is received by the acquisition unit 121-m of the secret collaborative learning device 12-m (where m = 1, …, M) (Fig. 3) and input to the control unit 126-m. The control unit 126-m that has received the instruction to start the secret aggregation process instructs the secret aggregation processing unit 122-m to start the secret aggregation process. In response to this, the secret aggregation processing unit 122-m reads a plurality of encrypted information [WM-n] (where n ∈ {1, …, N}) (encrypted information of information identifying a plurality of worker models) from the storage unit 127-m, and by performing secret calculations using these, without obtaining the plurality of worker models, identifies the encrypted information [GM] of the information GM that identifies the aggregated model obtained by aggregating the plurality of worker models. m is obtained and output. For example, if the information WM-n that identifies the worker model is the model parameter group {p1(n), …, p K (n)} of the worker model, for {n1, …, n max} ⊆ {1, …, N}, the model parameter groups {p1(n1), …, p K (n1)}, …, {p1(n max ), …, p K (n max )} are aggregated to obtain the model parameter group {p1, …, p K}, which becomes the information GM that identifies the aggregated model. For example, p k is a function value such as a weighted linear combination value or an average value of p k (n1), …, p k (n max ). Here, k is an index k = 1, …, K that identifies the model parameter, and K is a positive integer. The secret aggregation processing unit 122-m obtains and outputs the encrypted information [GM] of the information GM that identifies the aggregated model by secret calculations without restoring the information WM-n that identifies such a worker model or the information GM that identifies the aggregated model. m The encrypted information [GM] of the information GM that identifies the aggregated model m is sent to the providing unit 123-m (step S122-m).
[0026] The encrypted information [GM] m is input to the providing unit 123-m. The providing unit 123-m, via the control device 13, sends the encrypted information [GM]m It is transmitted (provided) to a plurality of model learning devices 11-n (where n ∈ {1, …, N}). For example, the providing unit 123-m transmits (provides) the confidential information [GM] via the control device 13 m to all the model learning devices 11-1, …, 11-N (step S123-m).
[0027] Confidential information [GM] m (where m ∈ {1, …, M}) the acquisition unit 112-n of the model learning device 11-n (Fig. 2) to which it is sent receives the confidential information [GM] m (confidential information of the information GM that identifies the aggregated model provided from the secret coalition learning device 12-m). The acquisition unit 112-n restores the confidential information [GM] m to obtain the information GM that identifies the aggregated model. For example, when the secret calculation method is a multi-party calculation method, the acquisition unit 112-n restores the information GM from a plurality of different confidential information [GM] m(1) , …, [GM] m(max) (where {m(1), …, m(max)} ⊆ {1, …, M}). When the secret calculation method is a homomorphic encryption method, the acquisition unit 112-n decrypts the confidential information [GM] m using the decryption key read from the storage unit 111-n to obtain the information GM. The information GM that identifies the aggregated model is stored in the storage unit 111-n (step S112-n).
[0028] [An example of asynchronous control (step S1331)] The processing of the federated learning system 1 by asynchronous control is exemplified. Note that the following asynchronous control is an example and does not limit the present invention. In asynchronous control, each model learning device 11-n executes local processing asynchronously with other model learning devices 11-n' (where n, n' ∈ {1, …, N}) included in the federated learning system 1. That is, the plurality of model learning devices 11-1, …, 11-N execute local processing asynchronously with each other. When asynchronous control is performed, the control unit 133 instructs each model learning device 11-n (Fig. 2) to execute local processing by asynchronous control. This instruction is acquired by the acquisition unit 112-n and sent to the control unit 116-n. The control unit 116-n executes asynchronous control.
[0029] In asynchronous control, each model learning device 11-n determines whether it is necessary to update the acquired aggregated model to obtain a new worker model. If the model learning device 11-n determines that this is necessary, it updates the aggregated model to obtain a new worker model. However, if it determines that it is not necessary, it does not update the aggregated model to obtain a new worker model, and after the waiting time has elapsed, it acquires the confidential information [GM] of the information GM for specifying a new aggregated model from each secret federated learning device 12-m. m Furthermore, in asynchronous control, each secret federated learning device 12-m determines whether it has obtained the confidential information [WM-n] of the information WM-n for specifying a worker model from a predetermined model learning device 11-n. It is not necessarily required to have obtained the confidential information [WM-n] of the information WM-n for specifying a worker model from all model learning devices 11-n. When each secret federated learning device 12-m determines that it has obtained the confidential information [WM-n] of the information WM-n for specifying a worker model, it uses the confidential information [WM-n] of the information WM-n for specifying a worker model to obtain the confidential information [GM] of the information GM for specifying an aggregated model obtained by aggregating worker models through secret calculation. The following shows a specific example. m m m m m m m m m m
[0030] When performing asynchronous processing, each model learning device 11-n (Fig. 2) executes the processes of steps S113-n, S114-n, and S115-n described above, and the secret federated learning device 12-m (Fig. 3) executes the process of step S121-m. However, in asynchronous processing, in step S115-n, the providing unit 115-n of the model learning device 11-n does not transmit the synchronization information to the control device 13.
[0031] Also, the determination unit 128-m of each secret federated learning device 12-m refers to the secret information [WM-n] stored in the storage unit 127-m m at a predetermined timing to determine whether the registration of the worker model is completed. For example, the determination unit 128-m may perform the determination periodically, or may perform the determination triggered by the fact that each piece of secret information [WM-n] m has been stored in the storage unit 127-m. Note that the completion of the registration of the worker model means that the secret information [WM-n1],..., [WM-n] max (where {n1,..., n max} = {1,..., N}) from a new worker model m ,..., [WM-n max m has been obtained. This does not necessarily mean that all of the secret information [WM-1],..., [WM-N] m of the information identifying the new worker model has been obtained. That is, the determination unit 128-m determines whether the secret information [WM-n1],..., [WM-n m has been obtained from at least some of the model learning devices 11-n1,..., 11-n max from , the secret information [WM-n1],..., [WM-n m ,..., [WM-n ma x m identifying the new worker model. The model learning devices 11-n1,..., 11-n max are the model learning devices 11-1,..., 11-N (i.e., {n1,..., n max} = {1,..., N}) that have already provided the secret information identifying the new worker model at the time of determination Alternatively, some pre-set model learning devices 11-n1, …, 11-n max (i.e., {n1, …, n max} ⊂ {1, …, N}) may also be used. Additionally, the confidential information [WM-n1] , …, [WM-n m max for identifying the new worker model m may be the confidential information [WM-n1] , …, [WM-n m max that has not yet been used in the secret aggregation process m , or it may be the confidential information [WM-n1] , …, [WM-n m max obtained after the previous secret aggregation process. However, since the confidential information [WM-n] m is in the form of shares of a secret sharing scheme or ciphertexts of a homomorphic encryption scheme, etc., it may not be possible to identify the model learning device 11-n that provided the confidential information [WM-n] m . In such a case, the determination unit 128-m may determine whether the registration of the worker model is complete based on the total data volume of the confidential information [WM-n] stored in the storage unit 127-m. For example, when the total data volume of the confidential information [WM-n] m stored in the storage unit 127-m matches the total data volume of the confidential information [W m M-n1] m , …, [WM-n m provided by the predetermined model learning devices 11-n1, …, 11-n max , the determination unit 128-m may determine that the registration of the worker model is complete, and if not, determine that the registration of the worker model is not complete. Alternatively, the determination unit 128-m may determine whether the total number of worker models corresponding to the confidential information [WM-n] m , …, [WM-n max m stored in the storage unit 127-m matches the total number of worker models provided by the predetermined model learning devices 11-n1, …, 11-n m max Total number of worker models n max If the worker model registration is completed, it is determined that the worker model registration is complete. If the worker model registration is not completed, it is determined that the worker model registration is not complete. For example, the information WM-n for identifying the worker model is a model parameter group, and the number of model parameters included in one worker model is N MP and the secret information [WM-n] stored in the memory unit 127-m m The total amount of data is N records. R If so, confidential information [WM-n] m The total number of worker models provided is N R / N MP In this case, the decision unit 128-m determines whether n max =N R / N MP The worker model registration is completed when If not, it may be determined that the registration of the worker model is not complete. Here, if it is determined that the registration of the worker model is not complete and a predetermined time has not elapsed since the reference time point (timeout has not occurred), the determination unit 128-m determines again at a predetermined opportunity whether the registration of this worker model is complete. For example, the determination unit 128-m may perform the determination again after a predetermined waiting time has elapsed, or may perform the determination again after any of the secret information [WM-n] has elapsed. m The determination may be performed again using the fact that the worker model registration is completed or that a predetermined time has elapsed (timeout) since the reference time point as a trigger. On the other hand, when it is determined that the worker model registration is completed or that a predetermined time has elapsed (timeout) since the reference time point, the determination unit 128-m sends an instruction to the control unit 126-m to start the secret aggregation process. An example of the reference time point for timeout is as explained in the above-mentioned synchronization control (step S128-m).
[0032] Upon receiving an instruction to start the secret aggregation process, the control unit 126-m instructs the secret aggregation processing unit 122-m to start the secret aggregation process. In response to this, the secret aggregation processing unit 122-m reads a plurality of confidential information [WM-n] (where n ∈ {1, …, N}) (confidential information of information identifying a plurality of worker models) from the storage unit 127-m, and by performing secret calculations using these, without obtaining the plurality of worker models, identifies the confidential information [GM] of the aggregation model that aggregates the plurality of worker models. m and outputs it. That is, when it is determined that the secret aggregation processing unit 122-m has obtained the confidential information [WM-n1] max from the predetermined model learning devices 11-n1, …, 11-n …, [WM-n m max m and has obtained the confidential information [WM-n1] …, [WM-n m max m of the information identifying the worker model, it aggregates the plurality of worker models by performing secret calculations using the confidential information [WM-n1] …, [WM-n m m of the information identifying the worker model, obtains the confidential information [GM] of the aggregation model that aggregates the plurality of worker models, and outputs it. The confidential information [GM] of the information identifying the aggregation model is sent to the providing unit 123-m (step S122’-m).
[0033] The acquisition unit 112-n of the model learning device 11-n (Fig. 2) accesses the providing unit 123-m of the secret federated learning device 12-m (where m ∈ {1, …, M}) (Fig. 3) at a predetermined timing, and acquires the confidential information [GM] of the information identifying the aggregation model from the providing unit 123-m. m The acquisition unit 112-n restores the acquired confidential information [GM] m to obtain the information GM that identifies the aggregation model. The information GM that identifies the aggregation model is stored in the storage unit 111-n (step S112’-n).
[0034] The determination unit 117-n determines whether it is necessary to update the aggregation model corresponding to the information GM stored in the storage unit 111-n to newly obtain a worker model. In other words, the determination unit 117-n determines whether it is necessary to update the latest aggregation model specified by the information GM by machine learning using the local learning data D-n to obtain a worker model. For example, when the aggregation model is the same as or approximated to the "latest aggregation model" that has already been used for generating the worker model (step S113-n), the determination unit 117-n determines that it is not necessary to update the aggregation model to newly obtain a worker model, and otherwise determines that it is necessary to update the aggregation model to newly obtain a worker model. Note that two aggregation models being approximated may mean, for example, that the distance between their model parameters is less than or equal to a predetermined value, or that the difference in the output distributions of the two aggregation models for a predetermined input group is less than or equal to a predetermined value (step S117a-n).
[0035] Here, when it is determined that it is not necessary to update the aggregation model to newly obtain a worker model, after the elapse of the waiting time without the learning unit 113-n updating the aggregation model to newly obtain a worker model, the acquisition unit 112-n acquires the confidential information [GM] of the information for specifying a new aggregation model from the secret joint learning device 12-m (where m ∈ {1,..., M}) (Fig. 3). m That is, without the learning unit 113-n obtaining a new worker model, after the elapse of the waiting time, the acquisition unit 112-n accesses the providing unit 123-m and acquires the confidential information [GM] of the information GM for specifying the aggregation model from the providing unit 123-m. m The acquisition unit 112-n restores the acquired confidential information [GM] m to obtain the information GM for specifying the aggregation model, stores this in the storage unit 111-n, and returns to step S117a-n (step S117b-n).
[0036] On the other hand, when it is determined that it is necessary to update the aggregated model to newly obtain a worker model, the learning unit 113-n reads the local learning data D-n stored in the storage unit 111-n and the latest information GM, and updates the latest aggregated model specified by the information GM by machine learning using the local learning data D-n to obtain a worker model, and outputs information WM-n for specifying the worker model (step S113-n). Thereafter, the processes after step S114-n described so far in this embodiment are executed again.
[0037] <Features of this embodiment> In this embodiment, the model learning device executes local processing of obtaining information for specifying a worker model by updating an aggregated model by machine learning using local learning data, and providing secret information of the information for specifying the worker model to the secret federated learning device. Further, the secret federated learning device obtains secret information of information for specifying a new aggregated model obtained by aggregating a plurality of worker models without obtaining the worker models, by performing secret calculation using the secret information of the information for specifying the obtained worker models, and executes secret aggregation processing of providing the information for specifying the new aggregated model or the secret information of the information for specifying the new aggregated model to a plurality of model learning devices. The control device compares the local processing time corresponding to the local processing with the aggregation processing time corresponding to the secret aggregation processing, and when the local processing time is longer than the aggregation processing time, or when the local processing time is equal to or longer than the aggregation processing time, performs asynchronous control to execute the local processing of a plurality of model learning devices asynchronously with each other, and when not performing asynchronous control, performs synchronous control to synchronize the local processing of a plurality of model learning devices with each other. In this case, since the secret federated learning device cannot obtain the worker model itself, it is also impossible to know the tendency of the learning data held by each model learning device based on the difference between the worker model and the aggregated model. Thereby, the security of federated learning can be improved. Furthermore, since synchronous control and asynchronous control are switched according to the relationship between the local processing time and the aggregation processing time, the processing of the entire federated learning system can be made efficient according to the processing speed of each model learning device.
[0038] [Second Embodiment] The second embodiment is a modification of the first embodiment. In the synchronization control of this embodiment, without using the confidential information of the information identifying the worker models whose contribution degree to the new aggregation model is less than or equal to the reference value, the confidential information of the information identifying the worker models whose contribution degree is greater than or equal to or exceeds the reference value is used to cause the secret aggregation processing device to execute the secret aggregation processing. Thereby, even in the synchronization process, the secret aggregation learning device can execute the secret aggregation process without waiting for the confidential information of the information identifying the worker models with low contribution degrees. The merit of starting the secret aggregation process earlier is greater than the merit of waiting for the confidential information of the information identifying the worker models with low contribution degrees, and the overall processing can be made more efficient. Hereinafter, the same reference numerals as those described so far will be used, and the description will be simplified.
[0039] <Configuration> As illustrated in FIG. 1, the federated learning system 2 of this embodiment includes N model learning devices 11-1,..., 11-N that perform model learning, M secret federated learning devices 22-1,..., 22-M that perform federated learning by secret calculation, and a control device 23 that controls the federated learning system 2.
[0040] As illustrated in FIG. 3, the secret federated learning device 22-m of this embodiment includes an acquisition unit 121-m, a secret aggregation processing unit 222-m, a provision unit 123-m, a control unit 126-m, a storage unit 127-m, and a determination unit 128-m. The secret federated learning device 22-m executes each process based on the control of the control unit 126-m, and the input information and the information obtained in each process are stored in the storage unit 127-m and read out and used as necessary.
[0041] As illustrated in FIG. 4, the control device 23 of this embodiment includes a measurement unit 131, a comparison unit 132, and a control unit 233.
[0042] <Preprocessing> The same as the first embodiment.
[0043] <Learning Process> The learning process of this embodiment is illustrated using FIG. 5. Also in this embodiment, the local process and the secret aggregation process are alternately repeated until the end conditions are satisfied. At this time, the new aggregation model corresponding to the confidential information [GM] obtained in the secret aggregation process is used as the aggregation model in the next local process to be executed. m
[0044] The measurement unit 231 (FIG. 4) of the control device 23 measures the local processing time T1 corresponding to the local process and the aggregation processing time T2 corresponding to the secret aggregation process in the federated learning system 2. The measured local processing time T1 and aggregation processing time T2 are sent to the comparison unit 132 (step S131).
[0045] The comparison unit 132 of the control device 23 compares the sent local processing time T1 and aggregation processing time T2 (step S132). Here, when the local processing time T1 is longer than the aggregation processing time T2 (T1 > T2), or when the local processing time T1 is greater than or equal to the aggregation processing time T2 (T1 ≧ T2), the control unit 233 performs asynchronous control to execute the local processes of the plurality of model learning devices 11-1,..., 11-N asynchronously with each other (step S1331). Then, it returns to step S131. On the other hand, when not performing asynchronous control, the control unit 233 performs synchronous control to synchronize the local processes of the plurality of model learning devices 11-1,..., 11-N with each other (step S1332). Further, in this embodiment, when performing synchronous control, the control unit 233 uses the confidential information of the information identifying the worker models whose contribution degree to the new aggregation model is less than or equal to the reference value without using the confidential information of the information identifying the worker models whose contribution degree to the new aggregation model is greater than or equal to the reference value or exceeds the reference value. A secret aggregation process (hereinafter referred to as "selective synchronous aggregation process") is executed on the secret federated learning devices 22-1,..., 22-M. The contribution degree of the worker model to the new aggregation model may be any index. For example, the Shapley value (for example, the index used in FedCoine) may be used as the contribution degree (step S233).
[0046] The asynchronous control (step S1331) of this embodiment is the same as the asynchronous control of embodiment 1. An example of the synchronous control (steps S1332, S233) involving the selective synchronous aggregation process of this embodiment will be described below.
[0047] [An example of synchronous control involving selective synchronous aggregation processing (steps S1332, S233)] The control unit 233 instructs each model learning device 11-n (FIG. 2) to execute local processing under synchronous control. Each model learning device 11-n that has been instructed to execute local processing under synchronous control executes steps S113-n and S114-n described in the first embodiment.
[0048] The provider 115-n stores confidential information [WM-n]={[WM-n]1, ..., [WM-n] of information WM-n for identifying a worker model. M The providing unit 115-n receives the confidential information [WM-n] of the information WM-n that identifies the worker model. m to the secret associative learning device 22-m (FIG. 3) (where m=1, . . . , M). Furthermore, the providing unit 115-n transmits (provides) the secret information [WM-n] to the secret associative learning device 12-m when the model learning device 11-n transmits the secret information [WM-n] to the secret associative learning device 12-m. m The synchronization information indicating that the command has been transmitted is sent to the control device 23 (step S115-n).
[0049] The acquisition unit 121-m of the secret association learning device 22-m (FIG. 3) acquires the worker model-specific information WM-n sent from the model learning device 11-n. m Receive the confidential information [WM-n] m is stored in storage unit 127-m (step S121-m).
[0050] The control unit 233 of the control device 23 that executes the selective synchronous aggregation process calculates the contribution of the worker model obtained by each model learning device 11-n (where n=1, . . . , N) to the aggregation model. For example, the control unit 233 receives information WM-n for identifying the worker model and information GM for identifying the aggregated model obtained by aggregating the worker models from each model learning device 11-n one by one or periodically, and calculates the contribution degree. Alternatively, each secret association learning device 22-m obtains the secret information [WM-n] of the information WM-n for identifying the worker model. m and the confidential information [WM-n] m Confidential information of GM, which identifies the aggregate model obtained using m may be used to calculate the secret information of the contribution degree by secret computation and transmit it to the control device 23. In this case, the control unit 233 restores the contribution degree from the secret information of the contribution degree sent from each secret associative learning device 22-m. Based on the obtained contribution degree and synchronization information, the control unit 233 calculates the secret information [WM-n]1, ..., [WM-n] corresponding to the worker model whose contribution degree is equal to or exceeds the reference value. M It is determined whether the secret information [WM-n]1, ..., [WM-n] corresponding to the worker model whose contribution degree is equal to or exceeds the reference value is transmitted to all the secret association learning devices 22-1, ..., 22-M (step S2332a). M has not been transmitted to all the secret association learning devices 22-1, ..., 22-M, and it is determined that a predetermined time has not elapsed since the reference time point (timeout has not occurred), the control unit 233 performs the determination in step S2332a at certain intervals. On the other hand, the secret information [WM-n]1, ..., [WM-n] corresponding to the worker model whose contribution degree is equal to or exceeds the reference value is M has been transmitted to all the secret association learning devices 22-1, ..., 22-M, or if it is determined that a predetermined time has elapsed (timed out) since the reference point in time, the control unit 233 sends an instruction to start the secret aggregation process to the secret association learning devices 22-1, ..., 22-M (S2332b).
[0051] An instruction to start the secret aggregation process is received by the acquisition unit 121-m of the secret federated learning device 22-m (where m = 1, …, M) (Fig. 3) and input to the control unit 126-m. The control unit 126-m that has received the instruction to start the secret aggregation process instructs the secret aggregation processing unit 222-m to start the secret aggregation process. In response to this, the secret aggregation processing unit 222-m reads a plurality of confidential information [WM-n] (where n ∈ {1, …, N}) (confidential information of information identifying a plurality of worker models) from the storage unit 127-m, and by performing secret calculations using these, without obtaining the plurality of worker models, identifies the confidential information [GM] of the information GM that identifies the aggregated model obtained by aggregating the plurality of worker models. m and outputs it. The confidential information [GM] of the information GM that identifies the aggregated model m is sent to the providing unit 123-m (step S222-m).
[0052] The confidential information [GM] m is input to the providing unit 123-m. The providing unit 123-m transmits (provides) the confidential information [GM] m to a plurality of model learning devices 11-n (where n ∈ {1, …, N}) via the control device 13 (step S123-m).
[0053] The confidential information [GM] m (where m ∈ {1, …, M}) sent to the acquisition unit 112-n of the model learning device 11-n (Fig. 2) receives the confidential information [GM] m (confidential information of the information GM that identifies the aggregated model provided by the secret federated learning device 22-m). The acquisition unit 112-n restores the confidential information [GM] m to obtain the information GM that identifies the aggregated model. The information GM that identifies the aggregated model is stored in the storage unit 111-n (step S112-n).
[0054] <Features of this embodiment> In this embodiment, the same effects as those of the first embodiment can be obtained. Furthermore, in the synchronization control of this embodiment, without using the confidential information of the information for identifying the worker models whose contribution degree to the new aggregation model is less than or equal to the reference value, the confidential information of the information for identifying the worker models whose contribution degree is greater than or equal to or exceeds the reference value is used to cause the secret aggregation processing device to execute the secret aggregation processing. Thereby, the efficiency can be further improved.
[0055] [Third Embodiment] The third embodiment is a modification of the first embodiment. In this embodiment, when performing asynchronous control and the local processing time is not longer than the aggregation processing time by a predetermined time or more, the secret aggregation processing device is caused to execute the secret aggregation processing without considering the contribution degree. In such a case, although there is more time margin on the side of the secret aggregation learning device than on the side of the model learning device, the difference is not so large. In such a case, it is more efficient as a whole to execute the secret aggregation processing without considering the contribution degree of the confidential information of the information for identifying the worker models. On the other hand, when performing asynchronous control and the local processing time is longer than the aggregation processing time by a predetermined time or more, without using the confidential information of the information for identifying the worker models whose contribution degree is less than or equal to the reference value, the confidential information of the information for identifying the worker models whose contribution degree is greater than or equal to or exceeds the reference value is used to cause the secret aggregation processing device to execute the secret aggregation processing. In such a case, the secret aggregation learning device side has a significantly larger time margin than the model learning device side. In such a case, the merit of starting the secret aggregation processing earlier is greater than the merit of waiting for the confidential information of the information for identifying the worker models with a low contribution degree, and the processing can be made more efficient as a whole.
[0056] <Configuration> As illustrated in FIG. 1, the federated learning system 3 of this embodiment includes N model learning devices 11-1,..., 11-N that perform model learning, M secret aggregation learning devices 32-1,..., 32-M that perform federated learning by secret calculation, and a control device 33 that controls the federated learning system 3.
[0057] As illustrated in FIG. 3, the secret federated learning device 32-m of this embodiment includes an acquisition unit 121-m, a secret aggregation processing unit 322-m, a provision unit 123-m, a control unit 126-m, a storage unit 127-m, and a determination unit 128-m. The secret federated learning device 32-m executes each process based on the control of the control unit 126-m, and the input information and the information obtained in each process are stored in the storage unit 127-m and read out and used as needed.
[0058] As illustrated in FIG. 4, the control device 33 of this embodiment includes a measurement unit 131, a comparison unit 332, and a control unit 333.
[0059] <Preprocessing> It is the same as the first embodiment.
[0060] <Learning process> The learning process of this embodiment is illustrated with reference to FIG. 5. In this embodiment as well, the local process and the secret aggregation process are alternately repeated until the end conditions are satisfied. At this time, the new aggregation model corresponding to the secret information [GM] obtained in the secret aggregation process is used as the aggregation model in the next local process to be executed. m
[0061] The comparison unit 332 of the control device 33 compares the sent local processing time T1 and the aggregation processing time T2 (step S132). Here, when the local processing time T1 is longer than the aggregation processing time T2 (T1 > T2), or when the local processing time T1 is greater than or equal to the aggregation processing time T2 (T1 ≧ T2), the control unit 333 performs asynchronous control to execute the local processes of the plurality of model learning devices 11-1,..., 11-N asynchronously with each other (step S1331). In this embodiment, when performing asynchronous control, the comparison unit 332 further determines whether the local processing time T1 is longer than the aggregation processing time T2 by a predetermined time A or more (that is, T1 ≧ T2 + Determine whether it is A). Here, A is a positive real number representing time. The comparison unit 332 may compare T1 with T2 + A and determine whether T1 ≥ T2 + A is satisfied, or compare T1 with a value α smaller than T2 + A and determine whether T1 > α is satisfied (step S332).
[0062] Here, when it is determined that the local processing time T1 is not longer than the predetermined time A or more than the aggregation processing time T2 (T2 < T1 < T2 + A or T2 ≤ T1 < T2 + A), the control unit 333 causes the secret aggregation processing that does not consider the contribution degree to be executed on the secret federated learning devices 32-1,..., 32-M. In such a case, although the secret federated learning devices 32-1,..., 32-M side has more time margin than the model learning devices 11-1,..., 11-N side, the difference is not so large. In such a case, it is more efficient as a whole to execute the secret aggregation processing without considering the contribution degree of the confidential information of the information specifying the worker model. Then, the process returns to step S131 (step S3331). On the other hand, when it is determined that the local processing time T1 is longer than the predetermined time A or more than the aggregation processing time T2 (T1 ≥ T2 + A), the control unit 333 executes, on the secret federated learning devices 32-1,..., 32-M, a secret aggregation process (hereinafter referred to as "selective asynchronous aggregation process") that uses the confidential information of the information specifying the worker model with a contribution degree equal to or higher than the reference value or exceeding the reference value without using the confidential information of the information specifying the worker model with a contribution degree lower than the reference value or equal to or lower than the reference value. In such a case, the secret federated learning devices 32-1,..., 32-M side has a significantly larger time margin than the model learning devices 11-1,..., 11-N side. In such a case, the merit of starting the secret aggregation process earlier is greater than the merit of waiting for the confidential information of the information specifying the worker model with a low contribution degree, and the processing can be made more efficient as a whole. Then, the process returns to step S131. (Step S3332).
[0063] On the other hand, when asynchronous control is not performed, the control unit 233 performs synchronization control to synchronize the local processing of the plurality of model learning devices 11-1, …, 11-N with each other (step S1332). Then, the process returns to step S131.
[0064] The synchronization control (step S1332) of the present embodiment is the same as the synchronization control of the first embodiment. Also, the secret aggregation process (step S3331) that does not consider the contribution degree in the asynchronous control is the same as the secret aggregation process in the asynchronous control described in the first embodiment. Hereinafter, an example of the selective asynchronous aggregation process (step S3332) of the present embodiment will be shown.
[0065] [Example of Selective Asynchronous Aggregation Process (Step S3332)] The control unit 333 instructs each model learning device 11-n (FIG. 2) to execute local processing by asynchronous control. Each model learning device 11-n instructed to execute local processing by synchronization control executes the processes of steps S113-n, S114-n, and S115-n described in the first embodiment, and the secret federated learning device 32-m (FIG. 3) executes the process of step S121-m. However, in the asynchronous process, in step S115-n, the providing unit 115-n of the model learning device 11-n does not transmit the synchronization information to the control device 33.
[0066] The control unit 333 of the control device 33 that executes the selective asynchronous aggregation process obtains the contribution degree of each model learning device 11-n (where n = 1, …, N) to the aggregated model of the worker models obtained. For example, the control unit 333 obtains the contribution degree by the method exemplified in the second embodiment. The contribution degree corresponding to each obtained secret information [WM-n] m is sent to each secret federated learning device 32-m. The contribution degree corresponding to each secret information [WM-n] m is received by the acquisition unit 121-m of each secret federated learning device 32-m and sent to the determination unit 128-m. The determination unit 128-m of each secret federated learning device 32-m, at a predetermined timing, the secret information [WM-n] stored in the storage unit 127-m mRefer to , and determine whether the registration of the worker model is completed. In the selective asynchronous aggregation process, the completion of the registration of the worker model means that all the confidential information [WM-n]1,…,[WM-n] corresponding to the worker models with a contribution degree equal to or higher than the reference value has been obtained. M Here, if it is determined that the registration of the worker model is not completed and that a predetermined time has not elapsed since the reference time point (i.e., it has not timed out), the determination unit 128-m will, on a predetermined occasion, determine again whether the registration of this worker model is completed. On the other hand, if it is determined that the registration of the worker model is completed or that a predetermined time has elapsed since the reference time point (i.e., it has timed out), the determination unit 128-m will send an instruction to the control unit 126-m to start the confidential aggregation process (step S328-m).
[0067] Upon receiving the instruction to start the confidential aggregation process, the control unit 126-m instructs the confidential aggregation processing unit 122-m to start the confidential aggregation process. In response to this, the confidential aggregation processing unit 122-m reads a plurality of pieces of confidential information [WM-n] (where n ∈ {1,…,N}) (confidential information of information identifying a plurality of worker models) from the storage unit 127-m, and through confidential calculation using these, without obtaining the plurality of worker models, identifies the confidential information [GM] of the information GM that identifies the aggregated model obtained by aggregating the plurality of worker models. m and outputs it. The confidential information [GM] of the information GM that identifies the aggregated model m is sent to the providing unit 123-m (step S322-m). Thereafter, the processing after step S112'-n of the asynchronous control of the first embodiment is executed.
[0068] <Features of this embodiment> In this embodiment, the same effects as those of the first embodiment can be obtained. Furthermore, in the asynchronous control of this embodiment, when the local processing time is not longer than the predetermined time or more than the aggregation processing time, the secret aggregation processing that does not consider the contribution degree is executed by the secret federated learning device. On the other hand, in the asynchronous control, when the local processing time is longer than the predetermined time or more than the aggregation processing time, the secret aggregation processing that uses the secret information of the information for identifying the worker model whose contribution degree is equal to or higher than the reference value without using the secret information of the information for identifying the worker model whose contribution degree is less than the reference value or less than or equal to the reference value is executed by the secret federated learning device. Thereby, the efficiency can be further improved.
[0069] [Modification Example 1 of the Third Embodiment] The second embodiment may be combined with the third embodiment. That is, the following processing may be performed. In the synchronous control, as described in the second embodiment, the secret aggregation processing that uses the secret information of the information for identifying the worker model whose contribution degree to the new aggregation model is less than the reference value or less than or equal to the reference value is executed by the secret federated learning device without using the secret information of the information for identifying the worker model whose contribution degree is equal to or higher than the reference value. Then, in the asynchronous control, when the local processing time is not longer than the predetermined time or more than the aggregation processing time, the secret aggregation processing that does not consider the contribution degree is executed by the secret federated learning device. On the other hand, in the asynchronous control, when the local processing time is longer than the predetermined time or more than the aggregation processing time, the secret aggregation processing that uses the secret information of the information for identifying the worker model whose contribution degree is equal to or higher than the reference value without using the secret information of the information for identifying the worker model whose contribution degree is less than the reference value or less than or equal to the reference value is executed by the secret federated learning device. Thereby, the efficiency can be further improved.
[0070] [Fourth Embodiment] In the secret aggregation processing of the above-described first, second, and third embodiments and Modification Example 1 of the third embodiment, the secret federated learning devices 12-m, 22-m, and 32-m identify the secret information [GM] of the information GM for a new aggregation model obtained by aggregating a plurality of worker models without obtaining the worker models by secret calculation using the obtained secret information [WM-n].m obtain the confidential information [GM] of information GM that identifies a new aggregated model m and provide it to a plurality of model learning devices 11-n. However, in this secret aggregation process, if any of the secret federated learning devices provides the confidential information [GM] of information GM that identifies a new aggregated model m instead, the information GM that identifies a new aggregated model may be provided to the model learning devices 11-1, …, 11-N. In this case, the secret federated learning device restores the information GM that identifies a new aggregated model from the confidential information [GM] of information GM that identifies a plurality of new aggregated models and provides it to the model learning devices 11-1, …, 11-N.
[0071] [Hardware Configuration] The model learning devices 11-m, 14-m, the federated learning devices 12-m, 22-m, 32-m, and the control devices 13, 23, 33 in each embodiment are general-purpose or dedicated computers equipped with, for example, a processor (hardware processor) such as a CPU (central processing unit) and a memory such as a RAM (random-access memory) and a ROM (read-only memory). That is, the model learning devices 11-m, 14-m, the federated learning devices 12-m, 22-m, 32-m, and the control devices 13, 23, 33 in each embodiment have, for example, a processing circuitry configured to implement each part they have. This computer may be equipped with one processor and memory, or may be equipped with a plurality of processors and memory. This program may be installed in the computer or may be recorded in a ROM or the like in advance. Also, instead of an electronic circuit that realizes a functional configuration by loading a program like a CPU, a part or all of the processing units may be configured using an electronic circuit that realizes a processing function alone. Also, the electronic circuit that constitutes one device may include a plurality of CPUs.
[0072] FIG. 6 is a block diagram illustrating the hardware configurations of the model learning devices 11-m, 14-m, the federated learning devices 12-m, 22-m, 32-m, and the control devices 13, 23, 33 in each embodiment. As illustrated in FIG. 6, the model learning devices 11-m, 14-m, the federated learning devices 12-m, 22-m, 32-m, and the control devices 13, 23, 33 in this example include a CPU (Central Processing Unit) 10a, an input unit 10b, an output unit 10c, a RAM (Random Access Memory) 10d, a ROM (Read Only Memory) 10e, and an auxiliary storage device 10f , a communication unit 10h, and a bus 10g. The CPU 10a in this example includes a control unit 10aa, an arithmetic unit 10ab, and a register 10ac, and executes various arithmetic processes according to various programs read into the register 10ac. The input unit 10b is an input terminal to which data is input, a keyboard, a mouse, a touch panel, or the like. The output unit 10c is an output terminal to which data is output, a display, or the like. The communication unit 10h is a LAN card or the like controlled by the CPU 10a that has read a predetermined program. The RAM 10d is an SRAM (Static Random Access Memory), a DRAM (Dynamic Random Access Memory), or the like, and includes a program area 10da in which a predetermined program is stored and various data It has a data area 10db to be stored. The auxiliary storage device 10f is, for example, a hard disk, an MO (Magneto-Optical disc), a semiconductor memory, etc., and has a program area 10fa where a predetermined program is stored and a data area 10fb where various data are stored. The bus 10g connects the CPU 10a, the input unit 10b, the output unit 10c, the RAM 10d, the ROM 10e, the communication unit 10h, and the auxiliary storage device 10f so that information can be exchanged. The CPU 10a writes the program stored in the program area 10fa of the auxiliary storage device 10f to the program area 10da of the RAM 10d according to the loaded OS (Operating System) program. Similarly, the CPU 10a writes various data stored in the data area 10fb of the auxiliary storage device 10f to the data area 10db of the RAM 10d. Then, the address on the RAM 10d where this program and data are written is stored in the register 10ac of the CPU 10a. The control unit 10aa of the CPU 10a sequentially reads out these addresses stored in the register 10ac, reads out the program and data from the area on the RAM 10d indicated by the read address, sequentially causes the arithmetic unit 10ab to execute the arithmetic indicated by the program, and stores the arithmetic result in the register 10ac. With such a configuration, the functional configurations of the model learning devices 11-m, 14-m, the collaborative learning devices 12-m, 22-m, 32-m, and the control devices 13, 23, 33 are realized.
[0073] The above program can be recorded on a computer-readable recording medium. Examples of computer-readable recording media are non-transitory recording media. Examples of such recording media are magnetic recording devices, optical discs, magneto-optical recording media, semiconductor memories, etc.
[0074] The distribution of this program is carried out, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Further, the program may be stored in a storage device of a server computer and transferred from the server computer to other computers via a network to distribute the program. As described above, a computer that executes such a program first stores, for example, a program recorded on a portable recording medium or a program transferred from a server computer in its own storage device. Then, at the time of executing the process, the computer reads the program stored in its own storage device and executes the process according to the read program. Also, as another execution form of this program, the computer may directly read the program from the portable recording medium and execute the process according to the program. Further, each time a program is transferred from the server computer to this computer, the computer may sequentially execute the process according to the received program. Also, the transfer of the program from the server computer to this computer may not be performed, and the above-described process may be executed by a so-called ASP (Application Service Provider) type service that realizes the processing function only by the execution instruction and result acquisition. Note that the program in this embodiment includes information used for processing by an electronic computer and similar to the program (data having a nature of defining the processing of the computer although not a direct instruction to the computer).
[0075] In each embodiment, the present apparatus is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized hardware-wise.
[0076] Note that the present invention is not limited to the above-described embodiments. For example, the various processes described above may be executed not only in time series according to the description, but also in parallel or individually according to the processing ability of the device that executes the processes, or as necessary. Needless to say, various modifications can be made without departing from the spirit of the present invention.
Explanation of Signs
[0077] 1, 2, 3 federated learning system 11-m, 14-m model learning device 111-n storage unit 112-n acquisition unit 113-n learning unit 114-n concealment unit 115-n providing unit 117-n determination unit 12-m, 22-m, 32-m 121-m acquisition unit 122-m, 222-m, 322-m secret aggregation processing unit 123-m providing unit 127-m storage unit 128-m determination unit 13, 23, 33 control device 131 measurement unit 132, 332 comparison unit 133, 233, 333 control unit
Claims
1. A control device for controlling a federated learning system including a plurality of model learning devices and one or more secure federated learning devices, wherein the model learning device executes local processing to obtain information for identifying a worker model by updating an aggregated model through machine learning using local learning data, and provides secret information of the information for identifying the worker model to the secure federated learning device, the secure federated learning device obtains secret information of information for identifying a new aggregated model obtained by aggregating a plurality of the worker models without obtaining the worker models, by performing secure computation using the obtained secret information of the information for identifying the worker model, and executes secure aggregation processing to provide the information for identifying the new aggregated model or the secret information of the information for identifying the new aggregated model to the plurality of model learning devices, the control device includes a comparison unit that compares a local processing time corresponding to the local processing with an aggregation processing time corresponding to the secure aggregation processing, and a control unit that performs asynchronous control to execute the local processing of the plurality of model learning devices asynchronously with each other when the local processing time is longer than the aggregation processing time, or when the local processing time is equal to or longer than the aggregation processing time, and performs synchronous control to synchronize the local processing of the plurality of model learning devices with each other when the asynchronous control is not performed. A control device having the above components.
2. The control device according to Claim 1, wherein the control unit, when performing the synchronous control, causes the secure federated learning device to execute the secure aggregation processing that uses the secret information of the information for identifying a worker model having a contribution degree to the new aggregated model equal to or higher than a reference value, without using the secret information of the information for identifying a worker model having a contribution degree lower than the reference value. A control device.
3. The control device according to Claim 2, wherein the control unit, when performing the asynchronous control and when the local processing time is not longer than the aggregation processing time by a predetermined time or more, causes the secure federated learning device to execute the secure aggregation processing that does not consider the contribution degree. A control device that causes the secret aggregation processing to be executed by the secret federated learning device, without using the confidential information of the information that identifies the worker model whose contribution degree is less than or equal to the reference value, when performing the asynchronous control and when the local processing time is longer than the aggregation processing time by a predetermined time or more, and using the confidential information of the information that identifies the worker model whose contribution degree is greater than or equal to the reference value or exceeds the reference value.
4. A model learning device of a federated learning system including a plurality of model learning devices and one or more secret federated learning devices, A storage unit that stores local learning data, An acquisition unit that obtains information for identifying an aggregated model or confidential information of the information for identifying the aggregated model from the secret federated learning device, A learning unit that updates the aggregated model by machine learning using the local learning data to obtain a worker model, A confidentiality unit that obtains confidential information of the information for identifying the worker model, And a providing unit that provides the confidential information of the information for identifying the worker model to the secret federated learning device, The local processing is a process in which the model learning device updates the aggregated model by machine learning using the local learning data to obtain information for identifying the worker model, and provides the confidential information of the information for identifying the worker model to the secret federated learning device. The local processing time is the time corresponding to the local processing. The secret aggregation processing is a process in which, by secret calculation using the confidential information of the information for identifying the worker model acquired by the secret federated learning device, without obtaining the worker model, the confidential information of the information for identifying a new aggregated model obtained by aggregating a plurality of the worker models is obtained, and the information for identifying the new aggregated model or the confidential information of the information for identifying the new aggregated model is provided to the plurality of model learning devices. The aggregation processing time is the time corresponding to the secret aggregation processing. The model learning device, When the local processing time is longer than the aggregation processing time or the local processing time is greater than or equal to the aggregation processing time, an instruction to execute the local processing by asynchronous control performed by the control device is acquired, and the local processing is executed asynchronously with other model learning devices included in the federated learning system. When not executing the local processing asynchronously, it acquires an instruction to execute the local processing by synchronous control performed by the control device, and executes the local processing synchronized with other model learning devices included in the federated learning system. A model learning device.
5. A secret federated learning device of a federated learning system including a plurality of model learning devices and one or more secret federated learning devices, An acquisition unit that obtains secret information of information for specifying a plurality of worker models from the plurality of model learning devices; A secret aggregation processing unit that obtains secret information of information for specifying a new aggregated model obtained by aggregating a plurality of the worker models without obtaining the worker models by secret calculation using the secret information of the information for specifying the plurality of worker models; A providing unit that provides the information for specifying the new aggregated model or the secret information of the information for specifying the new aggregated model to the plurality of model learning devices, The local processing is a process in which the model learning device updates an aggregated model by machine learning using local learning data to obtain information for specifying the worker model, and provides secret information of the information for specifying the worker model to the secret federated learning device. The local processing time is the time corresponding to the local processing. The secret aggregation processing is a process in which the secret federated learning device obtains secret information of information for specifying the new aggregated model obtained by aggregating a plurality of the worker models without obtaining the worker models by secret calculation using the secret information of the information for specifying the worker models obtained by the secret federated learning device, and provides the information for specifying the new aggregated model or the secret information of the information for specifying the new aggregated model to the plurality of model learning devices. The aggregation processing time is the time corresponding to the secret aggregation processing. The acquisition unit When the local processing time is longer than the aggregation processing time, or when the local processing time is equal to or longer than the aggregation processing time, it performs an asynchronous acquisition process of obtaining secret information of information for specifying the worker models obtained by performing the local processing asynchronously with each other in the plurality of model learning devices that have acquired an instruction to execute the local processing by asynchronous control performed by the control device. A secret collaborative learning device that performs a synchronization acquisition process to obtain secret information of information identifying a worker model obtained by performing local processing synchronized with each other in a plurality of the model learning devices that have obtained an instruction to execute the local processing by synchronization control performed by the control device when the asynchronous acquisition process is not performed.
6. A control method by a control device that controls a collaborative learning system including a plurality of model learning devices and one or more secret collaborative learning devices, The model learning device updates an aggregated model by machine learning using local learning data to obtain information identifying a worker model, and executes local processing to provide secret information of the information identifying the worker model to the secret collaborative learning device. The secret collaborative learning device uses secret calculation using the secret information of the information identifying the obtained worker model to obtain secret information of information identifying a new aggregated model obtained by aggregating a plurality of the worker models without obtaining the worker models, and executes a secret aggregation process of providing the information identifying the new aggregated model or the secret information of the information identifying the new aggregated model to the plurality of model learning devices. The control method includes: A comparison step in which a comparison unit of the control device compares a local processing time corresponding to the local processing with an aggregation processing time corresponding to the secret aggregation processing; A control step in which a control unit of the control device performs asynchronous control to cause the local processing of the plurality of model learning devices to be executed asynchronously with each other when the local processing time is longer than the aggregation processing time or when the local processing time is equal to or longer than the aggregation processing time, and performs synchronization control to synchronize the local processing of the plurality of model learning devices with each other when the asynchronous control is not performed; A control method having the above.
7. A model learning method by a model learning device of a collaborative learning system including a plurality of model learning devices and one or more secret collaborative learning devices, An acquisition step in which an acquisition unit of the model learning device obtains information identifying an aggregated model or secret information of the information identifying the aggregated model from the secret collaborative learning device; A learning step in which a learning unit of the model learning device updates the aggregated model by machine learning using local learning data stored in a storage unit of the model learning device to obtain a worker model. A secrecy step in which a secrecy unit of the model learning device obtains secrecy information of information identifying the worker model; A provision step in which a provision unit of the model learning device provides the secrecy information of the information identifying the worker model to the secret federated learning device, and Local processing is a process in which the model learning device updates the aggregated model by machine learning using the local learning data to obtain information identifying the worker model, and provides the secrecy information of the information identifying the worker model to the secret federated learning device; The local processing time is the time corresponding to the local processing; Secret aggregation processing is a process in which the secret federated learning device uses the secrecy information of the information identifying the worker model obtained by the secret federated learning device to obtain, without obtaining the worker model, the secrecy information of the information identifying a new aggregated model obtained by aggregating a plurality of the worker models, and provides the information identifying the new aggregated model or the secrecy information of the information identifying the new aggregated model to the plurality of model learning devices; The aggregation processing time is the time corresponding to the secret aggregation processing; The model learning device When the local processing time is longer than the aggregation processing time, or when the local processing time is equal to or longer than the aggregation processing time, obtains an instruction to execute the local processing by asynchronous control performed by a control device, and executes the local processing asynchronously with other model learning devices included in the federated learning system; When not executing the local processing asynchronously, obtains an instruction to execute the local processing by synchronous control performed by the control device, and executes the local processing synchronized with other model learning devices included in the federated learning system. A model learning method.
8. A secret federated learning method by a secret federated learning device of a federated learning system including a plurality of model learning devices and one or more secret federated learning devices, An acquisition step in which an acquisition unit of the secret federated learning device obtains secrecy information of information identifying a plurality of worker models from a plurality of the model learning devices; A secret aggregation processing step in which a secret aggregation processing unit of the secret federated learning device uses the secrecy information of the information identifying the plurality of worker models to obtain, without obtaining the worker models, the secrecy information of the information identifying a new aggregated model obtained by aggregating the plurality of worker models; A providing step in which a providing unit of the secret federated learning device provides the plurality of model learning devices with information for specifying the new aggregated model or secret information of the information for specifying the new aggregated model. Local processing is a process in which the model learning device updates an aggregated model by machine learning using local learning data to obtain information for specifying the worker model, and provides secret information of the information for specifying the worker model to the secret federated learning device. The local processing time is the time corresponding to the local processing. Secret aggregation processing is a process in which, without obtaining the worker models, using secret information of the information for specifying the worker models acquired by the secret federated learning device through secret calculation, secret information of the information for specifying the new aggregated model obtained by aggregating the plurality of worker models is obtained, and the information for specifying the new aggregated model or the secret information of the information for specifying the new aggregated model is provided to the plurality of model learning devices. The aggregation processing time is the time corresponding to the secret aggregation processing. The acquisition step is When the local processing time is longer than the aggregation processing time, or when the local processing time is equal to or longer than the aggregation processing time, asynchronous acquisition processing is performed to obtain secret information of the information for specifying the worker models obtained by performing asynchronous local processing among the plurality of model learning devices that have obtained an instruction to execute the local processing by asynchronous control performed by the control device. When the asynchronous acquisition processing is not performed, synchronous acquisition processing is performed to obtain secret information of the information for specifying the worker models obtained by performing synchronous local processing among the plurality of model learning devices that have obtained an instruction to execute the local processing by synchronous control performed by the control device. A secret federated learning method.
9. A program for causing a computer to function as the control device according to claim 1, the model learning device according to claim 4, or the secret federated learning device according to claim 5.
Citation Information
Patent Citations
Apparatuses, computer program products, and computer-implemented methods for privacy-preserving federated learning
US20210256309A1
Privacy-preserving asynchronous federated learning for vertical partitioned data
US20220004933A1
Distributed and federated learning using multi-layer machine learning models
US20220083917A1