Federative learning system, client device, server device, federative learning method and program

JP7927775B2Active Publication Date: 2026-10-01KK TOSHIBA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024007594
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-01-22
Publication Date
2026-10-01
Estimated Expiration
2044-01-22

Smart Images

  • Figure 0007927775000001
    Figure 0007927775000001
  • Figure 0007927775000002
    Figure 0007927775000002
  • Figure 0007927775000003
    Figure 0007927775000003
Patent Text Reader

Abstract

To detect an anomaly in association learning while balancing information security and reduced calculation cost.SOLUTION: An integration processing unit generates global learning model parameters by performing integration processing of association learning using multiple local learning model parameters. A server difference calculation unit performs server difference calculation processing that shows the differences between each of the multiple local learning model parameters and the global learning model parameters. A learning unit performs learning processing of local learning models using global learning model parameters and learning data. A local difference calculation unit performs a local difference calculation process that indicates the differences between global learning model parameters and local learning model parameters. A detection unit performs detection processing to detect abnormalities in the server device, client device, or communication channel by comparing the server difference and the local difference.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Embodiments of the present invention relate to a federated learning system, a client device, a server device, a federated learning method, and a program. [Background technology]

[0002] In recent years, the amount of data collected by organizations such as companies in their internal operations and services has been increasing due to the spread of the internet and digital technologies, and this collected data is now being treated as big data. Attention is being drawn to the potential of this data to create new value. In this context, a technology called federative learning has emerged as a way to enable data sharing and utilization across different organizations. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Kavita Kumari,et.al.,“BayBFed: Bayesian Backdoor Defense for Federated Learning”,arXiv:2301.09508v1 [cs.LG] 23 Jan 2023 [Overview of the Initiative] [Problems that the invention aims to solve]

[0004] However, with conventional technologies, it has been difficult to detect anomalies in associative learning while simultaneously ensuring information security and reducing computational costs. [Means for solving the problem]

[0005] The federated learning system of the embodiment comprises a plurality of client devices and a plurality of server devices connected to the plurality of client devices. Each of the server devices comprises a server receiving unit, an integration processing unit, a server difference calculation unit, and a server transmitting unit. The server receiving unit receives local learning model parameters of a plurality of local learning models transmitted from the plurality of client devices via a plurality of communication channels. The integration processing unit generates global learning model parameters of a global learning model by performing an integrated processing of federated learning using the plurality of local learning model parameters. The server difference calculation unit performs a server difference calculation process that shows the difference between each of the plurality of local learning model parameters and the global learning model parameter. The server transmitting unit transmits the global learning model parameter to each of the plurality of client devices. Each of the plurality of client devices comprises a client receiving unit, a learning unit, a local difference calculation unit, and a client transmitting unit. The client receiving unit receives the global learning model parameters transmitted from the server device via the communication channel. The learning unit performs a learning process of the local learning model using the global learning model parameters and the learning data stored in the client device. The local difference calculation unit performs a local difference calculation process that shows the difference between the global learning model parameters received from the server device and the local learning model parameters. The client transmission unit transmits the local learning model parameters to the server device. At least one of the multiple client devices and the multiple server devices is equipped with a detection unit that performs a detection process to detect if there is an abnormality in the server device, the client device, or the communication path by comparing the server difference with the local difference. When the detection process is performed on the client device, the server transmission unit further transmits the server difference to the client device on which the detection process is performed. When the detection process is performed on the server device, the client transmission unit further transmits the local difference to the server device on which the detection process is performed. [Brief explanation of the drawing]

[0006] [Figure 1] A diagram showing an example of the device configuration of the associative learning system of the first embodiment. [Figure 2] A diagram showing an example of the functional configuration of an integrated server device and client terminal according to the first embodiment. [Figure 3] A flowchart showing an example of an anomaly detection method according to the first embodiment. [Figure 4] A flowchart showing an example of an anomaly detection method according to the second embodiment. [Figure 5] A flowchart illustrating an example of an anomaly detection method according to the third embodiment. [Figure 6] A flowchart showing an example of an anomaly detection method according to the fourth embodiment. [Figure 7] A flowchart showing an example of an anomaly detection method according to the fifth embodiment. [Figure 8] A flowchart showing an example of an anomaly detection method according to the sixth embodiment. [Figure 9] A diagram showing examples of the hardware configurations of the integrated server device and client terminal according to the first to sixth embodiments. [Modes for carrying out the invention]

[0007] The embodiments of the federated learning system, client device, server device, federated learning method, and program will be described in detail below with reference to the attached drawings.

[0008] Federated learning is an effective technique for improving the accuracy of machine learning based on data across multiple organizations. In federated learning, a global learning model is distributed to training data stored at each site. The global learning model is updated using training data stored independently at each site. The global learning model is then updated by receiving the difference between the pre-update model and the post-update model back into the global learning model.

[0009] In other words, in associative learning, only the learning result data is provided from the client terminal to the integrated server device. Therefore, since there is no need to provide learning data that requires privacy considerations to the integrated server device, there is an advantage in that the leakage of learning data during data transmission and from the integrated server can be prevented. Furthermore, since only model parameters are transmitted and received, it is also possible to reduce the amount of communication data.

[0010] However, since federated learning is designed with the possibility of attackers such as integrated server devices and client terminals, it poses privacy threats, such as the possibility of learning data being inferred from information obtained during the learning process. To address these issues, methods for detecting attacks on learning models, such as the one described in Non-Patent Document 1, have been reported, but these methods are computationally expensive.

[0011] The following embodiments describe a federated learning system, client device, server device, federated learning method, and program that perform anomaly detection in federated learning while balancing information security and suppression of computational costs.

[0012] (First Embodiment) First, an example of the device configuration of the associative learning system 1 of the first embodiment will be described.

[0013] [Example of device configuration] Figure 1 shows an example of the device configuration of the federated learning system 1 of the first embodiment. The federated learning system 1 of the first embodiment comprises a plurality of integrated server devices 10 (example of a server device) and a plurality of client terminals 20 (example of a client device).

[0014] The federated learning system 1 of the first embodiment comprises a plurality of integrated server devices 10. This relates to the use of secret sharing as the encryption method in the first embodiment.

[0015] The client terminals 20 shown in Figure 1 are each installed on the client's network, such as within a factory or store. Each client terminal 20 is a computer located within the network that can access the data held by each client.

[0016] In the first embodiment, it is assumed that the data held by each client is configured in a network that prevents access from other client terminals 20. Furthermore, it is assumed that each client terminal 20 exists for each set of data used to train a machine learning model.

[0017] To distinguish between multiple client terminals, they are denoted as C1, C2, ..., C5, as shown in Figure 1. The total number of client terminals is N. ca When N ca is an integer greater than or equal to 2. Each client terminal 20 includes a management program for executing local learning of federated learning.

[0018] Similarly, in order to distinguish between multiple integrated server devices 10, the multiple integrated server devices 10 are denoted as T1, T2, T3, etc., as shown in Figure 1.

[0019] [Example of functional configuration] Figure 2 shows an example of the functional configuration of the integrated server device 10 and client terminal 20 of the first embodiment. In Figure 2, a pair of integrated server devices 10 and client terminals 20 included in the federated learning system 1 of the first embodiment will be used as an example for explanation.

[0020] The integrated server device 10 of the first embodiment includes a storage unit 101, a synchronization processing unit 102, an integrated processing unit 103, a communication unit 104, a difference calculation unit 105, and a detection unit 106.

[0021] The memory unit 101 stores information about the global learning model (for example, global learning model parameters).

[0022] The synchronization processing unit 102 synchronizes the local learning model on each client terminal 20 with the global learning model before allowing each client terminal 20 to train its respective local learning model. This synchronization process, performed before training, eliminates the need to adjust matching parameters between the local and global learning models in federated learning, based on the global learning model information held by the integrated server device 10.

[0023] The integration processing unit 103 generates a global learning model by performing an integrated processing of associative learning based on the learning results (local learning model parameters) received from each client terminal 20.

[0024] The communication unit 104 performs communication processing such as receiving learning results from multiple client terminals and sending global learning model parameters to each client terminal. In addition, when the communication unit 104 sends global learning model parameters to each client terminal, it also encrypts the global learning model parameters using a secret sharing method.

[0025] The difference calculation unit 105 performs a server difference calculation process that shows the difference between each of the multiple local learning model parameters and the global learning model parameters.

[0026] The detection unit 106 performs a detection process to detect abnormalities in the integrated server device 10, the client terminals 20, or the communication path between the integrated server device 10 and the client terminals 20 by comparing the server difference with the local difference calculated by each client terminal 20 (described later).

[0027] The client terminal 20 of the first embodiment includes a storage unit 201, a learning unit 202, an encryption unit 203, a decryption unit 204, a communication unit 205, a difference calculation unit 206, and a detection unit 207.

[0028] The memory unit 201 stores the local learning model parameters of the local learning model, as well as the training data used for machine learning of the local learning model.

[0029] The learning unit 202 uses the training data to perform machine learning on a locally learned model.

[0030] The encryption unit 203 encrypts the local learning model parameters of the local learning model obtained through learning.

[0031] The decryption unit 204 decrypts the encrypted global learning model parameters received from the integrated server device 10.

[0032] The communication unit 205 performs communication processing such as transmitting encrypted local learning model parameters to the integrated server device 10 and receiving global learning model parameters transmitted from the integrated server device 10.

[0033] The difference calculation unit 206 performs a local difference calculation process that shows the difference between the global learning model parameters received from the integrated server device 10 and the local learning model parameters.

[0034] The detection unit 207 performs a detection process to detect abnormalities in the integrated server device 10, the client terminal 20, or the communication path between the integrated server device 10 and the client terminal 20 by comparing the server difference with the local difference.

[0035] Next, the operation of the first embodiment will be described.

[0036] In the first embodiment, secret sharing is used to conceal data, and the secret sharing uses a secret sharing method based on Shamir's threshold secret sharing scheme. Shamir's threshold secret sharing scheme converts one piece of secret information into n shared values and distributes them to n devices. In addition, if a predetermined number of pieces of the distributed information are collected, the original secret information can be restored. The secret sharing method of the first embodiment, for integers n and k of 2 or more that satisfy k≦n, generates shared information obtained by sharing the original data into n pieces, and has both erasure resistance that allows the original data to be restored even if n-k pieces of shared information are lost, and confidentiality that does not allow any part of the original data to be restored from less than k pieces of shared information.

[0037] The encryption unit 203 encrypts the local learning model parameters generated by the client terminal 20 using the above secret sharing method. Therefore, the number of integrated server devices 10 is equal to the number n of distribution destinations when encryption is performed by secret sharing. In addition, the number of data required for decrypting the global learning model obtained by the integrated server devices 10 is k, where k is an integer satisfying 1<k≦n.

[0038] For example, a neural network is used as the algorithm of the learning model targeted for anomaly detection in the first embodiment. A neural network is composed of an input layer, an output layer, and one or more hidden layers, each layer includes a plurality of nodes, and the nodes are interconnected. For network parameters including weights and biases that constitute the neural network, the learning model reflecting the training data in each client terminal 20 is updated between the client terminals 20 and the integrated server device 10.

[0039] In addition, the learning algorithm is not limited to a neural network with a specific structure, and a neural network of any structure including a convolutional neural network (CNN), a recurrent neural network (RNN), and the like may be used.

[0040] An example of the anomaly detection method by the federated learning system 1 according to the first embodiment will be described below.

[0041] [Examples of anomaly detection methods] Figure 3 is a flowchart showing an example of an anomaly detection method according to the first embodiment. The example in Figure 3 shows a series of flows in the processing of each round. In the series of flows, first, the client terminal 20 receives the global learning model and updates the local learning model. The integrated server device 10 receives the updated local learning model from the client terminal 20 and updates the global learning model. Then, the integrated server device 10 transmits the global learning model to each client terminal 20.

[0042] In the first embodiment, anomaly detection is performed on the client terminal 20 during each round to detect a malicious integrated server device 10, or to detect anomalies in the learning model during communication via the communication path between the client terminal 20 and the integrated server device 10.

[0043] Here, Nca represents the number of client terminals 20 used in the federated learning system 1 of the first embodiment, and Nsa represents the number of integrated server devices 10, with both Nca and Nsa being integers of 2 or greater. Also, t represents the number of rounds, and t=0 at the start of step S001.

[0044] First, the communication unit 205 of each client terminal 20 receives the initial values ​​of the global learning model from the integrated server device 10 or the like before starting the learning process (steps S001 and S003). The initial global learning model may be provided by the integrated server device 10, or it may be provided by another server device that only provides the initial values.

[0045] Furthermore, in the first embodiment, the flowchart is drawn assuming that encryption is performed to provide the initial global learning model to each client terminal 20 more securely. However, the encryption method for the initial global learning model is not limited to secret sharing and can be any method.

[0046] In step S002, if the number of rounds t is 1 or less, the process branches to step S003; if the number of rounds is 2 or more, the process branches to step S004. That is, at the time of the 0th round, which is when the initial values ​​of the global learning model are sent immediately after step S001, the local learning model has not yet been generated on the client terminal 20, so the detection steps S004 to S011 are not performed.

[0047] The period up to the reception of the initial values ​​for this global learning model is considered round 0. When the round count is incremented by 1 in step S012, it becomes round 1. Therefore, from round 1 onward, step S012 serves as the starting point for round number transitions.

[0048] (Round 1) The operation of the first round will now be explained. First, the decryption unit 204 decrypts the global learning model (step S013). Specifically, in the first round, in step S003, each client terminal 20 receives the initial global learning model. After that, the decryption unit 204 decrypts the received global learning model. As mentioned above, the encryption method for the initial global learning model is not limited, so the decryption method is based on the encryption method used.

[0049] Next, the learning unit 202 takes the learning dataset held in the memory unit 201 of each client terminal 20 as input and updates (learns) the local learning model using the decoded global learning model as initial values ​​(step S014). The local learning model parameters obtained in step S014 are then set m i t Let i be the client number, and satisfy 1 ≤ i ≤ Nca. The local learning model parameters of the local learning model obtained in this way are encrypted by secret sharing in the encryption unit 203 (step S015). Hereafter, the value encrypted by secret sharing will be called the share. The share of the local learning model parameters sent to the integrated server device x is [m i,x tis denoted as such. Here, the share of plaintext z is denoted as [z]. The number x identifying the integrated server device 10 satisfies 1≦x≦Nsa.

[0050] Next, the communication unit 205 transmits the share of the local learning model parameter [m i,x t generated in step S015 to the integrated server device x (step S016).

[0051] Next, the communication unit 104 of the integrated server device x receives the share of the local learning model parameter [m i,x t from each client terminal 20 (step S017). Then, in step S018, the number of rounds is checked, and step S019 is skipped in the first round. Step S019 is a step of calculating parameters for detection performed by the client terminal 20, and in the first round, the global learning model parameter [m x t-1 from one round before that is used herein does not exist, which is why step S019 is skipped.

[0052] Next, the integration processing unit 103 uses the shares of the local learning models received from each client terminal 20 to generate the global learning model parameter [M x t for the global learning model (step S020). The calculation for model integration is performed by secure computation, and all parameter calculations used in model integration and detection performed from the second round onward are also performed by secure computation. For example, the integration processing unit 103 calculates FedAvg from the local learning model parameters of each client terminal 20 and the number of data pieces used for learning, and generates the global learning model. Although FedAvg is taken as an example of the model integration calculation herein, other integration algorithms may also be used.

[0053] Next, the communication unit 104 transmits the global learning model parameter [M i,x tSend ] (step S021). However, if the transmission in step S021 is [M i,x t ] is only true when round number t=1, and from t=2 onwards, [M i,x t In addition to the detection parameter ΔD global ΔD is also sent. global This will be explained in the instructions for the second round and beyond.

[0054] Also, [M i,x t ] and [M x t ] represents the same global learning model, and the number i represents the number that identifies client terminal 20. [M i,x t ] indicates that these are global learning model parameters sent to client terminal i. The notation is differentiated to account for the possibility that the values ​​of the global learning model parameters may have changed since transmission when client terminal i receives them due to an attack such as model tampering or other reasons between client terminal i and the integrated server device x.

[0055] Next, we return to step S002, and since t=1, we proceed to step S003 at this branch, and the share of the global learning model sent from the integrated server device x in step S021 above [M i,x t The program receives [ ] (step S003). After step S003, the program proceeds to step S012, where the round number is incremented by 1, changing t=1 to t=2, and the operation for the first round ends.

[0056] (From the second round onwards) Next, we will explain the operation from the second round onward. From the second round onward, the process begins from step S013. The decoding unit 204 receives the share of the global learning model [M] from the integrated server device x. i,x tThe global learning model is decoded using [ ] (step S013). Subsequent steps S014-S017 are the same as in the first round.

[0057] At the branch in step S018, in the second round and beyond, the next step is step S019. The difference calculation unit 105 calculates the difference ΔD between the global learning model and the local learning model, which is a parameter used when anomaly detection is performed on the client terminal i. global Calculate (i,x) (Step S019).

[0058] ΔD global (i,x) is ΔD global (i,x)=[M x t-1 ]-[m i,x t ] is defined as. Item 1 [M x t-1 ] represents the share of the global learning model generated in the integrated server device x in round t-1. For example, in round 2, [M obtained in round 1] x 1 ] but [M x t-1 This corresponds to [m]. Also, item 2 is [m i,x t ] is the share of the local learning model generated by the client terminal i in round t in the integrated server device x and received in step S017.

[0059] This ΔD global (i,x) is performed at the client terminals 20 participating in the integration process. For example, if all client terminals 20 participate in the integration process, Nca ΔD operations are performed at each integration server device 10. global (i,x) is generated.

[0060] Next, as in the first round, a global learning model is generated in step S020, and then the communication unit 104 sets the share [M] of the global learning model generated in steps S019 and S020. i,x t] and ΔD global Send (i,x) to each client terminal i (step S021).

[0061] Next, we return to step S002, and since the branching condition t>1 is satisfied, the process proceeds to step S004. That is, the communication unit 205 of the client terminal i receives the parameter ([M i,x t ] and ΔD global (i,x)) is received (step S004). At this time, ΔD global (i,x) is received from all integrated server devices x that participated in the t-1 round. For example, if all integrated server devices 10 participated in the integration process, Nsa ΔD global (i,x)) is received.

[0062] Steps S005 to S011 are the detection process, which is performed by the difference calculation unit 206 or the detection unit 207. First, the detection unit 207 calculates the difference ΔD between the global learning model and the local learning model, which is a parameter used when performing detection on the client terminal i. local Calculate (i,x) (Step S005).

[0063] ΔD local (i,x) is ΔD local (i,x)=[M i,x t-1 ]-[m i,x t ] is defined as. Item 1 [M i,x t-1 ] represents the share of the global learning model generated in the integrated server device x in round t-1. That is, item 1 [M i,x t-1 ] is the [M received in step S003 if step S004 is in the t-1 round or if t-1=1. i,x t-1 ]. Also, the second item is [m i,x tis a share of the local learning model generated in step S014 of the t-th round and encrypted in step S015. ΔD local used for calculating (i,x) [m i,x t has never been leaked to the outside from the client terminal i that generated said [m i,x t .

[0064] Next, the difference calculation unit 206 calculates ΔD calculated in step S005 local (i,x) and ΔD received in step S004 global calculates a difference from (i,x) (step S006).

[0065] ΔD local (i,x) and ΔD global each term of (i,x) is calculated in the same process, respectively. Here, ΔD local one item of (i,x) [M i,x t-1 is a value received in step S004 of the (t-1)-th round, or in step S003 if t-1=1. On the other hand, ΔD global one item of (i,x) [M i,x t-1 is ΔD received in step S004 of the t-th round global it is a value used for calculating (i,x). Although the timing of transmission from the integrated server device x to the client terminal i is different, these two [M i,x t-1 values should inherently be the same value.

[0066] Also, ΔD local the second item of (i,x) [m i,x t is a value that, after being generated in step S015 of the t-th round, has been stored in the internal storage unit 201 or the like of the client terminal i that generated said [m i,x t . On the other hand, ΔD global the second item of (i,x) [m i,x tIn step S017 of the tth round, the integrated server device x receives from the client terminal i, and in step S019, ΔD global These are the values ​​used to calculate (i,x). Therefore, these two [m i,x t The values ​​of ] should ideally be the same.

[0067] Therefore, unless there is tampering by an attacker, anomalies in network communication or server processing, ΔD local (i,x)-ΔD global (i,x)=0 (difference ΔD) local and ΔD global (They match). Here, the tampering by the attacker may occur between the client terminal i and the integrated server device x, and also on the integrated server device x.

[0068] Next, the detection unit 207 detects ΔD local (i,x)-ΔD global The number of integrated server devices satisfying (i,x)=0 is counted (step S007), and this count value is denoted as Ns. At this time, the detection unit 207 of the client terminal i is ΔD local (i,x)-ΔD global The value of (i,x) is recorded as either 0 or non-zero for each identification number x of the integrated server device 10, and the results are stored as list L.

[0069] Next, the detection unit 207 determines that the count value Ns is the global learning model [M i,x t Determine whether the number of k or more is necessary for decoding ] (Step S008).

[0070] If Ns ≥ k (step S008, Yes), the detection unit 207 selects the share of the integrated server device to be used for decoding the global learning model (step S011). For example, the detection unit 207 selects ΔD from list L. local (i,x)-ΔD global Share of the global learning model of k integrated server devices x satisfying (i,x)=0 [M i,x tis selected at random.

[0071] On the other hand, when Ns<k (step S008, No), the detection unit 207 restores the share of the global learning model and the list L to the state of the (t-1)-th round (step S009). Specifically, the detection unit 207 of the client terminal i acquires the share [M of the global learning model in the t-th round i,x t with the adopted share [M of the global learning model that satisfied Ns≧k in the (t-1)-th round i,x t-1 . That is, [M i,x t =[M i,x t-1 holds. This replacement is performed on all servers Nsa. Further, the detection unit 207 also replaces the values of the aforementioned list L with corresponding values in accordance with the replacement of the share of the global learning model.

[0072] Next, the detection unit 207 again counts the number of integrated server devices Ns that satisfy ΔD local -ΔD global =0 (step S010). Here, it is assumed that Ns≧k was satisfied in step S008 in the (t-1)-th round, and the process of step S011 is performed.

[0073] The share [M of the global learning model selected in this manner i,x t is decrypted in the next round. Therefore, the operation up to the selection of the share [M of this global learning model i,x t corresponds to the operation of the t-th round.

[0074] The above is the processing flow of anomaly detection in the first embodiment.

[0075] (Second Embodiment) Next, a second embodiment will be described. In the description of the second embodiment, explanations similar to those of the first embodiment will be omitted, and the differences from the first embodiment will be described. In the second embodiment, the case in which the integrated server device 10 performs anomaly detection will be described.

[0076] [Examples of anomaly detection methods] Figure 4 is a flowchart showing an example of an anomaly detection method according to the second embodiment. The example in Figure 4 shows a series of flows in the processing of each round. In the series of flows, first, the client terminal 20 receives the global learning model and updates the local learning model. The integrated server device 10 receives the updated local learning model from the client terminal 20 and updates the global learning model. Then, the integrated server device 10 transmits the global learning model to each client terminal 20.

[0077] In the second embodiment, during each round, the integrated server device 10 performs anomaly detection, detecting malicious client terminals 20 or detecting anomalies in the learning model in communication via the communication path between the client terminals 20 and the integrated server device 10.

[0078] Let t be the number of rounds, and set t=0 at the start of step S101.

[0079] Step S101 is the same as step S001 in the first embodiment, so its description is omitted.

[0080] In step S102, if the number of rounds t is less than 1, the process branches to step S103; if the number of rounds is 1 or greater, the process branches to step S104. This branching occurs because, in the initial case, the encryption method for the received global learning model is not limited. From round 1 onward, the received global learning model is encrypted using secret sharing, and the shared data [M i,x t It is received as ].

[0081] The period up to the reception of the initial values ​​for this global learning model is considered round 0. When the round count is incremented by 1 in step S105, it becomes round 1. Therefore, from round 1 onward, step S105 is the starting point for the round number change.

[0082] (Round 1) The operation of the first round will now be explained. First, the decryption unit 204 decrypts the global learning model (step S106). Specifically, in the first round, in step S103, each client terminal 20 receives the initial global learning model. After that, the decryption unit 204 decrypts the received global learning model. As mentioned above, the encryption method for the initial global learning model is not limited, so the decryption method is based on the encryption method used.

[0083] Steps S107 and S108 are the same as steps S014 and S015 of the first embodiment, so their description is omitted.

[0084] If the number of rounds in step S109 is 1 or less, proceed to step S110. Step S110 is the same as step S016 in the first embodiment, so its explanation is omitted.

[0085] Next, in the processing of the integrated server device x, if the number of rounds is 1, the branch in step S113 proceeds to step S114, and then step S119 is executed. The processing in steps S114 and S119 is the same as steps S017 and S020 in the first embodiment, so the explanation is omitted. Next, the communication unit 104 sends the global learning model parameter [M i,x t Send ] (step S120).

[0086] Also, similar to the first embodiment, [M i,x t ] and [M x t] represents the same global learning model, and the number i represents the number that identifies client terminal 20. [M i,x t ] indicates that it is a global learning model parameter of the global learning model sent to client terminal i.

[0087] Next, we return to step S102, and since t=1, we proceed to step S104 at this branch, and the share of the global learning model sent from the integrated server device x in step S120 above [M i,x t The program receives [ ] (step S104). After step S104, the program proceeds to step S105, where the round number is incremented by 1, changing t=1 to t=2, and the operation for the first round ends.

[0088] (From the second round onwards) Next, we will explain the operation from the second round onward. From the second round onward, the process begins at step S105. Step S106 of the second round is the same as step S013 of the second round in the first embodiment, so we will omit its explanation. Furthermore, steps S107 and S108 that follow will be the same as in the first round.

[0089] At the branching point in step S109, in the second round and beyond, the next step will be step S111.

[0090] The difference calculation unit 105 calculates the difference ΔD between the global learning model and the local learning model, which are parameters used when anomaly detection is performed on the integrated server device x. local Calculate (i,x) (Step S111). ΔD local The definition of (i,x) is the same as in the first embodiment. ΔD local (i,x) 1st item [M i,x t-1 ] is received in step S104 of round t-1 [M i,x t-1 ]. Also, ΔD local The second term (i,x) [m i,xt ] is the share of the locally learned model generated in step S107 of the tth round and encrypted in step S108.

[0091] Next, the communication unit 205 receives the ΔD calculated in step S111. local (i,x) and the share of the local learning model parameters [m i,x t Send ] to the integrated server device x (step S112).

[0092] Then, on the integrated server device x side, in step S113, since the number of rounds is 2 or more, the process proceeds to step S115, and the communication unit 104 receives ΔD from each client terminal i. local (i,x) and the share of the updated local learning model [m i,x t ] is received (step S115).

[0093] Next, the difference calculation unit 105 calculates the difference ΔD between the global learning model and the local learning model, which is a parameter used when anomaly detection is performed on the integrated server device x. global Calculate (i,x) (Step S019). ΔD global The definition of (i,x) is the same as in the first embodiment. Note that ΔD global The second term (i,x) [m i,x t ] is the share of the locally learned model generated by client terminal i in round t and received in step S115.

[0094] This ΔD global (i,x) is performed at the client terminals 20 participating in the integration process. For example, if all client terminals 20 participate in the integration process, Nca ΔD operations are performed at each integration server device 10. global (i,x) is generated.

[0095] Next, the difference calculation unit 105 calculates the ΔD received in step S115. local (i,x) and ΔD calculated in step S116global (Calculate the difference between i and x (step S117). ΔD local (i,x) and ΔD global Each term in (i,x) is a value calculated using the same process. Here, ΔD local (i,x) 1st item [M i,x t-1 ] is received in step S104 of round t-1 and ΔD received in step S115 of round t. local This is the value used in the calculation of (i,x). On the other hand, ΔD global (i,x) 1st item [M i,x t-1 ] is a value that was calculated in step S119 of round t-1 and then stored in the memory unit 101 of the integrated server device x. ΔD global (i,x) 1st item [M i,x t-1 ] is sent from the integrated server device x to the client terminal i in round t-1, and in step S111 ΔD local The value used in the calculation of (i,x) should ideally be the same as the value used in the calculation.

[0096] Also, ΔD received in step S115 local The second term (i,x) [m i,x t ] is [M received in step S104 of round t-1, in step S108 of round t-1. i,x t-1 It is generated using ]. On the other hand, ΔD global (i,x) [m i,x t ] is the value received by the integrated server device x in step S115 of the tth round. These two [m i,x t The values ​​of ] should ideally be the same.

[0097] Therefore, unless there is tampering by an attacker, anomalies in network communication or server processing, ΔD local (i,x)-ΔD global(i,x)=0. Here, the tampering by the attacker could occur between the integrated server device x and the client terminal i, and also on the client terminal i.

[0098] Next, the detection unit 106 detects the share of the local learning model used for integration processing [m i,x t Select ] (step S118). Specifically, the detection unit 106 is ΔD local (i,x)-ΔD global The number of client terminals satisfying (i,x)=0 is counted, and this count value is denoted as Nc. At this time, the detection unit 106 of the integrated server device x determines ΔD local (i,x)-ΔD global The value of (i,x) is recorded as either 0 or non-zero for each number i that identifies the client terminal 20, and the results are stored as list L. This list L is created in each integrated server device x and shared among the integrated server devices 10. The detection unit 106 then performs ΔD in all integrated server devices 10. local (i,x)-ΔD global Share of the local learning model received from client terminal 20 which was determined to satisfy (i,x)=0 [m i,x t Select ].

[0099] Next, the integrated processing unit 103 determines the share of the local learning model selected in step S118 [m i,x t Using ], model integration is performed using functions such as FedAvg, and the share of the globally learned model [M x t Generate ] (step S119).

[0100] Step S120 is the same as Step S120 in the first round, so we will omit the explanation.

[0101] Then, in step S102, since the number of rounds is 2 or more, we proceed to step S104, and the communication unit 205 communicates the share of the global learning model [M i,xt ] is received (step S104). This completes the operation of the second round. Note that the processing in each round from the third round onward is the same as the operation of the second round.

[0102] (Modification 1 of the second embodiment) Next, a modification 1 of the second embodiment will be described. In the description of modification 1, explanations similar to those of the second embodiment will be omitted, and the differences from the second embodiment will be described. In modification 1, in step S117 of the detection process performed by the integrated server device x from the second round onward in the second embodiment, ΔD is performed t1 times consecutively (a predetermined number of times consecutively). local (i,x)-ΔD global Client terminal i for which (i,x)≠0 is not included in subsequent integration processing. The number of consecutive iterations t1 is set to any value greater than or equal to 1. Also, the number of client terminals required for integration is 2 or more, and if this condition is met, the next round of learning is performed.

[0103] (Modification 2 of the second embodiment) Next, a modified example 2 of the second embodiment will be described. In the description of modified example 2, explanations similar to those of the second embodiment will be omitted, and the differences from the second embodiment will be described. In modified example 2, in step S117 of the detection process performed by the integrated server device x from the second round onward in the second embodiment, ΔD is performed a total of t2 times (a predetermined total number of times). local (i,x)-ΔD global Client terminal i for which (i,x)≠0 is not included in subsequent integration processing. The total number of iterations t2 is set to any value greater than or equal to 1. Also, the number of client terminals required for integration is 2 or more, and if this condition is met, the next round of learning is performed.

[0104] As described in the first and second embodiments above, the detection process for detecting an abnormality in either the integrated server device 10, the client terminal 20, or the communication (communication path) between the integrated server device 10 and the client terminal 20 may be performed by the detection unit 106 of the integrated server device 10 or by the detection unit 207 of the client terminal 20.

[0105] In other words, in each integrated server device 10 (example of a server device), the communication unit 104 (example of a server receiving unit) receives local learning model parameters of multiple local learning models transmitted from multiple client terminals 20 (example of a client device) via multiple communication channels. The integration processing unit 103 generates global learning model parameters of a global learning model by performing an integrated processing of federated learning using the multiple local learning model parameters. The difference calculation unit 105 (example of a server difference calculation unit) calculates the server difference (ΔD as described above) which shows the difference between each of the multiple local learning model parameters and the global learning model parameter. global The calculation process for the ) is executed. Then, the communication unit 104 (example of a server transmission unit) sends the global learning model parameters to each of the multiple client terminals 20.

[0106] Furthermore, in each of the multiple client terminals 20, a communication unit 205 (example of a client receiving unit) receives global learning model parameters transmitted from the integrated server device 10 via the communication channel. A learning unit 202 uses the global learning model parameters and the learning data stored in the client terminal 20 to perform the learning process of the local learning model. A difference calculation unit 206 (example of a local difference calculation unit) calculates the local difference (ΔD as described above) which shows the difference between the global learning model parameters received from the integrated server device 10 and the local learning model parameters. local The calculation process for the parameters is executed. Then, the communication unit 205 (example of a client transmission unit) sends the local learning model parameters to the integrated server device 10.

[0107] Multiple client terminals 20 and at least one of the multiple integrated server devices 10 are equipped with a detection unit 106 (206) that performs a detection process to detect if there is an abnormality in either the integrated server device 10, the client terminals 20, or the communication (communication path) between the integrated server device 10 and the client terminals 20 by comparing the server difference and the local difference.

[0108] Furthermore, when detection processing is performed on client terminal 20, the communication unit 104 (example of server transmission unit) further transmits the server difference to client terminal 20 where detection processing is performed. When detection processing is performed on integrated server device 10, the communication unit 205 (example of client transmission unit) further transmits the local difference to integrated server device 10 where detection processing is performed.

[0109] Furthermore, in the first and second embodiments described above, the integration process and the server difference calculation process described above involve each of the multiple local learning model parameters and the global learning model parameters, and an encrypted share ([m i,x t ] and [M i,x t It is executed by secure computation under the state of ]). In addition, the local learning model shares the encrypted global learning model parameters [M i,x t It is updated based on the parameters decoded from ].

[0110] According to the first or second embodiment described above, anomaly detection in federated learning can be performed while ensuring information security and reducing computational costs. Specifically, the above share [m i,x t ] and [M i,x tThe model parameters are encrypted, and the difference between the locally learned model generated on the client terminal 20 and the globally learned model generated by the integration processing on the integrated server device 10 is calculated in an encrypted state on both the client terminal 20 and the integrated server device 10. Whether or not the values ​​match determines whether or not tampering has occurred. This ensures higher security and reduces the computational cost required for detection compared to conventional federated learning detection methods.

[0111] Furthermore, in the first or second embodiment described above, the value used for detection is the share of the local learning model encrypted by secret sharing [m i,x t ] and the share of global learning models [M i,x t-1 This is the difference between ]. This difference is calculated by ΔD calculated on the client terminal 20 and the integrated server device 10, respectively. local and ΔD global It is. Share [m i,x t ], share [M i,x t-1 ], ΔD local and ΔD global Even if the data is stolen when it is sent to the client terminal 20 or the integrated server device 10, in addition to the need to collect the number of corresponding values ​​required for decryption, ΔD is required for detection. local and ΔD global Since the calculation is performed without decryption, it is possible to achieve both information security and reduced computational costs.

[0112] (Third embodiment) Next, a third embodiment will be described. In the description of the third embodiment, explanations similar to those of the first embodiment will be omitted, and the differences from the first embodiment will be described. In the third embodiment, the detection process is performed at the client terminal, and an operation method that further limits the location of the abnormality compared to the operation of the detection unit 207 in the first embodiment will be described.

[0113] [Examples of anomaly detection methods] Figure 5 is a flowchart showing an example of an anomaly detection method according to the third embodiment.

[0114] (Round 1) The operation in the first round is the same as in the first embodiment, so the explanation will be omitted.

[0115] (From the second round onwards) The second round begins in step S216, after the round number is added in step S215. Steps S216 to S220 correspond to steps S013 to S016 of the first embodiment, respectively, and perform the same operations, so their explanation is omitted. Also, steps S220 to S224, which are operations in the integrated server device 10, perform the same operations as steps S017 to S021 of the first embodiment, so their explanation is omitted.

[0116] Next, the process returns to step S202, and if t is 2 or greater, the process proceeds to step S204, and steps S205 to S207 are executed. Steps S205 to S207 operate in the same way as steps S005 to S007 of the first embodiment, so their explanation is omitted.

[0117] Next, the detection unit 207 determines whether the condition Ns = Nsa is satisfied (step S208). Here, Ns is ΔD local (i,x)-ΔD global This indicates the number of integrated server devices 10 that satisfy (i,x)=0. Nsa indicates the number of integrated server devices 10 used in the federated learning system 1 of the third embodiment.

[0118] If Ns = Nsa (Step S208, Yes), the decoding unit 204 arbitrarily selects k shares from Nsa shares (shares of all integrated server devices 10) that are necessary for decoding the global learning model (Step S212).

[0119] On the other hand, if Ns ≠ Nsa (step S208, No), the decoding unit 204 determines the share of the global learning model from Nsa integrated server devices 10 [M i,xt From there, the global learning model is decoded using the k elements required for decoding (step S209).

[0120] Then, the detection unit 207 compares the global learning model decoded in step S209 with the global learning model from round t-1 (step S210). If the detection unit 207 finds that there are model parameter elements that have changed by d% or more as a result of the comparison, it determines the share of outliers [M] in the k combinations. i,x t Determine that it contains ]

[0121] Share of outliers [M i,x t If it is determined that ] is included, the detection unit 207, in particular, among the Nsa items, the conditional expression ΔD of step S207 local (i,x)-ΔD global Shares that did not satisfy (i,x=0) [M i,x t We will check the following. For example, the share of global learning models [M i,x t By changing the combination of selecting ] and performing decoding multiple times, the location of the anomaly can be more precisely identified. Specifically, in step S207, ΔD local (i,x)-ΔD global From the Ns shares that satisfy (i,x)=0, k-1 and ΔD local (i,x)-ΔD global By selecting one share that does not satisfy (i,x)=0 and implementing a composite global learning model, anomalies can be detected more efficiently. For example, anomalies in the integrated server device x and anomalies in the communication path can be detected.

[0122] Also, in step S207, ΔD local (i,x)-ΔD global Share of the global learning model of integrated server device x that does not satisfy (i,x)=0 [M i,x tIt is possible that using ] for decryption would not cause any problems. In this case, if an attack is carried out that transforms some of the parameters transmitted and received in the communication channel, [M i,x t ] and ΔD global Of these, ΔD global It is highly likely that only that part was tampered with.

[0123] Conversely, if the decoding result in step S210 is determined to contain an abnormal value, [M i,x t ] and ΔD global Among them, [M i,x t Even if only the [ ] is tampered with, there is a high probability that some of the parameters transmitted and received in the communication channel have been altered.

[0124] On the other hand, if both steps S207 and S210 show abnormalities, it is highly likely that an abnormality has already occurred at the time of model integration. In this case, [M i,x t ] and ΔD global The values ​​included in the generation of both parameters are those of the local learning model [M i,x t The data is likely to have been tampered with during the communication path from the client terminal i to the integrated server device x. Alternatively, the integrated server device x itself is likely to be outputting inaccurate values ​​due to the attack. In addition, a list of numbers x identifying Ns' normal integrated server devices 10 that were determined to be free of abnormalities in step S210 is recorded, and the result is stored as list L'.

[0125] Next, the detection unit 207 determines whether the condition Ns'≧k is satisfied (step S211). Here, Ns' represents the number of normal integrated server devices x that were determined to be free of abnormalities in step S210. k is the share of the global learning model of the integrated server devices x required for decoding [M i,x t This indicates the number of [ ].

[0126] When Ns'≧k (Yes in step S211), the decryption unit 204 arbitrarily selects k shares required for decryption from the Ns' shares in the list L' (shares of normal integrated server devices 10) that are used for decrypting the global learning model (step S212).

[0127] On the other hand, when Ns'<k (No in step S211), the detection unit 207 restores the shares of the global learning model, the list L and the list L' to the state of the (t-1)-th round (step S213). Specifically, the detection unit 207 of the client terminal i obtains the share [M of the global learning model in the t-th round i,x t with the share [M of the adopted global learning model that satisfies Ns'≧k in the (t-1)-th round i,x t-1 . That is, [M i,x t =[M i,x t-1 . Furthermore, the detection unit 207 also replaces the value of the above-mentioned list L with a corresponding value in accordance with the replacement of the shares of the global learning model.

[0128] Next, the detection unit 207 again counts the number Ns' of integrated server devices that satisfies ΔD local -ΔD global =0 (step S214). Here, it is assumed that Ns≧k is satisfied in step S211 in the (t-1)-th round, and the processing of step S212 is performed.

[0129] The share [M of the global learning model selected in this way i,x t will be decrypted in the next round. Therefore, the operation up to the selection of the share [M of the global learning model i,x t is the operation of the t-th round.

[0130] The above is the processing flow of abnormality detection in the third embodiment.

[0131] (Fourth Embodiment) Next, a fourth embodiment will be described. In the description of the fourth embodiment, descriptions similar to those of the first embodiment will be omitted, and only portions different from the first embodiment will be described. In the fourth embodiment, an example operation of reading the global learning model decoded in the (t-1)-th round when round t does not satisfy Ns≧k in the second and subsequent rounds will be described.

[0132] [Example of anomaly detection method] FIG. 6 is a flowchart showing an example of the anomaly detection method according to the fourth embodiment. In the example of FIG. 6, in step S008 performed after the second round in the first embodiment, the process for the case where Ns<k is changed to the process of step S311. In step S311 of the fourth embodiment, the detection unit 207 detects the share M of the global learning model of the (t-1)-th round i,x t , and reads the plaintext global learning model decoded in the (t-1)-th round instead of using said share.

[0133] (Fifth Embodiment) Next, a fifth embodiment will be described. In the description of the fifth embodiment, descriptions similar to those of the first embodiment will be omitted, and only portions different from the first embodiment will be described. In the fifth embodiment, an example operation of stopping the use of an integrated server device x in which an anomaly has been detected a plurality of times will be described.

[0134] [Example of anomaly detection method] FIG. 7 is a flowchart showing an example of the anomaly detection method according to the fifth embodiment. In the example of FIG. 7, according to the fifth embodiment, the client terminal i performs a detection step. For each pair of the client terminal i and the integrated server device x, the detection unit 207 calculates ΔD local -ΔD global counts the number of times ΔD-ΔD≠0 is satisfied. When the counted number is defined as r1(i,x), the method is configured such that an integrated server device x satisfying r1(i,x)≧a is not used for updating the global learning model.

[0135] (First round) The operation in the first round will be the same as in the first embodiment.

[0136] (From the second round onwards) The second round begins from step S416, after the round number is added in step S415. Steps S416 to S419 correspond to steps S013 to S016 of the first embodiment, respectively, and perform the same operations. Steps S420 to S424, which are operations in the integrated server device 10, are the same as steps S017 to S021 of the first embodiment, so their explanation is omitted.

[0137] Next, the process returns to step S402, and if t is 2 or greater, the process proceeds to step S404, and steps S405 to S407 are executed. Steps S405 to S407 operate in the same way as steps S005 to S007 of the first embodiment, so their explanation is omitted.

[0138] Next, the detection unit 207 of the client terminal i detects ΔD local (i,x)-ΔD global The identification number list Lng of the integrated server device x for which (i,x)≠0 t Generate (step S408). This identification number list Lng t This records which of the Nsa integrated server devices x is experiencing an anomaly (i.e., a verification error) in the t-th round, with which integrated server device x is the client terminal i experiencing a connection problem (i.e., the verification is failing).

[0139] Next, the detection unit 207 processes this identification number list Lng t Based on this, ΔD by round t local (i,x)-ΔD global The number of times (i,x)≠0 is counted for each integrated server device x (step S409), and this count is defined as r1(i,x).

[0140] Next, the detection unit 207 counts integrated server devices x for which the count value r1(i,x) in step S409 satisfies r1(i,x)≧a (step S410), and sets the count value as Nngs.

[0141] Next, the detection unit 207 determines whether a value obtained by subtracting the count value Nngs from the total number of integrated server devices Ns (Nsa-N ngs ) is equal to or greater than the number k required to decrypt the share of the global learning model (step S411). That is, shares of the global learning model parameters received from integrated server devices x that satisfy r1(i,x)≧a (where a is an example of a count threshold) are not used for decrypting the global learning model parameters.

[0142] Nsa-N ngs When Nsa-N ngs ≧k holds (Yes in step S411), the decryption unit 204 arbitrarily selects k shares required for decryption from the Nsa-N ngs shares (shares of integrated server devices x satisfying r1(i,x)<a) for use in decrypting the global learning model (step S414).

[0143] On the other hand, when Nsa-N ngs <k holds (No in step S411), the number of shares required for decryption is insufficient. Therefore, the detection unit 207 restores the share of the global learning model [M ngs and the list L to the state where Nsa-N i,x t ≧k was satisfied in the (t-1)-th round (step S412). Thereafter, after replacing the values with those of the (t-1)-th round again, the process returns to step S407 and proceeds in order.

[0144] Note that, in round t-1, it has been confirmed that Nsa-N ngs ≧k holds, and if it can be trusted that the value replaced in step S412 matches the value used in the (t-1)-th round, the process may proceed directly from step S412 to step S414. In this case, in step S412, the list Lng tAlso, replace it with the value from round t-1. That is, Lng t Lng t-1 Let's assume that.

[0145] Step S414 is the last step of the process in round t, and the process continues from there into round t+1.

[0146] (Modified version of the fifth embodiment) Next, a modified version of the fifth embodiment will be described. In describing the modified version, explanations similar to those of the fifth embodiment will be omitted, and the differences from the fifth embodiment will be described. In the modified version, the operation of controlling how many attacks on the integrated server device 10 will be allowed will be described.

[0147] In a modified example, in step S410 of the fifth embodiment, ΔD is performed a times or more. local (i,x)-ΔD global Share of integrated server device x such that (i,x)≠0 [M i,x t ] is not used for decoding the global learning model. Specifically, if a is reached in round t, the decoding unit 204 will, from round t+1 onwards, use the share [M] sent from its integrated server device x. i,x t Do not use ] for decoding in the global learning model.

[0148] At that time, the detection unit 207 determines the number of shares that can be used for decoding, i.e., Nsa-N ngs If the number of shares required for decoding the global learning model falls below k, the associative learning system 1 will be stopped.

[0149] (Sixth Embodiment) Next, the sixth embodiment will be described. In the description of the sixth embodiment, explanations similar to those of the second embodiment will be omitted, and the differences from the second embodiment will be explained. In the sixth embodiment, an example of the operation of stopping (terminating) the operation of the integrated server device x will be described.

[0150] [Examples of anomaly detection methods] Figure 8 is a flowchart showing an example of an anomaly detection method according to the sixth embodiment. In the example in Figure 8, the operation of the detection process performed by the integrated server device x from step S117 onwards (in the example in Figure 8, the operation from step S517 onwards) is different from the second round onwards in the second embodiment. Specifically, the detection unit 106 performs ΔD local -ΔD global Step S518 counts the number of client terminals Nc that satisfy =0, and steps S519 and S523 are added to stop operation if the count is less than a certain value b.

[0151] Here, b (an example of a threshold for the number of clients) can be arbitrarily set within the range Nca ≥ b ≥ 2 for all client terminals Nca.

[0152] Finally, examples of the hardware configurations of the integrated server device 10 and client terminal 20 according to the first to sixth embodiments will be described.

[0153] [Example hardware configuration] Figure 9 shows examples of the hardware configuration of the integrated server device 10 and client terminal 20 according to the first to sixth embodiments. The integrated server device 10 and client terminal 20 according to the first to sixth embodiments include a processor 301, main memory 302, auxiliary storage 303, display device 304, input device 305, and communication device 306. The processor 301, main memory 302, auxiliary storage 303, display device 304, input device 305, and communication device 306 are connected via a bus 310.

[0154] Furthermore, the integrated server device 10 and client terminal 20 may not be equipped with some of the above-described components. For example, if the integrated server device 10 can utilize the input and display functions of an external device, the integrated server device 10 may not be equipped with a display device 304 and an input device 305.

[0155] The processor 301 executes the program read from the auxiliary storage device 303 into the main memory device 302. The main memory device 302 is memory such as ROM and RAM. The auxiliary storage device 303 is such as an HDD (Hard Disk Drive) and a memory card.

[0156] The display device 304 is, for example, a liquid crystal display. The input device 305 is an interface for operating the integrated server device 10 and the client terminal 20. The display device 304 and the input device 305 may be implemented by a touch panel or the like that has both display and input functions. The communication device 306 is an interface for communicating with other devices.

[0157] For example, programs executed on the integrated server device 10 and client terminals 20 are provided as computer program products, recorded in installable or executable file format on computer-readable storage media such as memory cards, hard disks, CD-RWs, CD-ROMs, CD-Rs, DVD-RAMs, and DVD-Rs.

[0158] Alternatively, for example, the integrated server device 10 and the client terminal 20 may be configured to be stored on a computer connected to a network such as the Internet, and provided by allowing users to download them via the network.

[0159] Alternatively, for example, the integrated server device 10 and client terminals 20 may be configured to provide programs via a network such as the Internet without requiring downloads. Specifically, the server computer may execute the processing using a so-called ASP (Application Service Provider) type service, which does not transfer programs but instead implements processing functions only by issuing execution commands and obtaining results.

[0160] Alternatively, for example, the programs for the integrated server device 10 and the client terminal 20 may be pre-installed and provided in ROM or the like.

[0161] The programs executed on the integrated server device 10 and client terminals 20 are configured as modules that include functions that can also be implemented by programs, as described above. In actual hardware terms, each of these functions is loaded onto the main memory 302 by the processor 301 reading and executing a program from a storage medium. In other words, each of these function blocks is generated on the main memory 302.

[0162] Furthermore, some or all of the above-mentioned functions may be implemented using hardware such as an IC (Integrated Circuit) instead of software.

[0163] Alternatively, each function may be implemented using multiple processors 301. In this case, each processor 301 may implement one of the functions, or it may implement two or more of the functions.

[0164] While several embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of Symbols]

[0165] 1. Associative Learning System 10. Integrated Server Device 20 client terminals 101 Storage section 102 Synchronization Processing Unit 103 Integrated Processing Unit 104 Communications Department 105 Difference calculation unit 106 Detection unit 201 Storage unit 202 Learning unit 203 Encryption unit 204 Decryption unit 205 Communication unit 206 Difference calculation unit 207 Detection unit 301 Processor 302 Main storage device 303 Auxiliary storage device 304 Display device 305 Input device 306 Communication device 310 Bus

Claims

1. The system comprises multiple client devices and multiple server devices connected to the multiple client devices, Each of the aforementioned server devices is: A server receiving unit that receives local learning model parameters of multiple local learning models transmitted from multiple client devices via multiple communication channels, An integration processing unit that generates global learning model parameters for a global learning model by performing an integrated process of federated learning using multiple local learning model parameters, A server difference calculation unit that performs a server difference calculation process that shows the difference between each of the multiple local learning model parameters and the global learning model parameters, The system includes a server transmission unit that transmits the global learning model parameters to each of the multiple client devices, Each of the aforementioned client devices is A client receiving unit that receives the global learning model parameters transmitted from the server device via the communication channel, A learning unit that performs the learning process of the local learning model using the global learning model parameters and the learning data stored in the client device, A local difference calculation unit that performs a local difference calculation process that shows the difference between the global learning model parameters received from the server device and the local learning model parameters, The system includes a client transmission unit that transmits the local learning model parameters to the server device, At least one of the multiple client devices and the multiple server devices is: The system includes a detection unit that performs a detection process to detect if there is an abnormality in the server device, the client device, or the communication path by comparing the server difference with the local difference. When the detection process is performed on the client device, the server transmission unit further transmits the server difference to the client device on which the detection process is performed. When the detection process is performed on the server device, the client transmission unit further transmits the local difference to the server device on which the detection process is performed. Associative learning system.

2. The aforementioned integration process and the server difference calculation process are performed by secure computation, which calculates each of the multiple local learning model parameters and the global learning model parameters in an encrypted share state. The local learning model is updated based on parameters decrypted from the shared encrypted global learning model parameters. The associative learning system according to claim 1.

3. The detection unit detects the anomaly if the server difference and the local difference do not match. The associative learning system according to claim 1.

4. The detection unit is provided in the client device, The detection unit identifies the location of the anomaly by changing the combination of the shares of the multiple global learning model parameters used when the global learning model parameters are decoded, based on the shares of the multiple global learning model parameters. The associative learning system according to claim 2 or 3.

5. The aforementioned client device The detection unit, The system includes a decoding unit that decodes the global learning model parameters from the share of the global learning model parameters received from the server device, where the server difference and the local difference match. If, as a result of performing detection in round t, the detection unit finds that the number of server devices whose server difference and local difference match is less than the number of shares required to decode the global learning model parameters, it restores the shares of the global learning model parameters to the shares of the global learning model parameters received from the server devices in the detection in round t-1. The associative learning system according to claim 2 or 3.

6. The decoding unit shall not use the share of the global learning model parameters received from the server device for decoding the global learning model parameters if, as a result of detection performed in multiple rounds, the number of times the server difference and the local difference did not match exceeds a threshold number. The associative learning system according to claim 5.

7. The detection unit is provided in the client device, The detection unit stops the federated learning system if the number of server devices whose server difference matches the local difference is less than the number of shares required to decode the global learning model parameters. The associative learning system according to claim 2 or 3.

8. The detection unit is provided in the server device, The detection unit stops the federated learning system if the number of client devices whose server difference matches the local difference is less than the client number threshold. The associative learning system according to claim 2 or 3.

9. The detection unit is provided in the server device, The integration processing unit generates a share of the global learning model parameters by performing an integration process of federated learning using the shares of a plurality of local learning model parameters, and the detection unit determines that the shares of the local learning model parameters received from the client device in which the server difference and the local difference do not match for a predetermined number of consecutive times are not used in the integration process. The associative learning system according to claim 2 or 3.

10. The detection unit is provided in the server device, The integration processing unit generates a share of the global learning model parameters by performing a federated learning integration process using the shares of the multiple local learning model parameters. The shares of the local learning model parameters received from the client device are not used in the integration process if the number of times the server difference and the local difference do not match reaches a predetermined total number. The associative learning system according to claim 2 or 3.

11. A server receiving unit that receives local learning model parameters of multiple local learning models transmitted from multiple client devices via multiple communication channels, An integration processing unit that generates global learning model parameters for a global learning model by performing an integrated process of federated learning using multiple local learning model parameters, A server difference calculation unit that performs a server difference calculation process that shows the difference between each of the multiple local learning model parameters and the global learning model parameters, The system includes a server transmission unit that transmits the global learning model parameters to each of the multiple client devices, The server receiving unit receives local differences from each of the plurality of client devices, which indicate the difference between the global learning model parameters and the local learning model parameters. A detection unit performs a detection process to detect an abnormality in the device itself, the client device, or the communication path by comparing the server difference with the local difference. A server device equipped with the following features.

12. A client receiving unit that receives global learning model parameters transmitted from the server device via a communication channel, A learning unit that uses the global learning model parameters and the learning data to perform the learning process of a local learning model, A local difference calculation unit that performs a local difference calculation process that shows the difference between the global learning model parameters received from the server device and the local learning model parameters of the local learning model, The system includes a client transmission unit that transmits the local learning model parameters to the server device, The client receiving unit receives a server difference from the server device, which indicates the difference between the global learning model parameters and the local learning model parameters. A detection unit performs a detection process to detect if there is an abnormality in the device itself, the server device, or the communication path by comparing the server difference with the local difference. A client device equipped with the following features.

13. A federated learning method for a federated learning system comprising a plurality of client devices and a plurality of server devices connected to the plurality of client devices, Each of the server devices receives local learning model parameters of multiple local learning models transmitted from the multiple client devices via multiple communication channels. Each of the server devices generates global learning model parameters for a global learning model by performing an integrated process of federated learning using a plurality of local learning model parameters. Each of the server devices performs a server difference calculation process that shows the difference between each of the plurality of local learning model parameters and the global learning model parameters. Each of the server devices transmits the global learning model parameters to each of the plurality of client devices. Each of the plurality of client devices receives the global learning model parameters transmitted from the server device via the communication channel, Each of the plurality of client devices performs the learning process of the local learning model using the global learning model parameters and the learning data stored in the client device. Each of the plurality of client devices performs a local difference calculation process that shows the difference between the global learning model parameters received from the server device and the local learning model parameters. Each of the aforementioned client devices transmits the local learning model parameters to the server device. The process includes the step of having at least one of the multiple client devices and the multiple server devices perform a detection process to detect an abnormality in the server device, the client device, or the communication path by comparing the server difference with the local difference, When the detection process is performed on the client device, each of the server devices further transmits the server difference to the client device on which the detection process is performed. When the detection process is performed on the server device, each of the plurality of client devices further transmits the local difference to the server device on which the detection process is performed, Associative learning methods.

14. The server device A server receiving unit that receives local learning model parameters of multiple local learning models transmitted from multiple client devices via multiple communication channels, An integration processing unit that generates global learning model parameters for a global learning model by performing an integrated process of federated learning using multiple local learning model parameters, A server difference calculation unit that performs a server difference calculation process that shows the difference between each of the multiple local learning model parameters and the global learning model parameters, The server transmits the global learning model parameters to each of the multiple client devices, The server receiving unit receives local differences from each of the plurality of client devices, which indicate the difference between the global learning model parameters and the local learning model parameters. A detection unit performs a detection process to detect if there is an abnormality in the server device, the client device, or the communication path by comparing the server difference with the local difference. A program designed to function as such.

15. The client device A client receiving unit that receives global learning model parameters transmitted from the server device via a communication channel, A learning unit that uses the global learning model parameters and the learning data to perform the learning process of a local learning model, A local difference calculation unit that performs a local difference calculation process that shows the difference between the global learning model parameters received from the server device and the local learning model parameters of the local learning model, The local learning model parameters are configured to function as a client transmission unit that transmits them to the server device. The client receiving unit receives a server difference from the server device, which indicates the difference between the global learning model parameters and the local learning model parameters. A detection unit performs a detection process to detect if there is an abnormality in the server device, the client device, or the communication path by comparing the server difference with the local difference. A program designed to function as such.

Citation Information

Patent Citations

  • Learning system, device and method

    JP2023042922A

  • Secure computing system, financial institution server, information processing system, secure computing method, and recording medium

    WO2022269699A1