Client device, server device, information processing system, information processing method, and program
By calculating and integrating the importance of global models from multiple client devices, the federated learning method improves the accuracy of models, addressing the issue of low accuracy due to model aggregation with low affinity.
Patent Information
- Application Number
- PCT/JP2023/044326
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-19
AI Technical Summary
Existing federated learning techniques struggle to fully utilize the knowledge possessed by each client device, resulting in models with low accuracy due to the aggregation of models with low affinity.
A method where client devices calculate a degree of importance for each global model based on multiple global models provided by the server, and then integrate these global models to generate a new local model, which is used to improve the accuracy of the model.
This approach enhances the accuracy of the model by effectively integrating diverse knowledge from multiple client devices, leading to a more robust and accurate global model.
Smart Images

Figure JP2023044326_19062025_PF_FP_ABST
Abstract
Description
Client device, server device, information processing system, information processing method, and program
[0001] The present disclosure relates to a client device, a server device, an information processing system, an information processing method, and a program.
[0002] A technique called federated learning is known as a machine learning technique for distributed data sets. In federated learning, each client device learns a model distributed from a server device and transmits it to the server device. The server device then aggregates the received models and distributes them to each client device. This process is repeated. Each client device can obtain a model that reflects the knowledge of other client devices without disclosing its own knowledge to the other client devices.
[0003] It is also known to divide multiple client devices into clusters and perform federated learning for each cluster so that models of client devices with low affinity are not aggregated (see, for example, Patent Document 1).
[0004] International Publication No. 2021 / 059607 Pamphlet
[0005] The technology described in Patent Document 1, for example, has a problem in that it is not possible to fully utilize the knowledge possessed by each client device, and as a result, the accuracy of the created model is low.
[0006] The present disclosure has been made in view of the above-mentioned problems, and an exemplary purpose thereof is to provide a technique that can improve the accuracy of a model.
[0007] A client device according to an exemplary aspect of the present disclosure includes a learning means for learning a local model to generate a learned local model that is used by a server device to generate a global model related to the client device; a first calculation means for calculating a first importance level indicating the degree of importance assigned to each global model based on a plurality of global models provided by the server device; and an integration means for integrating the plurality of global models based on the first importance level assigned to each global model to generate a new local model to be learned by the learning means.
[0008] An information processing system according to an exemplary aspect of the present disclosure is an information processing system including the above-described plurality of client devices and the server device, wherein the server device includes an acquisition means for acquiring the trained local model from each of the plurality of client devices, an aggregation means for generating, for each of a plurality of clusters into which the plurality of client devices are divided, a global model that aggregates the trained local models acquired from each client device belonging to that cluster, and a provision means for providing, to each of the plurality of client devices, the plurality of global models consisting of the global models generated for each of the plurality of clusters.
[0009] A server device according to an exemplary aspect of the present disclosure includes an acquisition means for acquiring the trained local model from each of the plurality of client devices described above, an aggregation means for generating a global model for each of a plurality of clusters into which the plurality of client devices are divided, aggregating the trained local models acquired from each client device belonging to that cluster, and a provision means for providing each of the plurality of client devices with the plurality of global models consisting of the global models generated for each of the plurality of clusters.
[0010] A client device according to an exemplary aspect of the present disclosure includes an input data acquisition means for acquiring input data, and an inference means for performing inference on the input data using a local model obtained by integrating multiple global models provided from a server device based on a first importance level indicating the degree of importance given to each global model.
[0011] An information processing method according to an exemplary aspect of the present disclosure includes: a learning process in which at least one processor learns a local model to generate a learned local model that is used by a server device to generate a global model related to the local device; a first calculation process in which the at least one processor calculates a first importance level indicating the degree of importance assigned to each global model based on a plurality of global models provided by the server device; and an integration process in which the at least one processor integrates the plurality of global models based on the first importance level assigned to each global model to generate a new local model to be learned by the learning process.
[0012] An information processing method according to an exemplary aspect of the present disclosure includes an acquisition process in which at least one processor acquires the trained local model from each of a plurality of client devices that execute the information processing method described in claim 11; an aggregation process in which the at least one processor generates, for each of a plurality of clusters into which the plurality of client devices are divided, a global model that aggregates the trained local models acquired from each client device that belongs to the cluster; and a provision process in which the at least one processor provides, to each of the plurality of client devices, the plurality of global models consisting of the global models generated for each of the plurality of clusters.
[0013] An information processing method according to an exemplary aspect of the present disclosure includes an input data acquisition process in which at least one processor acquires input data, and an inference process in which the at least one processor performs inference on the input data using a local model obtained by integrating multiple global models provided by a server device based on a first importance level indicating the degree of importance assigned to each global model.
[0014] A program according to an exemplary aspect of the present disclosure is a program for causing a computer to function as the above-described client device, causing the computer to function as the learning means, the first calculation means, and the integration means.
[0015] A program according to an exemplary aspect of the present disclosure is a program for causing a computer to function as the server device described above, causing the computer to function as the acquisition means, the aggregation means, and the provision means.
[0016] A program according to an exemplary aspect of the present disclosure is a program for causing a computer to function as the above-described client device, causing the computer to function as the input data acquisition means and the inference means.
[0017] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that a technology capable of improving the accuracy of a model can be provided.
[0018] FIG. 1 is a block diagram showing a configuration of a client device according to the present disclosure. FIG. 1 is a flow diagram showing a flow of an information processing method according to the present disclosure. ... server device according to the present disclosure. FIG. 1 is a flow diagram showing a flow of an information processing method according to the present disclosure. FIG. 1 is a block diagram showing a configuration of an information processing system according to the present disclosure. FIG. 1 is a flow diagram showing a flow of an information processing method according to the present disclosure. FIG. 1 is a schematic diagram explaining an overview of an information processing system according to the present disclosure. FIG. 1 is a schematic diagram explaining an example configuration of a model targeted by an information processing system according to the present disclosure. FIG. 1 is a block diagram showing a configuration of an information processing system according to the present disclosure. FIG. 1 is a flow diagram showing a flow of an information processing method according to the present disclosure. FIG. 1 is a block diagram showing a configuration of an information processing system according to the present disclosure. FIG. 1 is a diagram showing a schematic diagram of an example of a second importance level according to the present disclosure. FIG. 1 is a flow diagram showing a flow of an information processing method according to the present disclosure. FIG. 1 is a block diagram showing an example hardware configuration of a computer functioning as each device according to the present disclosure.
[0019] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.
[0020] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.
[0021] (Overview of Present Exemplary Embodiment) An overview of a client device 1, a client device 2, a server device 3, and an information processing system 100 according to this exemplary embodiment will be described. As an example, the client device 1 trains a local model based on multiple global models provided from the server device, and provides the trained local model to the server device so that it can be used to generate at least one new global model. Here, the provision of multiple global models from the server device, the training of the local model, and the provision of the trained local model may or may not be repeated. The server device may be the server device 3 described below, or may not be the server device 3.
[0022] The client device 2 is, for example, a device that executes inference processing using a local model. The local model may or may not be, for example, a local model generated by the client device 1.
[0023] The server device 3 aggregates the local models acquired from the multiple client devices 1 for each cluster to generate a global model, and provides the multiple global models consisting of the generated global models for each cluster to each client device 1. Here, the acquisition of local models from each client device 1, the generation of a global model for each cluster, and the provision of the multiple global models may or may not be repeated. A cluster may be a group in which some data overlaps with each other. For ease of explanation, such groups will hereinafter be referred to as a "cluster."
[0024] The information processing system 100 includes a plurality of client devices 1 and a server device 3. At least two of the plurality of client devices 1 may possess knowledge in different domains. The information processing system 100 can also be considered as a system that performs so-called federated learning, but this term does not limit the present exemplary embodiment.
[0025] In this exemplary embodiment, any model configured to be updatable (trainable) can be applied as the global model and the local model. For example, specific examples of the model include models based on machine learning algorithms such as a convolutional neural network (CNN) or a recurrent neural network (RNN), and combinations thereof.
[0026] (Configuration of Client Device 1) The configuration of the client device 1 will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the client device 1. As shown in FIG. 1, the client device 1 includes a learning unit 11, a first calculation unit 12, and an integration unit 13. The learning unit 11, the first calculation unit 12, and the integration unit 13 are an example of a configuration that realizes the learning means, the first calculation means, and the integration means. For example, if the client device 1 includes at least one processor, the learning unit 11, the first calculation unit 12, and the integration unit 13 are realized by the at least one processor executing a program.
[0027] The learning unit 11 trains the local model to generate a trained local model that is used by the server device to generate a global model related to the device itself. For example, the learning unit 11 may train the local model using data in a domain corresponding to the device itself. The global model related to the device itself may be a global model corresponding to a cluster to which the device itself belongs.
[0028] The first calculation unit 12 calculates a first degree of importance indicating the degree of importance attached to each global model based on multiple global models provided by the server device. For example, the multiple global models may include a global model related to the server device. As an example, the multiple global models may correspond to multiple clusters including the cluster to which the server device belongs. For example, the first calculation unit 12 may obtain input or pre-set information for each cluster as the first degree of importance of the global model corresponding to the cluster. Furthermore, for example, the first calculation unit 12 may calculate the first degree of importance based on the evaluation results of each global model. However, the method for calculating the first degree of importance is not limited to this.
[0029] The integrating unit 13 generates a local model by integrating multiple global models based on the first importance level assigned to each global model. For example, the integrating unit 13 may generate a local model by integrating each global model with a weight based on the first importance level. However, the method for generating a local model is not limited to this.
[0030] (Effects of Client Device 1) As described above, the client device 1 is configured to include the learning unit 11, the first calculation unit 12, and the integration unit 13. Therefore, the local model obtained in the client device 1 reflects not only the knowledge of the client device itself but also a variety of knowledge, such as knowledge of other client devices related to each of the multiple global models. As an example, the local model obtained in the client device 1 reflects not only the knowledge of other client devices that belong to the same cluster as the client device itself but also a variety of knowledge, such as knowledge of other client devices that belong to clusters other than the cluster to which the client device itself belongs. Therefore, the client device 1 creates a model that reflects a variety of knowledge, thereby achieving the effect of realizing a highly accurate model.
[0031] (Flow of Information Processing Method S1) Next, the flow of the information processing method S1 executed by the client device 1 will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of the information processing method S1. As shown in Fig. 2, the information processing method S1 includes a learning process S11, a first calculation process S12, and an integration process S13.
[0032] In the learning process S11, at least one processor (e.g., learning unit 11) trains a local model to generate a learned local model that is used by the server device to generate a global model related to the device itself.
[0033] In the first calculation process S12, at least one processor (e.g., the first calculation unit 12) calculates a first importance level indicating the degree of importance attached to each global model based on multiple global models provided by the server device.
[0034] In the integration process S13, at least one processor (e.g., integration unit 13) integrates multiple global models based on the first importance level for each global model to generate a new local model to be trained by the learning process.
[0035] (Effects of Information Processing Method S1) As described above, the information processing method S1 includes the learning process S11, the first calculation process S12, and the integration process S13. Therefore, the information processing method S1 can achieve the same effects as the client device 1.
[0036] (Configuration of Client Device 2) Next, the configuration of the client device 2 according to this exemplary embodiment will be described with reference to FIG. 3. FIG. 3 is a block diagram showing the configuration of the client device 2. As shown in FIG. 3, the client device 2 includes an input data acquisition unit 21 and an inference unit 22. The input data acquisition unit 21 and the inference unit 22 are examples of configurations that realize an acquisition means and an inference means. For example, if the client device 2 includes at least one processor, the input data acquisition unit 21 and the inference unit 22 are realized by the at least one processor executing a program.
[0037] The input data acquisition unit 21 acquires input data. Here, the input data may be, for example, image data, text data, or other data, or a combination thereof. The input data acquired by the input data acquisition unit 21 is, for example, treated as a target for inference processing by the client device 2 in the inference phase.
[0038] The inference unit 22 performs inference on the input data using a local model. As the local model, a model obtained by integrating multiple global models provided from the server device based on a first importance level indicating the degree of importance attached to each global model is applied. For example, the multiple global models may be global models corresponding to multiple clusters including the cluster to which the inference unit 22 belongs.
[0039] The local model used by the inference unit 22 for inference may be the model itself obtained by integrating the multiple global models. As an example of such a local model, a "new local model" generated by the integration unit 13 of the client device 1 may be applied. Furthermore, the local model used by the inference unit 22 for inference may be a trained model obtained by performing training based on the integrated model. As an example of such a local model, a "trained local model" obtained by training the "new local model" generated by the integration unit 13 of the client device 1 by the learning unit 11 may be applied.
[0040] (Effects of Client Device 2) As described above, the client device 2 is configured to include the above-described input data acquisition unit 21 and inference unit 22. Therefore, the local model used for inference processing in the client device 2 reflects not only the knowledge of the client device itself but also a variety of knowledge, such as knowledge of other client devices related to each of the multiple global models. As an example, the local model reflects not only the knowledge of other client devices that belong to the same cluster as the client device itself, but also the knowledge of other client devices that belong to clusters other than the cluster to which the client device itself belongs. Therefore, the client device 2 can achieve the effect of performing highly accurate inference by utilizing a variety of knowledge, such as knowledge of other client devices.
[0041] (Flow of Information Processing Method S2) Next, the flow of the information processing method S2 executed by the client device 2 will be described with reference to Fig. 4. Fig. 4 is a flow diagram showing the flow of the information processing method S2. As shown in Fig. 4, the information processing method S2 includes an input data acquisition process S21 and an inference process S22.
[0042] In the input data acquisition process S21, at least one processor (for example, the input data acquisition unit 21) acquires input data.
[0043] In the inference process S22, at least one processor (e.g., the inference unit 22) performs inference on the input data using a local model, which is a model obtained by integrating multiple global models provided by the server device based on a first importance level indicating the degree of importance attached to each global model.
[0044] (Effects of Information Processing Method S2) As described above, the information processing method S2 includes the input data acquisition process S21 and the inference process S22. Therefore, the information processing method S2 can achieve the same effects as the client device 2.
[0045] (Configuration of Server Device 3) Next, the configuration of the server device 3 according to this exemplary embodiment will be described with reference to FIG. 5. FIG. 5 is a block diagram showing the configuration of the server device 3. As shown in FIG. 5, the server device 3 includes an acquisition unit 31, an aggregation unit 32, and a provision unit 33. The acquisition unit 31, the aggregation unit 32, and the provision unit 33 are an example of a configuration that realizes the acquisition means, the aggregation means, and the provision means. For example, if the server device 3 includes at least one processor, the acquisition unit 31, the aggregation unit 32, and the provision unit 33 are realized by the at least one processor executing a program.
[0046] The acquisition unit 31 acquires a trained local model from each of the multiple client devices 1 .
[0047] The aggregating unit 32 generates a global model for each of a plurality of clusters into which the plurality of client devices 1 are divided, by aggregating the local models after training acquired from each of the client devices 1 belonging to that cluster.
[0048] The client devices 1 may be divided into a plurality of clusters in advance, or may be divided into a plurality of clusters by the server device 3A. When the division is performed by the server device 3A, the server device 3A may acquire related information for classifying each client device 1, and divide the client devices 1 based on the related information to determine a plurality of clusters.
[0049] Furthermore, the number of client devices 1 belonging to each cluster may be one or more. For example, when a cluster has a plurality of client devices 1, the aggregating unit 32 may calculate an average of parameters of the local models after learning acquired from each of the client devices 1 belonging to the cluster, thereby generating a global model including the calculated parameters. However, the method of aggregating a plurality of local models is not limited to this.
[0050] The providing unit 33 provides a plurality of global models, each of which is made up of a global model generated for each of a plurality of clusters, to each of a plurality of client devices 1 .
[0051] (Effects of the Server Device 3) As described above, the server device 3 is configured to include the above-described acquisition unit 31, aggregation unit 32, and provision unit 33. Therefore, the multiple global models provided from the server device 3 to each client device 1 include not only a global model corresponding to the cluster to which the client device 1 belongs, but also global models corresponding to other clusters. The global models belonging to the other clusters reflect the knowledge of other client devices belonging to the other clusters. Then, in the client device 1, a local model is generated that integrates such multiple global models based on the first importance level. Therefore, the configuration of the server device 3 provides an effect that each client device 1 can utilize the knowledge of other client devices belonging to clusters different from the client device itself.
[0052] (Flow of information processing method S3) Next, the flow of the information processing method S3 executed by the server device 3 will be described with reference to Fig. 6. Fig. 6 is a flow diagram showing the flow of the information processing method S3. As shown in Fig. 6, the information processing method S3 includes an acquisition process S31, an aggregation process S32, and a provision process S33.
[0053] In the acquisition process S31, at least one processor (for example, the acquisition unit 31) acquires a trained local model from each of the multiple client devices 1.
[0054] In the aggregation process S32, at least one processor (e.g., aggregation unit 32) generates a global model for each of the multiple clusters into which the multiple client devices 1 are divided, by aggregating the learned local models obtained from each client device 1 belonging to the cluster.
[0055] In a providing process S33, at least one processor (e.g., a providing unit 33) provides a plurality of global models, each of which is made up of a global model generated for each of a plurality of clusters, to each of a plurality of client devices 1.
[0056] (Effects of Information Processing Method S3) As described above, the information processing method S3 employs a configuration including the above-described acquisition process S31, aggregation process S32, and provision process S33. Therefore, the information processing method S3 can obtain the same effects as the server device 3.
[0057] (Configuration of Information Processing System 100) Next, the configuration of the information processing system 100 according to this exemplary embodiment will be described with reference to FIG. 7. FIG. 7 is a block diagram showing the configuration of the information processing system 100. As shown in FIG. 7, the information processing system 100 includes a plurality of client devices 1-1, 1-2, 1-3, 1-4, ... and a server device 3. The client devices 1-1, 1-2, 1-3, 1-4, ... each have the same configuration as the client device 1, but may possess knowledge in different domains. Furthermore, it is assumed that the client devices 1-1 and 1-2 belong to a first cluster, and the client devices 1-3, 1-4, ... belong to a second cluster different from the first cluster.
[0058] The above-described client devices 1 can be applied as the client devices 1-1, 1-2, 1-3, 1-4, etc., and when there is no need to particularly distinguish between them, each will be simply referred to as a client device 1. While four client devices 1 are shown in FIG. 7 , the number of client devices 1 included in the information processing system 100 may be any number, such as two, three, or four or more. While FIG. 7 shows an example in which multiple client devices 1 are divided into two clusters, the number of clusters may be any number, such as three or more. Although FIG. 7 shows an example in which each cluster includes the same number of client devices 1, the number of client devices 1 included in each cluster may be the same as or different from at least one of the other clusters.
[0059] 7, the client device 1 and the server device 3 are communicatively connected via a network N. The specific configuration of the network N does not limit the present exemplary embodiment, but may be, for example, a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination of these networks. The configurations of the client device 1 and the server device 3 are as described above, and therefore will not be described in detail again.
[0060] (Flow of information processing method S100) Next, the flow of the information processing method S100 executed by the information processing system 100 will be described with reference to Fig. 8. Fig. 8 is a flow diagram showing the flow of the information processing method S100. As shown in Fig. 8, the information processing method S100 includes steps S11 to S13 and steps S31 to S33.
[0061] In step S11 (learning process), at least one processor (e.g., learning unit 11) of each client device 1 (e.g., 1-1, 1-2, 1-3, 1-4) trains a local model. As a result, a trained local model is generated in each client device 1, which is used by the server device 3 to generate a global model related to that client device 1. In addition, at least one processor of each client device 1 transmits the trained local model to the server device 3.
[0062] In step S31 (acquisition process), at least one processor (for example, the acquisition unit 31) of the server device 3 acquires the trained local model from each client device 1.
[0063] In step S32 (aggregation processing), at least one processor (e.g., aggregator 32) of server device 3 generates a global model for each of the multiple clusters into which the multiple client devices 1 are divided, aggregating the trained local models acquired from each client device 1 belonging to that cluster. For example, a first global model is generated by aggregating the trained local models from client devices 1-1 and 1-2 belonging to a first cluster. Also, a second global model is generated by aggregating the trained local models from client devices 1-3 and 1-4 belonging to a second cluster.
[0064] In step S33 (provision process), at least one processor (e.g., the provision unit 33) of the server device 3 provides a plurality of global models, each consisting of a global model generated for each of the plurality of clusters, to each of the plurality of client devices 1. As an example, the above-described first global model and second global model are provided to each of the client devices 1-1, 1-2, 1-3, and 1-4.
[0065] In step S12 (first calculation process), at least one processor (e.g., first calculation unit 12) of each client device 1 calculates a first degree of importance indicating the degree of importance attached to each global model based on the multiple global models provided by server device 3. The multiple provided models consist of global models corresponding to multiple clusters including the cluster to which the client device 1 belongs. As an example, client device 1-1 calculates a first first degree of importance for the first global model and a second first degree of importance for the second global model. Client devices 1-2, 1-3, and 1-4 similarly calculate first and second first degrees of importance.
[0066] In step S13 (integration process), at least one processor (e.g., integration unit 13) of each client device 1 integrates multiple global models based on the first importance levels assigned to each global model to generate a new local model to be trained through the learning process. As an example, client device 1-1 integrates the first and second global models based on the first and second first importance levels calculated by the client device itself to generate a new local model. New local models are similarly generated in client devices 1-2, 1-3, and 1-4.
[0067] Thereafter, each client device 1 may or may not repeat the process from step S11. When the process from step S11 is repeated, in the newly executed step S11, new learning is performed based on the new local model generated in step S13, and a new trained local model is generated. Furthermore, the server device 3 may or may not repeat the process from step S31. When the process from step S31 is repeated, in the newly executed step S31, a new trained local model is acquired from each client device 1. Note that the repeated process can be terminated at any time.
[0068] (Effects of Information Processing System 100) As described above, the information processing system 100 is configured to include a plurality of the above-described client devices 1 and a server device 3. Therefore, a plurality of global models corresponding to the plurality of clusters are provided to each client device 1 belonging to any of the plurality of clusters. The global model corresponding to each cluster reflects the knowledge of the client device 1 belonging to that cluster. Then, in each client device 1, a local model is generated by integrating the plurality of global models based on the first importance level. In other words, the local model obtained in the client device 1 reflects the knowledge of other client devices belonging to clusters other than the cluster to which the client device 1 belongs. For example, the local model obtained in the client device 1-1 reflects not only the knowledge of the client device 1-2 belonging to the same cluster, but also the knowledge of the client devices 1-3, 1-4, etc. belonging to other clusters. Therefore, the configuration of the information processing system 100 provides the effect that each client device 1 can utilize the knowledge of other client devices belonging to clusters different from the client device 1.
[0069] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.
[0070] (Overview of Information Processing System 100A) First, the problem solved by the information processing system 100A according to this exemplary embodiment will be described. In federated learning, the aggregation of models from client devices with low affinity can lead to a deterioration in accuracy. Therefore, a server device may divide multiple client devices into multiple clusters and perform federated learning for each cluster. In this case, a first problem and a second problem can be considered. As described in the "Problem to be Solved by the Invention," the server device cannot fully utilize the knowledge of other client devices that belong to different clusters from its own device. The second problem is that, because the server device divides multiple client devices into multiple clusters, acquiring related information from each client device may pose a risk of information leakage. For example, if the model to be used in federated learning uses image data as input, an example of related information is information indicating the type of image data (e.g., color image, monochrome image, etc.). If confidential information can be inferred from the type of image data stored, it is undesirable to disclose the type of image data to the server device. The information processing system 100A is an example of an information processing system that solves the second problem in addition to solving the first problem as in the first exemplary embodiment.
[0071] An overview of such an information processing system 100A will be described with reference to FIG. 9. FIG. 9 is a schematic diagram illustrating an overview of the information processing system 100A. The information processing system 100A is a system in which a server device 3A that holds a source model and M client devices 1A that hold data in each target domain perform federated learning while performing domain adaptation. Note that M is a natural number of 2 or greater, and when describing each client device 1A separately, they will be given a subnumber and referred to as client device 1A-k (k is a natural number between 1 and M).
[0072] The source model has been trained in advance using data in the source domain (hereinafter also referred to as source domain data) as training data. Domain adaptation is training the source model to obtain a target model adapted to the target domain. Each client device 1A obtains a target model adapted to its own target domain through associative learning. The source domain data is not disclosed to each client device 1A. Performing domain adaptation without reference to the source domain data is also referred to as source-free domain adaptation. Furthermore, each client device 1A holds data in the target domain (hereinafter also referred to as target domain data) as training data for associative learning. The target domain data is not disclosed to other client devices 1A. The target domain of each client device 1A is different from the target domain of at least one other client device 1A. Note that, in this exemplary embodiment, it is assumed that the source domain data is labeled and the target domain data is unlabeled, but this is not limited to this.
[0073] As shown in FIG. 9, each client device 1A-k obtains a source model provided by the server device 3A as a local model f^k to be initially trained. The "^k" in "f^k" is an alternative notation for the superscript k. Similarly, "^" is used below as an alternative notation for a superscript. Each client device 1A-k trains a local model f^k using target domain data stored in the client device, and then transmits the trained local model f^k to the server. The server device 3A performs a clustering process to divide the M client devices 1A-k into C clusters c based on the received M local models f^k. c is a natural number between 1 and C, and C is a natural number greater than or equal to 2. For example, clustering processing is performed such that client devices 1A-1, 1A-2, and 1A-3 belong to cluster 1, client devices 1A-4, 1A-5, 1A-6, and 1A-7 belong to cluster 2, and client devices 1A-(M-1) and 1A-M belong to cluster C. However, the number of client devices 1A-k belonging to each cluster only needs to be one or more, and is not limited to the example of FIG.
[0074] The server device 3A aggregates the local models f^k for each cluster c to generate a global model f^c, and provides the C global models f^c to each client device 1A-k. The "_c" in "f_c" is an alternative notation for the subscript c. Similarly, "_" is used below as an alternative notation for a subscript. Each client device 1A-k calculates a first importance u^k_c indicating the degree of importance attached to each global model f^c, and integrates the C global models f^c based on the first importance u^k_c to generate a new local model f^k. The period from when each client device 1A-k obtains the local model f^k to be learned to when it generates a new local model f^k is referred to as one round. The new local model f^k is then used as the new learning target, and the next round is repeated. Note that the C clusters c determined in the first round may also be applied in the second round and subsequent rounds. In other words, in the second round and thereafter, the process of dividing the M client devices 1A-k into C clusters c can be omitted. The local model f^k obtained after an appropriate number of rounds has been repeated is applied as a domain-adapted target model.
[0075] An example configuration of a model f targeted by the information processing system 100A will be described with reference to Fig. 10. Fig. 10 is a schematic diagram illustrating an example configuration of a model f targeted by the information processing system 100A. Here, the source model, target model, local model f^k, and global model f_c are each simply referred to as "model f." The local model f^k and global model f_c are models that are generated in the process of generating a target model that has been domain-adapted based on a source model through associative learning, and therefore these models f have the same configuration except for some or all of the parameters being different.
[0076] As shown in FIG. 10 , model f is a model that infers a class corresponding to a feature extracted from input data, and includes, as an example, a feature extraction model F and a classification model G. The feature extraction model F extracts feature data from the input data. The classification model G outputs a class corresponding to the feature data as an inference result. The feature extraction model F is a neural network and is configured by parameters assigned to each node. The classification model G is configured by weight vectors corresponding to each class. Furthermore, learning of the local model f^k executed in the client device 1A-k is performed by training the feature extraction model F without updating the classification model G. Therefore, among the local models f^k trained by each client device 1A-k, the parameters constituting the feature extraction model F may differ from one another, but the classification model G is common. Therefore, among multiple global models f_c aggregated in the server device 3A, the parameters constituting the feature extraction model F may differ from one another, but the classification model G is common.
[0077] (Configuration of Information Processing System 100A) The configuration of the information processing system 100A will be described with reference to FIG. 11. FIG. 11 is a block diagram showing the configuration of the information processing system 100A. As shown in FIG. 11, the information processing system 100A includes a server device 3A and M client devices 1A-k. Also, as shown in FIG. 11, the server device 3A and each of the client devices 1A-k are communicatively connected via a network N. A specific example of the network N is as described in the first exemplary embodiment, and therefore detailed description will not be repeated.
[0078] (Client Device 1A-k) The configuration of client device 1A-k will be described with reference to FIG. 11. As shown in FIG. 11, client device 1A-k includes a control unit 110, a storage unit 120, and a communication unit 130. The control unit 110 controls each unit of client device 1A. The storage unit 120 stores various data used by the control unit 110. For example, the storage unit 120 stores target domain data. The communication unit 130 communicates with server device 3A via network N under the control of the control unit 110.
[0079] The control unit 110 includes a learning unit 11, a first calculation unit 12, and an integration unit 13 included in the client device 1. The learning unit 11 is configured similarly to the first exemplary embodiment, but is also configured as follows: In the first round of federated learning, the learning unit 11 sets the source model provided by the server device 3A as the local model f^k to be learned. In the second round and subsequent rounds of federated learning, the learning unit 11 sets the new local model f^k generated by the integration unit 13 as the learning target. The learning unit 11 also trains the local model f^k using target domain data of its own device. In one round, the learning unit 11 performs several epochs of learning based on the local model f^k to generate the trained local model f^k. Note that this exemplary embodiment assumes that the target domain data is not labeled. Known techniques can be applied to machine learning techniques using unlabeled data.
[0080] The first calculation unit 12 is configured similarly to the first exemplary embodiment, but is also configured as follows. The first calculation unit 12 calculates a first degree of importance based on an evaluation result obtained by evaluating each global model f_c using evaluation data for the domain corresponding to the local device. As the evaluation data, part or all of the target domain data stored in the storage unit 120 is used. The evaluation data may be the same as or different from the target domain data used for learning by the learning unit 11. This allows the calculation of a first degree of importance according to the degree to which each global model f_c is adapted to the target domain corresponding to the local device. As a result, it is expected that the new local model f^k integrated by the integration unit 13 using such first degrees of importance will appropriately reflect the knowledge of client devices 1A-j (j≠k) belonging to other clusters.
[0081] As another example, the first calculation unit 12 may evaluate each global model f_c based on the degree to which feature amounts extracted by each global model f_c from the evaluation data converge to a class. For example, the degree of convergence of feature amounts to a class can be calculated based on the similarity between a feature amount vector extracted from the evaluation data by the feature amount extraction model F and a weight vector corresponding to a class inferred from the feature amount vector in the classification model G. As an example, a first importance level u^k_c, which is the degree to which the client device 1A-k values the global model f_c, is calculated using the following equations (1) to (3).
[0082] In equation (1), I^k_c indicates the evaluation result of the global model f_c by the client device 1A-k. x_i indicates N pieces of evaluation data used by the client device 1A-k. i is a natural number between 1 and N, and N is a natural number greater than or equal to 2. f_c(x_i) indicates the feature vector of the evaluation data x_i extracted by the feature extraction model F included in the global model f_c. Wl(i) indicates a weight vector corresponding to the class inferred from the feature vector f_c(x_i) in the classification model G.
[0083] Dist represents an arbitrary distance function. Note that in formulas (1) and (2), Dist is assumed to be a distance function (e.g., Euclidean distance) in which the function value decreases as the similarity increases. In other words, the smaller the value of the evaluation result I^k_c, the higher the evaluation. If a distance function (e.g., cosine similarity) in which the function value increases as the similarity increases is used as Dist, the larger the value of the evaluation result I^k_c, the higher the evaluation. In this case, formula (3) is slightly modified as described below.
[0084] l(i) indicates a class inferred from the feature vector f_c(x_i) by the classification model G. For example, l(i) is the class that has the smallest distance from the feature vector f_c(x_i) among the weight vectors Wl corresponding to the classes indicated by the classification model G, as shown in equation (2).
[0085] In equation (3), T represents a temperature parameter. As shown in equation (3), the first degree of importance u^k_c is calculated by normalizing the evaluation result I^k_c for each global model f_c calculated using equations (1) and (2) using the softmax function. The negative sign "-" on the right side of equation (3) is added so that the higher the evaluation result, the greater the first degree of importance u^k_c calculated for the evaluation result I^k_c, where the smaller the value, the higher the evaluation. As described above, when cosine similarity or the like is used as Dist in equations (1) and (2), equation (3) is modified as follows. In this case, the larger the value of the evaluation result I^k_c, the higher the evaluation, so the negative sign "-" on the right side of equation (3) is unnecessary.
[0086] Note that the technique for evaluating the performance of a model using unlabeled evaluation data is not limited to the method using equations (1), (2), and (3), and any known technique can be applied. An example of the known technique is SND (Soft Neighborhood Density), but is not limited to this.
[0087] The integrating unit 13 is configured similarly to the first exemplary embodiment, and is also configured as follows. As an example, a new local model f^k generated by the integrating unit 13 in the client device 1A-k may be calculated by the following equation (4). The new local model f^k indicated by equation (4) is applied as a target to be learned by the learning unit 11.
[0088] (Server Device 3A) The configuration of the server device 3A will be described with reference to FIG. 11. As shown in FIG. 11, the server device 3A includes a control unit 310, a storage unit 320, and a communication unit 330. The control unit 310 controls each unit of the server device 3A. The storage unit 320 stores various data used by the control unit 310. For example, the storage unit 320 stores a trained source model. The communication unit 330 communicates with each client device 1A-k via the network N under the control of the control unit 310.
[0089] The control unit 310 includes a clustering unit 34 in addition to the acquisition unit 31, aggregation unit 32, and provision unit 33 included in the server device 3. The acquisition unit 31, aggregation unit 32, and provision unit 33 are configured as described in the first exemplary embodiment, and are also configured as follows: The aggregation unit 32 applies multiple clusters determined by the clustering unit 34 (described later) as multiple clusters for which a global model f_c should be generated for each cluster. Furthermore, the provision unit 33 provides the trained source model to each client device 1A-k at the start of federated learning.
[0090] The clustering unit 34 constitutes an example of clustering means. If the server device 3A includes at least one processor, the clustering unit 34 is realized by the at least one processor executing a program. The clustering unit 34 divides the multiple client devices 1A-k into multiple clusters c based on some or all of the parameters constituting the local model f^k acquired from each of the multiple client devices 1A-k. This makes it possible to perform a process (clustering process) of dividing the multiple client devices 1A-k into multiple clusters c without acquiring information related to the target domain, which may pose a risk of information leakage.
[0091] For example, the clustering unit 34 may refer to the parameters of each local model f^k and perform the clustering process using a clustering method that does not require the number of clusters to be set in advance. This has the advantage of not requiring the number of clusters to be set in advance. A known technique can be applied as a clustering technique that does not require the number of clusters to be set in advance. Such a known technique includes, but is not limited to, FINCH (First Integer Neighbor Clustering Hierarchy).
[0092] An example of the clustering process performed by the clustering unit 34 will be described with reference to FIG. 10 . For example, as shown in FIG. 10 , the clustering unit 34 may perform the clustering process by referring only to parameter p1 included in the input-side layer of the feature extraction model F. The input-side layer includes at least the first layer counting from the input side, and may also include the second layer and beyond. This is because knowledge of the target domain tends to be more strongly reflected in layers closer to the input side of the feature extraction model F. This allows for highly accurate clustering process while saving computational costs compared to performing clustering process by referring to all of the parameters that make up the local model f^k.
[0093] (Flow of Information Processing Method S100A) Next, the flow of information processing method S100A executed by information processing system 100A will be described with reference to Fig. 12. Fig. 12 is a flow diagram showing the flow of information processing method S100A. As shown in Fig. 12, information processing method S100A includes steps S30A, S311A, S34A, and S14A in addition to steps substantially similar to those of information processing method S100. Furthermore, information processing method S100A includes step S12A instead of step S12 included in information processing method S100.
[0094] In step S30A, the providing unit 33 of the server device 3A provides the source model to each of the client devices 1A-k.
[0095] Step S11 (learning process) is described in the same manner as in the information processing method S100. More specifically, in step S11, the learning unit 11 of the client device 1A-k performs several epochs of learning using the target domain data of the client device 1A-k, with the local model f^k as the source model. The learning unit 11 also transmits the learned local model f^k to the server device 3A.
[0096] Step S31 (acquisition process) is explained in the same manner as in the information processing method S100. As a result, the server device 3A obtains a plurality of local models f^k.
[0097] Step S311A is an example of clustering processing. In step S311A, the clustering unit 34 of the server device 3A divides the multiple client devices 1A-k into multiple clusters based on some or all of the parameters constituting the local model f^k acquired from each client device 1A-k.
[0098] Steps S32 (aggregation processing) and S33 (provision processing) are described in the same manner as in the information processing method S100. As a result, multiple global models f_c generated for each cluster c are provided from the server device 3A to each client device 1A-k.
[0099] In step S12A, the first calculation unit 12 of the client device 1A-k calculates a first degree of importance f^k_c based on an evaluation result obtained by evaluating each global model f_c using the target domain data as evaluation data. More specifically, the first calculation unit 12 may evaluate each global model f_c based on the degree to which feature quantities extracted from the evaluation data by the global model f_c converge to a class. For example, the first calculation unit 12 may calculate the first degree of importance u^k_c using equations (1) to (3).
[0100] Step S13 (integration process) is described in the same manner as in the information processing method S100. More specifically, the integration unit 13 of the client device 1A-k may integrate multiple global models f_c using equation (4). This results in a new local model f^k.
[0101] In step S14A, the control unit 110 of the client device 1A-k determines whether a termination condition is met. For example, the termination condition may be, but is not limited to, receiving a termination notification from the server device 3A.
[0102] If the determination in step S14A is No, the control unit 110 repeats the process from step S11. In step S11 in the next round, learning is performed based on the new local model f^k generated in the most recently executed step S13.
[0103] If the determination in step S14A is Yes, the processing of the client device 1A-k ends. The new local model f^k generated in the most recently executed step S13 may be used as the domain-adapted target model. Alternatively, the trained local model f^k generated in the most recently executed step S11 may be used as the domain-adapted target model. In this case, step S14A, which determines whether the termination condition is satisfied, may be performed after step S11.
[0104] In step S34A, the control unit 310 of the server device 3A determines whether a termination condition is satisfied. For example, the termination condition may be, but is not limited to, that the number of times the round is repeated exceeds a threshold, with the execution of steps S11, S12A, S13, S31, S311A, S32, and S33 being considered as one round.
[0105] If the determination in step S34 is No, the control unit 310 repeats the process from step S31. If the determination in step S34 is Yes, the process of the control unit 310 ends. In this case, the control unit 310 may end the process after notifying each client device 1A-k that the federated learning has ended.
[0106] (Effects of Information Processing System 100A) As described above, the information processing system 100A employs a configuration in which the first calculation unit 12 of the client device 1A-k calculates a first degree of importance based on the evaluation results obtained by evaluating each global model f_c using evaluation data for the domain corresponding to the client device. Therefore, in addition to the effects achieved by the information processing system 100, the information processing system 100A can calculate a first degree of importance according to the degree to which each global model f_c is adapted to the domain of the client device. As a result, a new local model f^k obtained by integrating the global model f_c using such a first degree of importance can be expected to appropriately reflect the knowledge of other client devices 1A-j that belong to a different cluster from the client device.
[0107] Furthermore, in the information processing system 100A, each global model f_c is a model that infers a class corresponding to a feature extracted from the input data, and the first calculation unit 12 evaluates each global model f_c based on the degree to which the feature extracted from the evaluation data by the global model f_c converges to the class. Therefore, in addition to the effects achieved by the information processing system 100, the information processing system 100A has the effect of being able to accurately evaluate each global model f_c when a label is not associated with the evaluation data.
[0108] Furthermore, the information processing system 100A employs a configuration in which the server device 3A further includes a clustering unit 34 that divides the multiple client devices 1A-k into multiple clusters c based on some or all of the parameters constituting the local model f^k acquired from each of the multiple client devices 1A. Therefore, in addition to the effects achieved by the information processing system 100, the information processing system 100A has the effect of eliminating the need to acquire information related to the target domain, which may pose a risk of information leakage due to the division of the multiple client devices 1A-k into multiple clusters.
[0109] [Variation 1] An information processing system 100B, which is a variation of the information processing system 100A, will be described. In the information processing system 100B, a second degree of importance is used in addition to the first degree of importance when integrating multiple global models f_c. The second degree of importance indicates the degree to which each cluster emphasizes other clusters among multiple clusters.
[0110] (Configuration of Information Processing System 100B) The configuration of information processing system 100B will be described with reference to FIG. 13. FIG. 13 is a block diagram showing the configuration of information processing system 100B. As shown in FIG. 13, information processing system 100B includes a server device 3B and M client devices 1B-k. In the following description, differences between information processing system 100B and information processing system 100A will be mainly described. Other points will be described in the same manner as in the description of information processing system 100A, except that the suffix A in the reference symbols assigned to each element is replaced with B, and therefore description will not be repeated.
[0111] The control unit 110 of the server device 3B has the same configuration as the control unit 110 of the server device 3A, and further includes a second calculation unit 35. The second calculation unit 35 constitutes an example of second calculation means. When the server device 3B includes at least one processor, the second calculation unit 35 is realized by the at least one processor executing a program. The second calculation unit 35 calculates a second importance level α_c→c' indicating the degree to which each cluster c emphasizes another cluster c' among multiple clusters c.
[0112] Here, an example of the second level of importance α_c→c' will be described with reference to Fig. 14. Fig. 14 is a diagram schematically showing an example of the second level of importance α_c→c'. In the example of Fig. 14, the number C of clusters c is 3. Six second levels of importance α_1→2, α_2→1, α_1→3, α_3→1, α_2→3, α_3→2 are calculated between three clusters.
[0113] For example, the second calculation unit 35 may calculate the second degree of importance α_c→c′ between the multiple clusters c based on the first degree of importance u^k_c for each global model f_c acquired from each of the multiple client devices 1B-k. As an example, the second calculation unit 35 may calculate the second degree of importance α_c→c′ using the following equation (5):
[0114] In equation (5), u^k_c indicates the first importance level used by each client device 1B-k in the previous round. Furthermore, K(c) indicates the set of client devices 1B-k belonging to cluster c. Furthermore, n_k indicates the number of target domain data used by client device 1B-k in the learning process of the previous round. Note that the sum of the second importance levels α_c→c', which are the degree to which each of the other clusters c values one cluster c', is assumed to be 1.
[0115] The client device 1B-k has a configuration similar to that of the client device 1A-k, but differs in the details of the integration unit 13. The integration unit 13 integrates multiple global models f_c based on the second level of importance α_c→c' in addition to the first level of importance u^k_c. As an example, the integration unit 13 may integrate multiple global models f_c using the following equation (6). Based on the new local model f^k calculated by equation (6), the learning unit 11 performs a learning process in the next round.
[0116] According to equation (6), for example, in the example of FIG. 14 , the weight assigned to the global model f_1 of cluster 1 is calculated as the sum of (i) the product of the degree to which cluster 2 emphasizes cluster 1, α_2→1, and the degree to which the device itself emphasizes cluster 2, u^k_2, (ii) the product of the degree to which cluster 3 emphasizes cluster 1, α_3→1, and the degree to which the device itself emphasizes cluster 3, u^k_3, and (iii) the product of the degree to which cluster 1 emphasizes cluster 1, α_1→1, and the degree to which the device itself emphasizes cluster 1, u^k_1. The weight assigned to the global model f_2 of cluster 2 and the weight assigned to the global model f_3 of cluster 3 are similarly described. In the client device 1B-k, weights are assigned to the global models f_1, f_2, and f_3 in this way and integrated, thereby generating a new local model f^k.
[0117] Here, for example, when integrating multiple global models f_c, it is assumed that only the first degree of importance is used without using the second degree of importance. In this assumption, for example, if client device 1B-k belonging to cluster 1 does not emphasize other clusters 2 and 3, then integration is performed in client device 1B-k with emphasis on only the global model f_1 of cluster 1. However, in this assumption, for example, if client device 1B-j belonging to cluster 2 emphasizes cluster 1, it is desirable that client device 1B-k belonging to cluster 1 also performs integration with increased emphasis on the global model f_2 of cluster 2. However, in this assumption, as described above, integration is performed in client device 1B-k with emphasis on only the global model f_1 of cluster 1.
[0118] In this modified example, the integration process is performed using the second degree of importance in addition to the first degree of importance, so that the new local model f^k generated in each client device 1B-k is generated taking into consideration not only the degree to which the client device itself emphasizes each cluster (from the perspective of the client device itself), but also the degree to which each of the other client devices 1B-j belonging to the other clusters emphasizes each cluster (from the perspective of the other devices). As a result, each client device 1B-k can perform the integration process by incorporating the perspectives of the other devices without being biased toward the first degree of importance calculated only from the perspective of the client device itself, and as a result, it is possible to utilize the knowledge of each of the other client devices 1B-j to a more appropriate extent.
[0119] (Flow of Information Processing Method S100B) Next, the flow of information processing method S100B executed by information processing system 100B will be described with reference to FIG. 15. FIG. 15 is a flow diagram showing the flow of information processing method S100B. As shown in FIG. 15, information processing method S100B includes steps S111B and S321B in addition to steps substantially similar to those of information processing method S100A. Furthermore, information processing method S100B includes step S13B instead of S13 included in information processing method S100A. The following will mainly describe the differences between information processing method S100B and information processing method S100A. As for other points, they will be described in the same manner as in the description of information processing method S100A, except that the suffix A in the reference symbols assigned to each element is replaced with B, and therefore description will not be repeated.
[0120] After step S11 (learning process), client device 1B-k executes step S111B. In step S111B, control unit 110 of client device 1B-k transmits the first importance level u^k_c for each global model f_c calculated in step S12A of the previous round to server device 3B. Note that this step is not executed in the first round.
[0121] The server device 3B executes step S321B following step S32 (aggregation processing). In step S321B, the second calculation unit 35 calculates the second degree of importance α_c→c' between the multiple clusters c. For example, the second calculation unit 35 may calculate the second degree of importance α_c→c' between the multiple clusters c based on the first degree of importance u^k_c for each global model f_c acquired from each of the multiple client devices 1B-k. As an example, the second calculation unit 35 may calculate the second degree of importance α_c→c' using the above-mentioned formula (5).
[0122] In step S13B, the integration unit 13 of the client device 1B-k integrates the multiple global models f_c based on the second level of importance α_c→c′ in addition to the first level of importance u^k_c. As an example, the integration unit 13 may calculate the second level of importance α_c→c′ using the above-mentioned formula (6).
[0123] As described above, in the information processing system 100B, the server device 3B further includes a second calculation unit 35 that calculates a second importance level α_c→c' indicating the degree to which each cluster c among the plurality of clusters c values other clusters c', and the integration unit 13 of the client device 1B-k integrates the plurality of global models f_c based on the second importance level α_c→c' in addition to the first importance level u^k_c. Therefore, according to the information processing system 100B, in addition to the effects achieved by the information processing systems 100 and 100A, each client device 1B-k can perform the integration process by taking into account the degree to which each cluster values each other without being biased toward the first importance level based only on the perspective of the client device itself, and can utilize the knowledge of the other client devices 1B-j to a more appropriate degree.
[0124] Furthermore, the information processing system 100B employs a configuration in which the second calculation unit 35 of the server device 3B calculates the second degree of importance among the multiple clusters c based on the first degree of importance u^k_c for each global model f_c acquired from each of the multiple client devices 1B-k. Therefore, in addition to the effects achieved by the information processing systems 100 and 100A, the information processing system 100B achieves the effect that each client device 1B-k can perform integration processing by taking into account the perspectives of the other client devices 1B-j reflected in the second degrees of importance, without being biased toward the first degree of importance based only on the perspective of the client device 1B-k, and can utilize the knowledge of the other client devices 1B-j to a more appropriate degree.
[0125] [Variation 2] In the information processing system 100A or 100B, the integration unit 13 of the client device 1A-k or 1B-k can be further modified as follows: The integration unit 13 may generate a new local model by integrating an integrated model obtained by integrating multiple global models f_c with a global model f_c corresponding to a cluster to which the client device 1A-k or 1B-k belongs, based on the usefulness of the integrated model and the global model.
[0126] As an example, the integration unit 13 may calculate the new local model f^k' using the following equation (7).
[0127] In equation (7), f^k indicates an "integrated model obtained by integrating multiple global models f_c" and is calculated, for example, using equation (4) or equation (6). f_c(k) indicates a "global model corresponding to the cluster to which the client device (client device 1A-k or 1B-k) belongs." β^k is the usefulness of the integrated model f^k and the global model f_c(k). The usefulness may be calculated, for example, by evaluating each of the integrated model f^k and the global model f_c(k) using target domain data held by the client device. Such usefulness can be calculated, for example, by applying a temperature parameter and a softmax function to a value calculated using the SND technique described above, as in equation (3), but is not limited to this. In this modification, the learning unit 11 performs a learning process in the next round based on the new local model f^k' calculated using equation (7).
[0128] In this manner, in this modified example, the integrator 13 of the client device 1A-k or 1B-k generates a new local model f^k' by integrating an integrated model f^k obtained by integrating a plurality of global models f_c and a global model f_c(k) corresponding to the cluster to which the client device belongs, based on the usefulness β^k of the integrated model f^k and the global model f_c(k). Therefore, according to this modified example, in addition to the effects achieved by the information processing system 100, 100A, or 100B, each client device 1A-k or 1B-k can obtain an even more accurate model as a local model to be learned in the next round or as a target model.
[0129] [Application Examples] Application examples of the information processing systems 100, 100A, and 100B according to each exemplary embodiment will be described below. The information processing system 100A can be applied to a variety of industries, and several examples will be described below. However, these examples do not limit the exemplary embodiment, and the information processing systems 100, 100A, and 100B can, of course, be applied to other industries. Furthermore, the information processing systems 100, 100A, and 100B can also be applied across several industries.
[0130] (Example 1: Autonomous Driving Assistance) The information processing system 100, 100A, or 100B may be applied to the field of autonomous driving assistance, for example.
[0131] For example, the information processing system 100, 100A, or 100B can be applied when multiple companies in different locations jointly create a model that supports autonomous driving based on images from an in-vehicle camera. For example, multiple client devices 1, 1A-k, or 1B-k are each located in each company. Furthermore, the server device 3, 3A, or 3B is located in an organization that provides a model generation service using federated learning. In the following, for ease of explanation, "client devices 1, 1A-k, or 1B-k owned by each company" will simply be referred to as "each company."
[0132] Here, we assume that a source model trained using "in-vehicle camera images in urban areas" (source domain) has already been generated, and the data used for training (source domain data) has not been made public.
[0133] The multiple companies located in different locations include a company located in a snowy region, a company located in a coastal area, and a company located in a mountainous region. The company located in the snowy region has "in-vehicle camera images in the snowy region" (target domain 1). The company located in the coastal area has "in-vehicle camera images in the coastal area" (target domain 2). The company located in the mountainous area has "in-vehicle camera images in the mountainous area" (target domain 3). When training a source model to obtain a target model adapted to its own target domain, each company also wants to utilize knowledge in the target domains of other companies. However, since in-vehicle camera images may contain personal information, disclosing them to other companies may raise privacy issues and is not desirable.
[0134] In such an application example, the information processing system 100, 100A, or 100B performs a clustering process to divide multiple companies into multiple clusters based on the local model obtained through the federated learning process. As a result, it is expected that companies located in different locations (such as snowy regions, coastal areas, or mountainous areas) will be divided into different clusters. Note that for this clustering process, each company does not need to disclose related information related to the in-vehicle camera images.
[0135] Furthermore, the information processing system 100, 100A, or 100B generates a global model that aggregates local models for each of such multiple clusters, so that local models with low affinity, for example, a local model generated by a company located in a snowy region and a local model generated by a company located in a coastal area, are not aggregated together.
[0136] Then, each company performs an integration process to integrate the global model for snowy regions, the global model for coastal regions, and the global model for mountainous regions based at least on a first importance level, which indicates the degree to which the company attaches importance to each global model, to obtain a new local model. Therefore, for example, even if a company is located in a snowy region, it can obtain a new local model as a target model that appropriately reflects not only the knowledge of companies located in snowy regions like the company itself, but also the knowledge of companies located in coastal and mountainous regions. Furthermore, if the integration process uses a second importance level in addition to the first importance level, the perspectives of other companies are also taken into account without being biased toward the company's own perspective alone, so it is possible to obtain a target model that appropriately reflects the knowledge of companies in not only the company's cluster but also other clusters.
[0137] (Example 2: Finance-related) The information processing system 100, 100A, or 100B according to each exemplary embodiment may be applied to the finance-related field, for example.
[0138] For example, the information processing system 100, 100A, or 100B can be applied when multiple bank branches jointly create a model that predicts loan default risk based on borrower characteristics (loans, business conditions). For example, multiple client devices 1, 1A-k, or 1B-k are deployed at each branch. Furthermore, the server device 3, 3A, or 3B is deployed in an organization that provides a model generation service using federated learning. For ease of explanation, the "client device 1, 1A-k, or 1B-k owned by each branch" will be referred to simply as "each branch." It is also possible to apply each bank affiliate instead of each bank branch. In this case, the following explanation will be the same if "branch" is replaced with "affiliate."
[0139] Here, a source model trained using characteristic data (source domain) of borrowers at one of the stores has already been generated and stored in the server devices (3, 3A, 3B). Note that the data (source domain data) used to train the source model is not made public.
[0140] For example, each store has its own borrower feature data (target domain). When training a source model to acquire a target model adapted to its own target domain, each store wants to utilize knowledge of other stores' target domains. However, because borrower feature data includes personal information, it is not desirable for multiple stores to share this data with each other.
[0141] In such an application example, the information processing system 100, 100A, or 100B performs a clustering process to divide multiple stores into multiple clusters based on a local model obtained through federated learning. For example, stores with many small borrowers and stores with many large borrowers are expected to be divided into different clusters. Note that for this clustering process, each store does not need to disclose related information related to the borrower's characteristic data.
[0142] Furthermore, the information processing system 100, 100A, or 100B generates a global model that aggregates local models for each of such multiple clusters, so that local models with low affinity, for example, a local model generated by a store with many small borrowers and a local model generated by a store with many large borrowers, are not aggregated.
[0143] Then, each store performs an integration process to integrate the global models of each cluster based at least on the first importance level, which is the degree of importance the store places on each cluster, to obtain a new local model. Thus, a store with many small-lot borrowers can obtain a new local model as a target model that appropriately reflects not only the knowledge of other stores with many small-lot borrowers, but also the knowledge of other stores with many large-lot borrowers. Furthermore, if the integration process uses the second importance level in addition to the first importance level, the perspectives of other stores are also taken into account without being biased toward the store's own perspective alone, so that a target model can be obtained that appropriately reflects the knowledge of stores in not only its own cluster but also other clusters.
[0144] As another example of application to the financial field, the information processing system 100, 100A, or 100B can be applied when multiple insurance company branches jointly create a model to predict insurance premiums based on customer health data, such as medical history, age, blood pressure, and genetic information. In this case, the above-mentioned example of a "model to predict default risk" can be explained in the same way by replacing "borrower characteristic data" and "default risk" with "customer health data" and "insurance premiums," respectively. It is also possible to apply each of multiple insurance companies instead of each insurance company branch. In this case, the same explanation can be obtained by further replacing "branch" with "insurance company." It is expected that the clustering process will result in, for example, branches with a large number of young customers and branches with a large number of senior customers being divided into different clusters.
[0145] This allows federated learning to be performed by dividing multiple stores into clusters without the need for each store to disclose related information related to "customer health data." This prevents local models with low affinity from being aggregated, such as a local model generated by a store with many young customers and a local model generated by a store with many senior customers. Furthermore, for example, a store with many young customers can obtain a new local model as a target model that appropriately reflects not only the knowledge of other stores with many young customers but also the knowledge of other stores with many senior customers. Furthermore, when a second importance level is used in addition to the first importance level in the integration process, the perspectives of other stores are also taken into account without being biased toward the perspective of the store itself, thereby enabling a target model to be obtained that appropriately reflects the knowledge of stores in not only its own cluster but also other clusters.
[0146] The results of prediction of loan default risk or insurance premiums using the above model are examples of information that assists the user in making decisions according to this exemplary embodiment.
[0147] (Example 3: Medical and Healthcare Related) The information processing system 100A according to this exemplary embodiment may be applied to the medical field, for example.
[0148] For example, the information processing system 100, 100A, or 100B can be applied when each clinic jointly creates a model that suggests the cause or treatment of a disease based on patient symptom data recorded in a medical record or the like. In this case, in the above example of the "model for predicting default risk," the "store," "borrower characteristic data," and "default risk" can be read as "clinic," "patient symptom data," and "cause or treatment of disease," respectively, and the explanation can be similar. Note that the clustering process is expected to divide clinics with different medical departments, such as internal medicine, ophthalmology, and dentistry, into different clusters.
[0149] This allows federated learning to be performed by dividing multiple clinics into clusters without the need for each clinic to disclose related information related to "patient symptom data." This prevents local models with low affinity from being aggregated, such as a local model generated by an ophthalmology clinic and a local model generated by a dental clinic. Furthermore, for example, an ophthalmology clinic can obtain a new local model as a target model that appropriately reflects not only the knowledge of other ophthalmology clinics but also the knowledge of clinics in other medical departments. Furthermore, when a second importance level is used in addition to the first importance level in the integration process, the perspectives of other clinics are also taken into account without being biased toward the perspective of only the clinic itself, thereby obtaining a target model that appropriately reflects the knowledge of not only its own cluster but also the clinics in other clusters.
[0150] In another example of application in the medical and healthcare fields, the information processing systems 100, 100A, and 100B can be used when multiple pharmaceutical companies jointly create a model that predicts the activity of compounds based on data such as compound (ligand) structures and protein structures. In this case, in the above example of a "model for predicting default risk," the terms "store," "borrower characteristic data," and "default risk" can be replaced with "pharmaceutical company," "data such as compound (ligand) structures and protein structures," and "compound activity," respectively, to achieve the same explanation. Note that data such as compound structures and protein structures held by each pharmaceutical company is confidential information and should not be disclosed to each other. Furthermore, clustering processing is expected to divide, for example, pharmaceutical companies that handle different types of compounds into different clusters.
[0151] This allows federated learning to be performed by dividing multiple pharmaceutical companies into clusters, without the need for each pharmaceutical company to disclose related information related to "data such as compound (ligand) structures and protein structures." This prevents local models with low affinity, such as local models generated by pharmaceutical companies that handle different types of compounds, from being aggregated. Furthermore, for example, a pharmaceutical company that handles a certain type of compound can obtain a new local model as a target model that appropriately reflects not only the knowledge of other pharmaceutical companies that handle the same type of compound, but also the knowledge of other pharmaceutical companies that handle different types of compounds. Furthermore, when a second importance level is used in addition to the first importance level in the integration process, the perspectives of other pharmaceutical companies are also taken into account without being biased toward the company's own perspective alone, thereby obtaining a target model that appropriately reflects the knowledge of pharmaceutical companies in not only its own cluster but also other clusters.
[0152] The above-described model's suggestions for the cause or treatment of a disease and prediction results for the activity of a compound are examples of information that assists a user in making decisions according to this exemplary embodiment.
[0153] (Example 4: Machine-related) The information processing system 100A according to this exemplary embodiment may be applied to a machine-related field, for example.
[0154] For example, the information processing system 100, 100A, or 100B can be applied when each factory jointly creates a model for controlling the operation of a robot based on factory status data such as manufacturing status, transportation status, etc. In this case, in the above example of the "model for predicting default risk," the "store," "borrower characteristic data," and "default risk" can be read as "factory," "factory status data," and "control signal for controlling the operation of a robot," respectively, and the explanation can be similar. Note that, for example, it is expected that factories that produce different types of products will be divided into different clusters through the clustering process.
[0155] This allows federated learning to be performed by dividing multiple factories into clusters without the need for each factory to disclose related information related to "factory status data." This prevents local models with low affinity from being aggregated, such as a local model generated by a semiconductor parts factory and a local model generated by a clothing factory. Furthermore, for example, a semiconductor parts factory can obtain a new local model as a target model that appropriately reflects not only the knowledge of other factories that produce the same semiconductor parts, but also the knowledge of factories that produce other types of products. Furthermore, when a second importance level is used in addition to the first importance level in the integration process, the perspectives of other factories are also taken into account without being biased toward the perspective of only the factory itself, thereby enabling a target model to be obtained that appropriately reflects the knowledge of factories in not only its own cluster but also other clusters.
[0156] As another example of application to machinery-related fields, the information processing systems 100, 100A, or 100B can be applied when multiple companies manufacturing transportation equipment jointly create a model for controlling transportation equipment (such as automobiles, airplanes, or ships) based on data on scenery, equipment measurement status, and congestion. In this case, the above example of a "model for predicting default risk" can be explained in the same way by replacing "store," "borrower characteristic data," and "default risk" with "company," "scenery, equipment measurement status, and congestion data," and "control signal for controlling transportation equipment," respectively. Furthermore, it is expected that the clustering process will divide, for example, companies manufacturing different types of transportation equipment into different clusters.
[0157] This allows federated learning to be performed by dividing multiple companies into clusters without each company having to disclose related information related to "scenery, equipment measurement status, and congestion data." This prevents local models with low affinity from being aggregated, such as a local model generated by a company that manufactures automobiles and a local model generated by a company that manufactures ships. Furthermore, for example, a company that manufactures automobiles can obtain a new local model as a target model that appropriately reflects not only the knowledge of other companies that manufacture the same automobiles, but also the knowledge of other companies that manufacture other types of transportation equipment. Furthermore, when a second importance level is used in addition to the first importance level in the integration process, the perspectives of other companies are also taken into account without being biased toward the company's own perspective alone, thereby obtaining a target model that appropriately reflects the knowledge of not only the company's own cluster but also the knowledge of companies in other clusters.
[0158] As another example of application to machinery-related fields, the information processing system 100, 100A, or 100B can be applied when multiple logistics companies jointly create a model for calculating transportation routes based on data on the status of goods and transportation equipment. In this case, in the above example of a "model for predicting default risk," the "store," "borrower characteristic data," and "default risk" can be replaced with "logistics company," "data on the status of goods and transportation equipment," and "transportation route," respectively, to achieve the same explanation. Furthermore, clustering processing is expected to divide, for example, logistics companies using different types of transportation equipment into different clusters.
[0159] This allows federated learning to be performed by dividing multiple logistics companies into clusters without each logistics company having to disclose related information related to "data on the status of transported goods and transportation equipment." This prevents local models with low affinity from being aggregated, such as a local model generated by a logistics company that uses airplanes and a local model generated by a logistics company that uses trucks. Furthermore, for example, a logistics company that uses airplanes can obtain a new local model as a target model that appropriately reflects not only the knowledge of other logistics companies that also use airplanes, but also the knowledge of other logistics companies that use other types of transportation equipment. Furthermore, when a second importance level is used in addition to the first importance level in the integration process, the perspectives of other companies are also taken into account without being biased toward the company's own perspective alone, thereby obtaining a target model that appropriately reflects the knowledge of logistics companies in not only its own cluster but also other clusters.
[0160] Note that the control of the robot's operation, the control of the transport equipment, or the calculation result of the transport route using the above model is an example of information that supports the user's decision-making according to this exemplary embodiment.
[0161] (Example 5: Court-Related) The information processing system 100A according to this exemplary embodiment may be applied to the court-related field, for example.
[0162] For example, the information processing system 100, 100A, or 100B can be applied when multiple courts jointly create a model to predict sentences, etc., based on data that forms the basis of the trial, such as the circumstances of the crime, the circumstances of the evidence, laws, and court precedents. In this case, in the above example of a "model to predict default risk," the explanation can be similarly achieved by replacing "store," "borrower characteristic data," and "default risk" with "court," "data that forms the basis of the trial," and "sentences, etc." Furthermore, it is expected that different types of courts will be divided into different clusters through clustering processing.
[0163] As a result, federated learning is performed by dividing multiple courts into clusters, without each court having to disclose related information related to the "data underlying the trial." This prevents local models with low affinity, such as local models created by different types of courts, from being aggregated. Furthermore, for example, a certain type of court can obtain a new local model as a target model that appropriately reflects not only the knowledge of the same type of court but also the knowledge of other types of courts. Furthermore, when a second importance level is used in addition to the first importance level in the integration process, the perspectives of other courts are also taken into account without being biased toward the perspective of only the court itself, thereby obtaining a target model that appropriately reflects the knowledge of not only the courts in its own cluster but also the courts in other clusters.
[0164] The predicted results of sentencing and the like using the above model are an example of information that supports the user's decision-making according to this exemplary embodiment.
[0165] (Notes on each exemplary embodiment) The configurations described in each exemplary embodiment are not limited to the examples described above. The following configurations may be included to resolve some secondary issues that may arise when implementing federated learning in practice.
[0166] For example, when transmitting model parameters from each client device to a server device, it is preferable to have a configuration that ensures the confidentiality of the model parameters. For example, the information processing system described in each exemplary embodiment may have a configuration related to homomorphic encryption or the like that allows calculations to be performed while keeping the model parameters confidential, so that the confidentiality of the model parameters themselves can be ensured.
[0167] As an example, each client device (1, 1A-k, 1B-k) may be configured to include an encryption unit that encrypts model parameters of the local model after training using homomorphic encryption or the like, and the aggregation unit 32 of the server device (3, 3A, 3B) may aggregate the encrypted model parameters while keeping them confidential. Alternatively, the server device (3, 3A, 3B) may be configured to include an encryption unit that encrypts model parameters of the global model using homomorphic encryption or the like, and the model parameters are decrypted in each client device (1, 1A-k, 1B-k).
[0168] Furthermore, it is preferable to have a configuration that can minimize the data size of model parameters when they are transmitted from each client device to the server device. For example, each client device (1, 1A-k, 1B-k) and the server device (3, 3A, 3B) may be configured to compress the model parameters. Furthermore, at least one of each client device (1, 1A-k, 1B-k) and the server device (3, 3A, 3B) may be configured to include a model reconfiguration unit that reduces the size of the model by reconfiguring the model (generating a distilled model, a derived model, a pseudo model, or a higher-level model). For example, such a configuration is suitable in situations where the processing performance of the client device is limited.
[0169] Furthermore, as an example, when each client device (1, 1A-k, 1B-k) is implemented as a wearable device, the client device (1, 1A-k, 1B-k) may be configured to control the transmission of model parameters to the server device (3, 3A, 3B) according to the remaining battery charge. As an example, when the remaining battery charge is below a predetermined level, only model parameters that have changed from their previous values by a predetermined value (percentage) or more may be transmitted to the server device (3, 3A, 3B). Alternatively, the client device (1, 1A-k, 1B-k) may be configured to apply a sampling process to the acquired data (for example, extracting only 10% by random sampling) and perform a learning process by the learning unit 11 using only the sampled data. Alternatively, the client device (1, 1A-k, 1B-k) may be configured to store the acquired data in another device (for example, the server device 3, 3A, 3B).
[0170] Furthermore, since there is a possibility that noise may be present in each learning data and each parameter, each client device (1, 1A-k, 1B-k) and server device (3, 3A, 3B) may have a configuration for improving resistance to the noise. As an example, each client device (1, 1A-k, 1B-k) and server device (3, 3A, 3B) may have a configuration for performing error correction on the learning data or model parameters.
[0171] [Example of Implementation by Software] Some or all of the functions of the client devices 1, 1A-k, 1B-k, 2 and the server devices 3, 3A, 3B (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as an integrated circuit (IC chip), or by software.
[0172] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 16. Figure 16 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.
[0173] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.
[0174] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0175] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.
[0176] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0177] [Appendix 1] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0178] (Supplementary Note 1) A client device comprising: a learning means for generating a trained local model by training a local model, the trained local model being used by a server device to generate a global model related to the client device; a first calculation means for calculating a first importance level indicating a degree of importance given to each global model based on a plurality of global models provided by the server device; and an integration means for integrating the plurality of global models based on the first importance level for each global model to generate a new local model to be trained by the learning means. Note that the plurality of global models provided by the server device may include a global model related to the client device.
[0179] (Supplementary Note 2) The client device according to Supplementary Note 1, wherein the plurality of global models are global models corresponding to each of a plurality of clusters including a cluster to which the client device belongs. Note that the global model related to the client device may be a global model corresponding to the cluster to which the client device belongs.
[0180] (Supplementary Note 3) The client device according to Supplementary Note 1 or 2, wherein the first calculation means calculates the first importance level based on an evaluation result obtained by evaluating each global model using evaluation data of a domain corresponding to the client device.
[0181] (Supplementary Note 4) The client device according to Supplementary Note 3, wherein each of the global models is a model that infers a class corresponding to a feature extracted from input data, and the first calculation means evaluates each global model based on the degree to which the feature extracted by each global model from the evaluation data converges to the class.
[0182] (Supplementary Note 5) The client device described in any one of Supplementary Notes 1 to 4, wherein the integration means generates the new local model by integrating an integrated model obtained by integrating the multiple global models with a global model related to the client device based on the usefulness of the integrated model and the global model.
[0183] (Supplementary Note 6) An information processing system comprising a plurality of client devices according to any one of Supplementary Notes 1 to 5 and the server device, wherein the server device comprises: an acquisition means for acquiring the trained local model from each of the plurality of client devices; an aggregation means for generating, for each of a plurality of clusters into which the plurality of client devices are divided, a global model that aggregates the trained local models acquired from each client device belonging to the cluster; and a provision means for providing, to each of the plurality of client devices, the plurality of global models consisting of the global models generated for each of the plurality of clusters.
[0184] (Supplementary Note 7) The information processing system according to Supplementary Note 6, wherein the server device further comprises a clustering means for dividing the plurality of client devices into a plurality of clusters based on some or all of parameters constituting a local model acquired from each of the plurality of client devices.
[0185] (Appendix 8) The information processing system described in Appendix 6 or 7, wherein the server device further includes a second calculation means for calculating a second degree of importance indicating the degree to which each cluster emphasizes other clusters among the plurality of clusters, and the integration means of the client device integrates the plurality of global models based on the second degree of importance in addition to the first degree of importance.
[0186] (Supplementary Note 9) The information processing system according to Supplementary Note 8, wherein the second calculation means of the server device calculates the second degree of importance among the plurality of clusters based on the first degree of importance for each global model acquired from each of the plurality of client devices.
[0187] (Supplementary Note 10) A server device comprising: an acquisition means for acquiring the trained local model from each of a plurality of client devices described in any one of Supplementary Notes 1 to 5; an aggregation means for generating, for each of a plurality of clusters into which the plurality of client devices are divided, a global model that aggregates the trained local models acquired from each client device belonging to the cluster; and a provision means for providing, to each of the plurality of client devices, the plurality of global models consisting of the global models generated for each of the plurality of clusters.
[0188] (Supplementary Note 11) A client device comprising: an input data acquisition means for acquiring input data; and an inference means for performing inference on the input data using a local model obtained by integrating multiple global models provided from a server device based on a first importance level indicating the degree of importance given to each global model.
[0189] (Supplementary Note 12) An information processing method including: a learning process in which at least one processor learns a local model to generate a learned local model to be used by a server device to generate a global model related to the local device; a first calculation process in which the at least one processor calculates a first importance level indicating a degree of importance to each global model based on a plurality of global models provided by the server device; and an integration process in which the at least one processor integrates the plurality of global models based on the first importance level to each global model to generate a new local model to be learned by the learning process.
[0190] (Supplementary Note 13) An information processing method including: an acquisition process in which at least one processor acquires the trained local model from each of a plurality of client devices that executes the information processing method described in Supplementary Note 12; an aggregation process in which the at least one processor generates, for each of a plurality of clusters into which the plurality of client devices are divided, a global model that aggregates the trained local models acquired from each client device that belongs to the cluster; and a provision process in which the at least one processor provides, to each of the plurality of client devices, the plurality of global models consisting of the global models generated for each of the plurality of clusters.
[0191] (Supplementary Note 14) An information processing method including: an input data acquisition process in which at least one processor acquires input data; and an inference process in which the at least one processor performs inference on the input data using a local model obtained by integrating multiple global models provided from a server device based on a first importance level indicating the degree of importance given to each global model.
[0192] (Supplementary Note 15) A program for causing a computer to function as the client device according to any one of Supplementary Notes 1 to 5, the program causing a computer to function as the learning means, the first calculation means, and the integration means.
[0193] (Supplementary Note 16) A program for causing a computer to function as the server device according to Supplementary Note 10, the program causing a computer to function as the acquiring means, the aggregating means, and the providing means.
[0194] (Supplementary Note 17) A program for causing a computer to function as the client device according to Supplementary Note 11, the program causing the computer to function as the input data acquisition means and the inference means.
[0195] 1, 1-1, 1-2, 1-3, 1-4, 1A-k, 3B-k Client device 3, 3A, 3B Server device 100, 100A, 100B Information processing system 11 Learning unit 12 First calculation unit 13 Integration unit 21 Input data acquisition unit 22 Inference unit 31 Acquisition unit 32 Aggregation unit 33 Provision unit 34 Clustering unit 35 Second calculation unit 110, 310 Control unit 120, 320 Storage unit 130, 330 Communication unit C1 Processor C2 Memory x_i Evaluation data
Claims
1. A learning means for generating a learned local model to be used by a server device to generate a global model related to the own device by learning a local model; a first calculating means for calculating a first degree of importance indicating the degree of importance of each global model based on a plurality of global models provided from the server device; and an integrating means for generating a new local model to be learned by the learning means by integrating the plurality of global models based on the first degree of importance for each global model. A client device comprising the above components.
2. The client device according to claim 1, wherein the plurality of global models are global models corresponding to each of a plurality of clusters including the cluster to which the own device belongs.
3. The client device according to claim 1 or 2, wherein the first calculating means calculates the first degree of importance based on an evaluation result obtained by evaluating each global model using evaluation data of a domain corresponding to the own device.
4. Each of the global models is a model for inferring a class corresponding to a feature amount extracted from input data, and the first calculating means evaluates the global model based on the degree to which the feature amount extracted from the evaluation data by each global model converges to the class. The client device according to claim 3.
5. The integrating means generates the new local model by integrating an integrated model obtained by integrating the plurality of global models and the global model related to the own device based on the degree of usefulness regarding the integrated model and the global model. The client device according to any one of claims 1 to 4.
6. An information processing system comprising the plurality of client devices according to any one of claims 1 to 5 and the server device, wherein the server device includes: an acquisition means for acquiring the learned local model from each of the plurality of client devices; an aggregation means for generating a global model by aggregating the learned local models acquired from the client devices belonging to each of the plurality of clusters into which the plurality of client devices are divided; and a provision means for providing the plurality of global models each composed of the global models generated for each of the plurality of clusters to each of the plurality of client devices.
7. The server device further includes a clustering means for dividing the plurality of client devices into a plurality of clusters based on a part or all of the parameters constituting the local model acquired from each of the plurality of client devices. The information processing system according to claim 6.
8. The server device further includes a second calculation means for calculating a second degree of importance indicating the degree to which each cluster values other clusters among the plurality of clusters. The integration means of the client device integrates the plurality of global models based on the second degree of importance in addition to based on the first degree of importance. The information processing system according to claim 6 or 7.
9. The second calculation means of the server device calculates the second degree of importance among the plurality of clusters based on the first degree of importance for each global model acquired from each of the plurality of client devices. The information processing system according to claim 8.
10. An acquisition means for acquiring the learned local model from each of the plurality of client devices according to any one of claims 1 to 5; an aggregation means for generating a global model obtained by aggregating the learned local models acquired from each client device belonging to a corresponding one of the plurality of clusters into which the plurality of client devices are divided; and a provision means for providing the plurality of global models each composed of the global models generated for each of the plurality of clusters to each of the plurality of client devices. A server device comprising the above.
11. An input data acquisition means for acquiring input data; an inference means for performing inference on the input data using a local model obtained by integrating a plurality of global models provided from a server device based on a first degree of importance indicating the degree of importance of each global model. A client device comprising the above.
12. An information processing method including: a learning process in which at least one processor generates a learned local model used by a server device to generate a global model related to the own device by learning the local model; a first calculation process in which the at least one processor calculates a first degree of importance indicating the degree of importance of each global model based on the plurality of global models provided from the server device; and an integration process in which the at least one processor generates a new local model to be learned by the learning process by integrating the plurality of global models based on the first degree of importance for each global model.
13. An information processing method including: an acquisition process in which at least one processor acquires the learned local model from each of a plurality of client devices that execute the information processing method according to claim 12; an aggregation process in which the at least one processor generates a global model by aggregating the learned local models acquired from each client device belonging to each of a plurality of clusters into which the plurality of client devices are divided; and a provision process in which the at least one processor provides each of the plurality of client devices with the plurality of global models each composed of the global models generated for each of the plurality of clusters.
14. An information processing method including: an input data acquisition process in which at least one processor acquires input data; and an inference process in which the at least one processor performs inference on the input data using a local model obtained by integrating a plurality of global models provided from a server device based on a first degree of importance indicating the degree of emphasis on each global model.
15. A program for causing a computer to function as the client device according to any one of claims 1 to 5, the program for causing a computer to function as the learning means, the first calculation means, and the integration means.
16. A program for causing a computer to function as the server device according to claim 10, the program for causing a computer to function as the acquisition means, the aggregation means, and the provision means.
17. A program for causing a computer to function as the client device according to claim 11, the program for causing a computer to function as the input data acquisition means and the inference means.