Learning device, learning system, learning method, and recording medium

The learning device and system address the challenge of integrating parameter values of rear-stage models by using a conversion mechanism to estimate output data across differing front-stage models, resulting in improved accuracy and efficiency of model learning.

WO2025104827A1PCT designated stage expired Publication Date: 2025-05-22NEC CORP

Patent Information

Application Number
PCT/JP2023/040985
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-14
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing learning methods face challenges in efficiently integrating parameter values of rear-stage models across devices with differing front-stage models, which affects the accuracy and efficiency of model learning.

Method used

A learning device and system that include a conversion mechanism to estimate output data of one front-stage model from another, allowing for the learning of rear-stage models using this estimated data, and a server that integrates parameter values from multiple client devices.

Benefits of technology

This approach enables efficient integration of parameter values for rear-stage models even when front-stage models differ, improving the accuracy of model learning while maintaining data confidentiality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023040985_22052025_PF_FP_ABST
    Figure JP2023040985_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A learning device according to the present invention comprises: a conversion means for calculating estimation data, which is an estimated value of output data of a first pre-stage model with respect to input of training data, from output data of a second pre-stage model with respect to input of the training data; and a training means for performing training on a post-stage model by using the estimation data.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, learning system, learning method, and recording medium

[0001] The present invention relates to a learning device, a learning system, a learning method, and a recording medium.

[0002] A learning method known as federated learning has been proposed, in which multiple models having the same structure are trained and the parameter values ​​of these multiple models are integrated. Patent Document 1 also describes integrating parameter values ​​of multiple local models with different input image sizes. Patent Document 1 describes preparing a global model that can obtain feature maps with different resolutions, and sharing parameters between the global model and the local models. Patent Document 1 also describes integrating parameter sets for each local model in the global model, and transmitting the updated parameter set of the global model or a subset thereof to the local models.

[0003] JP 2023-042922 A

[0004] Each of the multiple devices has a model that is a combination of a trained front-end model and a rear-end model to be trained, and even if the front-end models differ depending on the device, it is preferable that the parameter values ​​of the rear-end model can be integrated as efficiently as possible.

[0005] An example of an object of the present invention is to provide a learning device, a learning system, a learning method, and a recording medium that can solve the above-mentioned problems.

[0006] According to a first aspect of the present invention, a learning device includes a conversion means for calculating estimated data, which is an estimated value of output data of a first front-end model in response to input training data, from output data of a second front-end model in response to input training data, and a learning means for learning a rear-end model using the estimated data.

[0007] According to a second aspect of the present invention, a learning system includes a plurality of client devices including a first client device and a second client device, and a server device, wherein the first client device includes a first learning means for learning a first subsequent model using output data of a first previous model in response to input of first training data, a first transmission means for transmitting parameter values ​​of the first subsequent model obtained by learning to the server device, and a first setting means for setting the parameter values ​​transmitted from the server device to the first subsequent model, and the second client device includes a first learning means for learning the output data of a second previous model in response to input of second training data, and a first setting means for setting the parameter values ​​transmitted from the server device to the first subsequent model. the server device includes a conversion means for converting the parameter values ​​of the second subsequent model obtained by learning into estimated data that is a constant value; a second learning means for using the estimated data to learn a second subsequent model; a second transmission means for transmitting to the server device parameter values ​​of the second subsequent model obtained by learning; and a second setting means for setting the parameter values ​​transmitted from the server device to the second subsequent model, and the server device includes an integration means for calculating a parameter value by integrating parameter values ​​transmitted from a plurality of client devices including the first client device and the second client device, and a server-side transmission means for transmitting the calculated parameter value to each of the plurality of client devices including the first client device and the second client device.

[0008] According to a third aspect of the present invention, a learning method includes a computer calculating estimated data, which is an estimate of output data of a first front-end model in response to input training data, from output data of a second front-end model in response to input training data, and learning a rear-end model using the estimated data.

[0009] According to a fourth aspect of the present invention, a learning method is provided for a learning system including a plurality of client devices, including a first client device and a second client device, and a server device, the method including: the first client device learning a first subsequent model using output data of a first pre-stage model in response to input of first training data, transmitting parameter values ​​of the first subsequent model obtained by learning to the server device, and setting the parameter values ​​transmitted from the server device to the first subsequent model; the second client device converting output data of a second pre-stage model in response to input of second training data into estimated data which is an estimate of the output data of the first pre-stage model in response to input of the second training data, learning a second subsequent model using the estimated data, transmitting parameter values ​​of the second subsequent model obtained by learning to the server device, and setting the parameter values ​​transmitted from the server device to the second subsequent model; and the server device calculating a parameter value by integrating parameter values ​​transmitted from the plurality of client devices, including the first client device and the second client device, and transmitting the calculated parameter value to each of the plurality of client devices, including the first client device and the second client device.

[0010] According to a fifth aspect of the present invention, the recording medium is a recording medium having recorded thereon a program that causes a computer to calculate estimated data, which is an estimate of output data of a first front-end model in response to input of training data, from output data of a second front-end model in response to input of that training data, and to learn a rear-end model using the estimated data.

[0011] According to the present invention, each of multiple devices has a model that is a combination of a trained upstream model and a downstream model to be trained, and even if the upstream models differ depending on the device, the parameter values ​​of the downstream model can be integrated relatively efficiently.

[0012] 1 is a diagram illustrating an example of the configuration of a learning system according to at least one embodiment; FIG. 2 is a diagram illustrating a first example of the configuration of a model included in a client device according to at least one embodiment; FIG. 3 is a diagram illustrating an example of the configuration of a client device according to at least one embodiment; FIG. 4 is a diagram illustrating an example of the configuration of a server device according to at least one embodiment; FIG. 5 is a diagram illustrating a second example of the configuration of a model included in a client device according to at least one embodiment; FIG. 6 is a diagram illustrating an example of the input and output of data in a server device according to at least one embodiment; FIG. 7 is a diagram illustrating a third example of the configuration of a model included in a client device according to at least one embodiment; FIG. 8 is a diagram illustrating an example of the input and output of data in a model when a client device according to at least one embodiment trains a converter; FIG. 9 is a diagram illustrating an example of the input and output of data in a model when a learning system according to at least one embodiment trains a converter by federated learning; FIG. 10 is a diagram illustrating an example of the combination of converters according to at least one embodiment; FIG. 11 is a diagram illustrating an example of the configuration of an estimation device according to at least one embodiment; FIG. 12 is a diagram illustrating an example of the input and output of data in a model when an estimation device according to at least one embodiment performs estimation; FIG. 13 is a diagram illustrating an example of the procedure of processing when a client device according to at least one embodiment trains a converter; FIG. 14 is a diagram illustrating an example of the procedure of processing when a client device according to at least one embodiment acquires input data for a subsequent model. FIG. 1 is a diagram showing an example of a processing procedure performed by a client device in federated learning of a subsequent-stage model by a learning system according to at least one embodiment. FIG. 2 is a diagram showing an example of a processing procedure performed by a server device in federated learning of a subsequent-stage model by a learning system according to at least one embodiment. FIG. 3 is a diagram showing an example of the configuration of a learning device according to at least one embodiment. FIG. 4 is a diagram showing an example of the configuration of a learning system according to at least one embodiment. FIG. 5 is a diagram showing an example of the configuration of an estimation device according to at least one embodiment. FIG. 6 is a diagram showing an example of the processing procedure in a learning method according to at least one embodiment. FIG. 7 is a diagram showing an example of the processing procedure in an estimation method according to at least one embodiment.FIG. 1 is a schematic block diagram illustrating the configuration of a computer according to at least one embodiment.

[0013] The following describes embodiments of the present invention, but the following embodiments do not limit the scope of the invention as claimed. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.

[0014] First Embodiment Fig. 1 is a diagram showing an example of the configuration of a learning system according to at least one embodiment. In the configuration shown in Fig. 1, learning system 1 includes n client devices 100 and a server device 200. Here, n is an integer greater than or equal to 2 and indicates the number of client devices 100. When distinguishing between the n client devices 100, they are also referred to as client devices 100-1, 100-2, ..., 100-n.

[0015] The learning system 1 learns a model held by the client device 100. Specifically, each client device 100 learns a model, and the server device 200 integrates parameter values ​​of the model obtained by learning for each client device 100. Each client device 100 updates the model by setting the integrated parameter values ​​(parameter values ​​obtained by integration) to the model.

[0016] The client device 100 is an example of a learning device. The parameter values ​​after integration are also referred to as integrated parameter values. Both the client device 100 and the server device 200 may be configured using a computer. Furthermore, both the client device 100 and the server device 200 may be configured as a single device or as a combination of multiple devices.

[0017] Model learning here refers to updating the parameter values ​​of a model (machine learning model) based on training data. Model learning here can be considered as model training. Note that the parameter values ​​of the model here may be a combination of multiple parameter values. For example, the parameter values ​​of the model may be represented as vector data.

[0018] Each client device 100 has a model that is a combination of a front-end model, which is a trained machine learning model, and a back-end model, which is a machine learning model to be trained. A model that is a combination of a front-end model and a back-end model is also called a combined model.

[0019] An example of a combined model configured by combining an early model and a later model is one in which the early model extracts features of input data, and the later model performs some kind of processing such as image recognition, person authentication, or class classification using the extracted features. Another example of a combined model configured by combining an early model and a later model is one in which the early model receives input text and performs morphological analysis, and the later model uses the analysis results to make inferences about the input text. However, the use of the combined model possessed by the client device 100 is not limited to a specific use.

[0020] The client device 100 acquires training data for the combined model and trains a subsequent model, which is a model to be trained among the combined models. In training the subsequent model, the client device 100 acquires output data of the previous model when input data to the combined model among the training data is input to the previous model.

[0021] The client device 100 uses a combination of this output data and data from the original training data other than the input data to the combined model, such as the correct answer labels, as training data for the subsequent model to train the subsequent model. The server device 200 integrates the parameter values ​​of the subsequent model.

[0022] The output data of the previous model is also referred to as intermediate data. Furthermore, inputting the input data of the training data to the model into the model is also referred to as inputting the training data to the model.

[0023] The combination of the learning of the subsequent model performed by the client device 100, the integration of the parameter values ​​of the subsequent model performed by the server device 200, and the setting of the integrated parameter values ​​to the subsequent model performed by the client device is also referred to as federated learning of the subsequent model. The federated learning here is a learning method in which multiple models are trained, parameter values ​​common to the multiple models are calculated based on the parameter values ​​of the multiple models, and the calculated parameter values ​​are set to one or more of the multiple models.

[0024] The associative learning of the subsequent model can also be regarded as the associative learning of a combined model that combines the subsequent model and the previous model. Calculating parameter values ​​common to multiple models is also called parameter value integration.

[0025] The learning system 1 can improve the accuracy of the model while ensuring data confidentiality. Specifically, each client device 100 can learn the subsequent-stage model, and there is no need to exchange training data between the client devices 100. In this respect, the learning system 1 can ensure data confidentiality. For example, consider a case where each client device 100 is owned by a different company, and the training data contains information that requires confidentiality, such as personal information or trade secrets. In this case, each client device 100 can learn the subsequent-stage model using training data obtained by its own company (the company that owns that client device 100), and there is no need to provide training data to other companies.

[0026] Furthermore, by integrating the parameter values, the server device 200 can reflect the training data used in learning for multiple client devices 100 in the integrated parameter values. In this way, the learning system 1 can learn a model using a relatively large amount of data, and in this respect, the accuracy of the model obtained by learning is expected to be relatively high.

[0027] The following describes an example in which the server device 200 integrates parameter values ​​of multiple subsequent models having the same structure. In this case, the server device 200 may use a parameter value integration method in a federated learning algorithm that targets multiple models having the same structure, such as FedAvg (Federated Averaging). For example, the server device 200 may integrate parameter values ​​by averaging the values ​​of parameters at the same position in the model structure.

[0028] However, the method by which the server device 200 integrates the parameter values ​​is not limited to a specific method. The technology of the present disclosure can also be applied to a case in which the server device 200 integrates the parameter values ​​of multiple subsequent models having different structures.

[0029] When multiple client devices 100 have subsequent models with different structures, at least one client device 100 is configured to also have a subsequent model with the same structure as the subsequent models possessed by the other client devices 100. This allows the multiple client devices 100 to learn subsequent models with the same structure, and allows the server device 200 to integrate parameter values ​​of subsequent models with the same structure.

[0030] Furthermore, even if multiple client devices 100 have subsequent models with the same structure, if the previous models are different, it is conceivable that the accuracy of the subsequent models will actually decrease if the server device 200 integrates the parameter values ​​of the subsequent models. Even in such a case, one client device 100 may have multiple subsequent models and switch the subsequent model to be used depending on the previous model.

[0031] Fig. 2 is a diagram showing a first example of the configuration of models possessed by the client device 100. Fig. 2 shows an example of the configuration of models and the input and output of data in those models when there are two client devices 100. In the example of Fig. 2, client devices 100-1 and 100-2 each have one front-stage model and two back-stage models.

[0032] One of the latter-stage models possessed by each client device 100 is provided for configuring a combined model in combination with the former-stage model. Of the latter-stage models possessed by each client device 100, the other latter-stage models are provided for performing federated learning of the combined models possessed by the other client devices 100.

[0033] The subscript number of "later model" indicates an index that identifies the later model, and indicates the correspondence of the later model in the integration of parameter values. This index (the subscript number of "later model") is also called the index of the later model.

[0034] The server device 200 integrates the indexes of subsequent models trained using the output data or an estimated value thereof of the same previous model and having the same structure. Subsequent models with the same index are assumed to have been trained using the output data or an estimated value thereof of the same previous model and to have the same structure.

[0035] However, as described above, the technology disclosed herein can also be applied when the server device 200 integrates parameter values ​​of multiple subsequent models having different structures (i.e., when the learning system 1 performs federated learning of multiple subsequent models having different structures).

[0036] In the example of FIG. 2, the server device 200 is a subsequent model of the client device 100-1. 1 and the parameter values ​​of the subsequent model of the client device 100-2. 1 The parameter values ​​of the second-stage model of the client device 100-1 are integrated. 1 and the subsequent stage model of the client device 100-2. 1 are assumed to have the same structure.

[0037] The server device 200 is a subsequent model of the client device 100-1. 2 and the parameter values ​​of the subsequent model of the client device 100-2. 2The parameter values ​​of the second-stage model of the client device 100-1 are integrated. 2 and the subsequent stage model of the client device 100-2. 2 are assumed to have the same structure.

[0038] Even if the structure of the subsequent model is the same, the parameter values ​​obtained by learning will differ due to differences in the training data, and these will be subject to integration of the parameter values ​​by the server device 200. The index of the training data used may be enclosed in < > and shown as a superscript.

[0039] For example, in the example of FIG. 2, the client device 100-1 receives training data 1 The client device 100-2 receives the training data 2 The subscript numbers in "training data" indicate indexes that identify the training data.

[0040] For any of the subsequent models that the client device 100-1 has, training data 1 The latter model is trained using 1 <1> and later models 2 <1> The client device 100-2 uses the training data 2 For any of the subsequent models that the client device 100-2 has, the training data 2 The latter model is trained using 1 <2> and later models 2 <2> In this way, in the case of a subsequent model, the superscript number enclosed in < > indicates the index of the training data used to learn the subsequent model.

[0041] Each of the client devices 100-1 and 100-2 has one pre-stage model. 1 The client device 100-2 has a front-end model 2The subscript number in the "previous model" indicates an index that identifies the previous model.

[0042] Furthermore, the index of the previous model and the index of the next model indicate the correspondence between the previous model and the next model. Specifically, a combination of a previous model and a next model with the same index number is used as a combined model. In the example of FIG. 2, the previous model 1 and the later model 1 The combination of these is used as a combined model, and the pre-model 2 and the later model 2 The combination of is used as a binding model.

[0043] It is assumed that the indices of the previous model are determined so that the above-described correspondence relationship is established between the indices of the previous model and the indices of the subsequent model. In any combined model, a subsequent model with the same structure may be used for the same previous model so that the above-described correspondence relationship is established.

[0044] Alternatively, when subsequent models with different structures are used for the same previous model, different index numbers may be assigned depending on the structure of the subsequent model, even if the previous model is the same. In this case, the index of the previous model can be considered as an index that identifies the combined model.

[0045] Pre-stage model 1 training data 1 The output data obtained by inputting is intermediate data. 1 <1> Also, the previous model 2 training data 2 The output data obtained by inputting is intermediate data. 2 <2> In this way, the subscript number of "intermediate data" indicates the index of the previous model that output the intermediate data. The subscript number of "intermediate data" is also referred to as the index of the previous model in the intermediate data.

[0046] As above, the superscript numbers in < > indicate the indices of the training data used. In the case of intermediate data, the superscript numbers in < > indicate the indices of the training data input to the previous model so that the previous model outputs the intermediate data.

[0047] The client device 100-1 is a front-end model. 1 whereas the latter model 2 <1> In order to learn 2 Therefore, the client device 100-1 needs the output data of the converter 1,2 It has a converter. 1,2 is the previous model 1 Intermediate data output by 1 The previous model 2 Intermediate data output by 2 Estimation data, which is an estimate of 2 Here, the converter is also called an encoder or a mapping model.

[0048] In this way, the converter i,j The notation indicates that the converter is a front-end model. i Intermediate data output by i The previous model j Intermediate data output by j Estimation data, which is an estimate of j The fact that the output data of the converter is called estimated data means that the output data of the converter does not have to completely match the output data of the previous model. A trained model may be used as the converter.

[0049] The subscript number in the "estimated data" indicates the index of the previous model in the intermediate data indicated by the estimated data. i is the previous model i Intermediate data output by iThe subscript number in "estimation data" is also referred to as the index of the previous model in the estimation data. As above, the superscript number in < > indicates the index of the training data used. In the case of estimation data, the superscript number in < > indicates the index of the training data in the intermediate data input to the converter so that the converter outputs the estimation data.

[0050] Similarly, the client device 100-2 is a front-end model. 2 whereas the latter model 1 <2> In order to learn 1 Therefore, the client device 100-2 needs the output data of the converter 2,1 It has a converter. 2,1 is the previous model 2 Intermediate data output by 2 The previous model 1 Intermediate data output by 1 Estimation data, which is an estimate of 1 Convert to.

[0051] Later model 1 The parameter values ​​of the model obtained by training are 1 In this way, the subscript number in the "parameter value" indicates that it is the parameter value of the subsequent model identified by the index of that number. i parameter value of i The subscript number in the "parameter value" is also referred to as the index of the subsequent model in the parameter value.

[0052] As mentioned above, the superscript numbers enclosed in < > indicate the indices of the training data used. In the case of parameter values, the superscript numbers enclosed in < > indicate that the parameter values ​​are obtained by learning using the training data identified by the index of that number.

[0053] Parameter Value 1 <1> is the latter model1 <1> Parameter value 2 <1> is the latter model 2 <1> Parameter value 2 <2> is the latter model 2 <2> Parameter value 1 <2> is the latter model 1 <2> The parameter values ​​are shown below.

[0054] As described above, the server device 200 is a downstream model of the client device 100-1. 1 and the parameter values ​​of the subsequent model of the client device 100-2. 1 Specifically, the server device 200 integrates the parameter value 1 <1> and parameter values 1 <2> Integrate with.

[0055] The client device 100-1 uses the integrated parameter values ​​as the second-stage model. 1 <1> By setting 1 <1> The client device 100-2 updates the integrated parameter value to the next-stage model. 1 <2> By setting 1 <2> When the parameter values ​​after integration are set, the subsequent model 1 <1> and the later model 1 <2> is a model with the same structure and the same parameter values.

[0056] After that, the client device 100-1 receives the intermediate data 1 <1> Using the latter model 1 <1> The client device 100-2 further performs learning of the estimated data 1 <2>Using the latter model 1 <2> This further trains the next model. 1 <1> and the later model 1 <2> and will again have different parameter values.

[0057] The training of the subsequent model, the integration of parameter values, and the setting of the integrated parameter values ​​in the subsequent model, or the repetition of these steps, is also referred to as associative learning of the subsequent model. As described above, associative learning here is a learning method in which multiple models are trained, parameter values ​​common to the multiple models are calculated based on the parameter values ​​of the multiple models, and the calculated parameter values ​​are set in one or more of the multiple models.

[0058] A combination of the previous model and the subsequent model after the federated learning is completed can be used as a trained combined model. The client device 100 may be configured to execute processing using a combined model formed by combining the previous model and the subsequent model after the federated learning is completed. Alternatively, a combined model formed by combining the previous model and the subsequent model after the federated learning is completed may be installed in a device other than the client device 100. As described above, the use of the combined model is not limited to a specific use.

[0059] In order to train the subsequent model, the client device 100 may acquire output data or estimated data of the previous model, using a previous model in addition to the method using a converter. For example, when the client device 100-1 acquires output data or estimated data of the previous model, 2 and the training data 1 The previous model 2 Enter intermediate data 2 <1> and the subsequent model 2 <1> It can also be used for learning.

[0060] However, with this method, if the scale of the pre-model is large, it takes time to input the training data into the pre-model and calculate the intermediate data, and the load on the client device 100 becomes heavy.

[0061] On the other hand, when the scale of the previous-stage model is large, it is considered possible to make the converter smaller than the scale of the previous-stage model. In this case, by using the converter in the client device 100, it is expected that the process of inputting intermediate data into the converter and calculating estimated data can be completed in a shorter time than when inputting training data into the previous-stage model to generate intermediate data. Furthermore, in this case, by using the converter in the client device 100, it is expected that the load on the client device 100 will be relatively small.

[0062] For example, the client device 100-1 is a front-end model. 2 and the training data 1 The previous model 2 Enter intermediate data 2 <1> Rather than calculating the intermediate data 1 <1> The converter 1,2 Enter the estimated data 2 <1> It is expected that the time required for the process will be shorter when calculating . Furthermore, it is expected that the load on the client device 100-1 will be smaller in the latter case than in the former case. In this respect, the client device 100 can relatively efficiently integrate the parameter values ​​of the subsequent model.

[0063] The pre-model, the post-model, and the combiner are not limited to a specific type of model, and for example, the pre-model, the post-model, and the combiner may be configured using a neural network, but are not limited to this.

[0064] Fig. 3 is a diagram showing an example of the configuration of client device 100. In the configuration shown in Fig. 3, client device 100 includes a client-side communication unit 110, a client-side display unit 120, a client-side operation input unit 130, a client-side storage unit 180, and a client-side processing unit 190. Client-side processing unit 190 includes a training data acquisition unit 191, a model calculation unit 192, a learning execution unit 196, a learning result transmission processing unit 197, and a parameter value update unit 198. Model calculation unit 192 includes a first-stage model calculation unit 193, a conversion unit 194, and a second-stage model calculation unit 195.

[0065] The client-side communication unit 110 communicates with other devices. For example, the client-side communication unit 110 transmits parameter values ​​obtained by learning the subsequent-stage model to the server device 200. The client-side communication unit 110 also receives integrated parameter values ​​from the server device 200.

[0066] The client-side display unit 120 has a display screen such as a liquid crystal panel or an LED (Light Emitting Diode) panel, and displays various images. For example, the client-side display unit 120 may display information related to the learning of the subsequent-stage model, such as the progress of the learning of the subsequent-stage model or the accuracy of the combined model.

[0067] The client-side operation input unit 130 includes input devices such as a keyboard and a mouse, and receives user operations. For example, the client-side operation input unit 130 may receive user operations for making settings related to the learning of the subsequent-stage model, such as setting a learning rate in the learning of the subsequent-stage model.

[0068] The client-side storage unit 180 stores various data. For example, the client-side storage unit 180 may store models such as a front-end model, a back-end model, and a converter. The client-side storage unit 180 may also store training data for model learning. The client-side storage unit 180 is configured using a storage device provided in the client device 100.

[0069] The client-side processing unit 190 performs various processes by controlling each unit of the client device 100. The functions of the client-side processing unit 190 may be performed by a CPU (Central Processing Unit) included in the client device 100 reading and executing a program from the client-side storage unit 180.

[0070] The training data acquisition unit 191 acquires training data. For example, if the client-side storage unit 180 stores the training data, the training data acquisition unit 191 may read the training data from the client-side storage unit 180.

[0071] The model calculation unit 192 performs calculations using various models. The previous-stage model calculation unit 193 inputs data to the previous-stage model and calculates intermediate data, which is output data of the previous-stage model. The conversion unit 194 inputs the intermediate data to a converter and calculates estimated data.

[0072] The subsequent model calculation unit 195 inputs intermediate data or estimated data to the subsequent model and calculates output data of the subsequent model. The subsequent model calculation unit 195 calculates output data of the subsequent model when training the subsequent model. Furthermore, when the client device 100 performs processing using a trained combined model, the subsequent model calculation unit 195 inputs intermediate data of the trained subsequent model and calculates output data of the subsequent model.

[0073] The learning execution unit 196 performs model learning. In particular, the learning execution unit 196 performs learning of a subsequent-stage model. As described above, model learning here refers to updating the parameter values ​​of the model based on training data.

[0074] For example, the learning execution unit 196 uses, as training data for the subsequent model, a combination of output data from the previous model when input data to the combined model out of the training data for the combined model is input to the previous model, and data other than the input data to the combined model, such as a correct answer label, from the original training data (training data for the combined model), to learn the subsequent model. The method used by the learning execution unit 196 to learn the model is not limited to a specific method. For example, the learning execution unit 196 may learn the model using a learning method such as backpropagation, but is not limited to this.

[0075] The learning result transmission processing unit 197 transmits the parameter values ​​of the subsequent model obtained by learning the subsequent model to the server device 200 via the client-side communication unit 110. The parameter values ​​of the subsequent model obtained by learning the subsequent model are also referred to as the learning results of the subsequent model.

[0076] The parameter value update unit 198 sets the integrated parameter values ​​calculated by the server device 200 in the subsequent model. Setting of the integrated parameter values ​​in the subsequent model by the parameter value update unit 198 can be considered as updating the parameter values ​​of the subsequent model. Setting of the integrated parameter values ​​in the subsequent model by the parameter value update unit 198 can also be considered as updating the subsequent model.

[0077] Fig. 4 is a diagram showing an example of the configuration of the server device 200. In the configuration shown in Fig. 4, the server device 200 includes a server-side communication unit 210, a server-side display unit 220, a server-side operation input unit 230, a server-side storage unit 280, and a server-side processing unit 290. The server-side processing unit 290 includes a parameter value acquisition unit 291, an integration unit 292, and an integration result transmission processing unit 293.

[0078] The server-side communication unit 210 communicates with other devices. For example, the server-side communication unit 210 receives parameter values ​​obtained by training the subsequent-stage model from the client device 100. The server-side communication unit 210 also transmits the integrated parameter values ​​to the client device 100.

[0079] The server-side display unit 220 has a display screen such as a liquid crystal panel or an LED panel, and displays various images. For example, the server-side display unit 220 may display information related to the integration of parameter values, such as the number of times the server device 200 has integrated parameters.

[0080] The server-side operation input unit 230 includes input devices such as a keyboard and a mouse, and receives user operations. For example, the server-side operation input unit 230 may receive user operations for setting parameters related to the integration of parameter values, such as a weighting coefficient value when the server device 200 performs a weighted average of parameter values ​​as the integration of parameter values.

[0081] The server-side storage unit 280 stores various data. For example, the server-side storage unit 280 may store parameter values ​​transmitted from each client device 100. The server-side storage unit 280 is configured using a storage device provided in the server device 200.

[0082] The server-side processing unit 290 performs various processes by controlling each unit of the server device 200. The functions of the server-side processing unit 290 may be performed by a CPU included in the server device 200 reading and executing a program from the server-side storage unit 280.

[0083] The parameter value acquisition unit 291 acquires parameter values ​​transmitted from each client device 100. For example, the parameter value acquisition unit 291 extracts parameter values ​​from the reception data that the server-side communication unit 210 receives from each client device 100.

[0084] The integrating unit 292 integrates the parameter values ​​transmitted from each client device 100. In particular, the integrating unit 292 integrates the parameter values ​​transmitted from each client device 100 for each group of parameter values. The group of parameter values ​​here may be a group obtained by grouping the parameter values ​​for each previous-stage model.

[0085] For example, in the case of FIG. 2, the integration unit 292 1 <1> and parameter values1 <2> and parameter values 1 <1> and parameter values 1 <2> Both are pre-stage models. 1 Specifically, the parameter value is 1 <1> is the previous model 1 The intermediate data is the output data of 1 <1> The latter model using 1 <1> The parameter values ​​obtained by learning are 1 <2> is the previous model 1 The estimated data is the estimated value of the output data 1 <2> The latter model using 1 <2> These are the parameter values ​​obtained by learning.

[0086] The integration unit 292 also 2 <1> and parameter values 2 <2> and parameter values 2 <1> and parameter values 2 <2> Both are pre-stage models. 2 Specifically, the parameter value is 2 <1> is the previous model 2 The estimated data is the estimated value of the output data 2 <1> The latter model using 2 <1> The parameter values ​​obtained by learning are 2 <2> is the previous model 2 The intermediate data is the output data of 2 <2> The latter model using 2 <2> These are the parameter values ​​obtained by learning.

[0087] Alternatively, when the combined model is configured using subsequent models with different structures by the client device 100, even if the previous model is the same, the groups of parameter values ​​may be groups obtained by grouping the parameter values ​​for each previous model and for each structure of the subsequent model. In this case, the client device 100 may use the same intermediate data or the same estimated data to train multiple subsequent models with different structures.

[0088] To group the parameter values, the client device 100 may attach labels indicating an index of the previous model and an index of the next model (an index that identifies the structure of the next model) to the parameter values ​​and transmit the labels to the server device 200. The integration unit 292 may then refer to the labels attached to the parameter values ​​and group the parameter values ​​by previous model and by structure of the next model.

[0089] A group of parameter values ​​can also be considered as a group of subsequent models to which the parameter values ​​are set. Hereinafter, a group of parameter values ​​will also be treated as a group of subsequent models, and will also be referred to as a group of subsequent models.

[0090] The integrating unit 292 may integrate parameter values ​​by weighting them, such as by taking a weighted average of parameter values ​​at the same position in the model structure. For example, the integrating unit 292 may set a larger weight for a parameter value that has been used in learning in a larger number of training data.

[0091] Furthermore, the integration unit 292 may set a larger weight for the parameter value obtained by learning using intermediate data compared with the parameter value obtained by learning using estimated data.

[0092] In order for the integrating unit 292 to weight and integrate the parameter values, the client device 100 may attach a label indicating information for weighting to the parameter values ​​and transmit them to the server device 200.

[0093] The integration result transmission processing unit 293 transmits the integrated parameter values ​​to the client device 100 via the server side communication unit 210. The integrated parameter values ​​are also referred to as an integration result of the parameter values.

[0094] 5 is a diagram showing a second example of the configuration of a model held by the client device 100. FIG. 5 shows an example of related models and data input / output in those models for one group of parameter values. In FIG. 5, the previous model used to generate the parameter values ​​of the group being displayed is shown as the previous model. 1 It is expressed as follows.

[0095] m is the previous model 1 , and m is an integer greater than or equal to 1, indicating the number of client devices 100 having the 1 The client devices 100 having the same predecessor model are represented as client devices 100-1 to 100-m. n indicates the number of client devices provided in the learning system 1, and is an integer of n≧2. Also, n>m. p indicates the number of predecessor models provided in the learning system 1 when the same predecessor model is counted as one, and is an integer of p≧2. Also, p≦m.

[0096] If i is an integer in the range of 1≦i≦m, the client device 100-i is a pre-stage model 1 training data i Enter the intermediate data 1 <i> Then, the client device 100-i calculates the intermediate data 1 <i> Using the latter model 1 <i> The parameter values ​​obtained by learning are 1 <i> is transmitted to the server device 200.

[0097] Here, j is an integer satisfying m+1≦j≦n, and the previous model held by the client device 100-j is represented as previous model k, where k is an integer satisfying 2≦k≦p.

[0098] The client device 100-j is a front-end modelk training data j Enter the intermediate data k <j> Then, the client device 100-j calculates the intermediate data k <j> The converter k,1 Enter the estimated data 1 <j> The client device 100-j calculates the estimated data 1 <j> Using the latter model 1 <j> The parameter values ​​obtained by learning are 1 <j> The server device 200 transmits the parameter value 1 <1> to parameter value 1 <n> Integrate up to.

[0099] There may be a client device 100 that does not participate in federated learning for parameter values ​​included in a certain group. For example, in the example of FIG. 1 There may be a client device 100 that does not perform learning.

[0100] Of the subsequent models included in one group, the subsequent model included in the combined model is also referred to as the first subsequent model, and the other subsequent models are also referred to as the second subsequent model. 1 <1> From the later model 1 <m> Up to this point, this corresponds to the first latter stage model example, and the latter stage model 1 <m+1> From the later model 1 <n> The above corresponds to an example of the second latter stage model.

[0101] For one group of later-stage models, a client device 100 having a first later-stage model is referred to as a first client device, and a client device 100 having a second later-stage model is referred to as a second client device. In the example of Figure 5, client devices 100-1 to 100-m are examples of the first client device, and client devices 100-m+1 to 100-n are examples of the second client device.

[0102] The upstream model possessed by the first client device is also referred to as a first upstream model, and the upstream model possessed by the second client device is also referred to as a second upstream model. 1 corresponds to an example of the first pre-stage model, and pre-stage model 2 to pre-stage model p The above corresponds to an example of the second pre-stage model.

[0103] The output data of the first pre-stage model is also referred to as the first output model, and the output data of the second pre-stage model is also referred to as the second output data. 1 <1> From intermediate data 1 <m> Up to this point corresponds to an example of the first output data, and 2 <m+1> From intermediate data p <n> The above corresponds to an example of the second output data.

[0104] The learning execution unit 196 of the first client device is also referred to as first learning means. The learning execution unit 196 of the second client device is also referred to as second learning means. The combination of the learning result transmission processing unit 197 and the client-side communication unit 110 of the first client device is also referred to as first transmission means. The combination of the learning result transmission processing unit 197 and the client-side communication unit 110 of the second client device is also referred to as second transmission means. The parameter value update unit 198 of the first client device is also referred to as first setting means. The parameter value update unit 198 of the second client device is also referred to as second setting means.

[0105] The training data for the combined model of the first client device is also referred to as first training data.1 The training data m corresponds to an example of the first training data. The training data for the combined model of the second client device is also referred to as the second training data. In the example of FIG. 5, the training data m+1 The training data n through n correspond to examples of second training data. Note that Fig. 5 does not show the subsequent model included in the combined model of the second client device.

[0106] 6 is a diagram showing an example of data input / output in the server device 200. FIG. 6 shows an example of data input / output in the server device 200 for one group of parameter values. In the example of FIG. 6, the server device 200 i <1> to parameter value i <n> These parameter values ​​are integrated based on the inputs up to

[0107] Here, n is an integer greater than or equal to 2, indicating the number of client devices 100 included in the learning system 1. i is an integer greater than or equal to 1, indicating the index of the subsequent model that is the target of parameter value integration shown in FIG. 6.

[0108] The server device 200 receives the parameter value i <1> to parameter value i <n> The parameter value is calculated by integrating these parameter values. i <1> to parameter value i <n> The parameter value is the parameter value i ’ It can also be written as:

[0109] The server device 200 receives the integrated parameter value i ’ , the parameter value i <1> to parameter value i <n> The data is transmitted to each of the client devices 100 as the transmission source, and the subsequent model i <1> From the later modeli <n> The integrated parameter values ​​for each i ’ Set the following.

[0110] Fig. 7 is a diagram showing a third example of the configuration of models possessed by the client device 100. Fig. 7 shows an example of models possessed by one client device 100 and data input / output in those models. The target client device shown in Fig. 7 is designated as client device 100-i.

[0111] In the example of FIG. 7, the client device 100-i is a front-end model j and the converter j,1 From the converter j,j-1 Up to and converter j,j+1 From the converter j,p Up to and later models 1 <i> From the later model p <i> It has up to.

[0112] Here, p is an integer such that 2≦p≦n indicates the number of previous models that the learning system 1 has, when the same previous model is counted as one. Here, n is an integer such that n≧2 indicates the number of client devices 100 that the learning system 1 has. Here, i is an integer such that 1≦i≦n. Here, j is an integer such that 1≦j≦p.

[0113] The training data acquisition unit 191 acquires training data i and obtain the training data i to the pre-stage model calculation unit 193. The pre-stage model calculation unit 193 outputs the training data i The previous model j Enter intermediate data i Calculate the intermediate data i are output to the conversion unit 194 and the learning execution unit 196.

[0114] The conversion unit 194 converts the training data i The converter j,1 From the converter j,j-1 Up to and converter j,j+1 From the converter j,pEnter the estimated data 1 <i> Estimated data from j-1 <i> Up to and estimated data j+1 <i> Estimated data from p <i> The conversion unit 194 calculates the estimated data 1 <i> Estimated data from j-1 <i> Up to and estimated data j+1 <i> Estimated data from p <i> are output to the learning execution unit 196.

[0115] The learning execution unit 196 uses the intermediate data j <i> Using the latter model j <i> The learning execution unit 196 also performs learning of the subsequent model. 1 <i> From the later model j-1 <i> Up to and after models j+1 <i> From the later model p <i> For the period up to k <i> Using the latter model k <i> Here, k is an integer such that 1≦k≦p and k≠j.

[0116] The learning execution unit 196 causes the subsequent model calculation unit 195 to calculate output data of the subsequent model, and performs learning of the subsequent model using the obtained output data. 1 <i> to parameter value p <i> The learning result transmission processing unit 197 outputs the parameter values ​​obtained by learning. 1 <i> to parameter value p <i> The above is transmitted to the server device 200 via the client-side communication unit 110.

[0117] The client device 100 may train the converter. In this case, the client device 100 trains the converter before starting training of the subsequent model.

[0118] In relation to one group of the subsequent model, when the client device 100-i in FIG. 7 corresponds to an example of the first client device, the training data i corresponds to the first training data example, and the previous model j corresponds to an example of the first front-end model. j <i> corresponds to an example of a first subsequent stage model. Furthermore, the learning execution unit 196 of the client device 100-i corresponds to an example of a first learning means, and the combination of the client-side communication unit 110 of the client device 100-i and the learning result transmission processing unit 197 corresponds to an example of a first transmission means.

[0119] In relation to one group of the subsequent model, when the client device 100-i in FIG. 7 corresponds to an example of the second client device, the training data i corresponds to the second training data example, and the previous model j corresponds to an example of the second front-end model. 1 <i> From the later model j-1 <i> Up to and after the model j+1 <i> From the later model p <i> Any one of the above corresponds to an example of a second subsequent stage model. In addition, the learning execution unit 196 of the client device 100-i corresponds to an example of a second learning means, and the combination of the client-side communication unit 110 of the client device 100-i and the learning result transmission processing unit 197 corresponds to an example of a second transmission means.

[0120] 8 is a diagram showing an example of input and output of data in a model when the client device 100 performs learning of a converter. j,kIn the example shown, when the client device 100-i uses multiple converters during training of the subsequent model, the client device 100-i may perform training of each of the multiple converters in advance.

[0121] In the example of FIG. 8, the client device 100-i is a front-end model j And the previous model k and the converter j,k The previous model j and the converter j,k is a model that the client device 100-i also uses when learning the subsequent model. k The client device 100-i does not use the previous model when training the converter. k may be temporarily included.

[0122] The first-stage model calculation unit 193 calculates the first-stage model j training data i Enter the intermediate data j <i> The previous model calculation unit 193 calculates the previous model k training data i Enter the intermediate data k <i> The learning execution unit 196 calculates the following. j,k Intermediate data j <i> The converter when you input j,k The estimated data is the output data of k <i> However, the intermediate data k <i> The converters j and k are trained so that:

[0123] The learning execution unit 196 causes the conversion unit 194 to calculate output data of the converter, and uses the obtained output data to train the converter. As described above, the method used by the learning execution unit 196 for training the converter is not limited to a specific method. For example, the learning execution unit 196 may train the model using a learning method such as backpropagation, but is not limited to this.

[0124] Intermediate data calculated by the conversion unit 194 during converter learning j <i> The client-side storage unit 180 may store the intermediate data . Then, when the subsequent model is being learned, the client device 100-i (for example, the previous model calculation unit 193) reads the intermediate data . j <i> may be read out and used. This allows the preceding model calculation unit 193 to omit, or a part of, the process of inputting training data to the preceding model and calculating output data when learning the subsequent model. This is expected to reduce the calculation time and the load on the preceding model calculation unit 193.

[0125] Pre-stage model j When there are multiple client devices 100 having a combined model including (i.e., (ii)), only one or some of these multiple client devices 100 may be configured to train the combiner. Then, the client device 100 that has trained the combiner may transmit (provide) the trained combiner to the client device 100 that has not trained the combiner. This reduces the load of combiner training on the entire learning system 1.

[0126] The server device 200 may perform converter training, and one device may perform converter training collectively for use by multiple client devices 100. In this case, it is possible that the training data acquired by each client device 100 cannot be used for converter training for some reason, such as data confidentiality. When the training data acquired by each client device 100 cannot be used for converter training, open data (publicly available data) or composite data in which temporary values ​​are set for confidential parts may be used to train the converter.

[0127] The learning system 1 may be configured to perform learning of the converter by federated learning. Fig. 9 is a diagram showing an example of input and output of data in a model when the learning system 1 performs learning of the converter by federated learning.j,k This shows an example of performing federated learning.

[0128] Here, j and k are integers satisfying 1≦j, k≦n and j≠k. Here, n is an integer satisfying n≧2, which indicates the number of client devices 100 included in the learning system 1. Note that FIG. 9 shows the converters of all the client devices 100 included in the learning system 1. j,k This shows an example of learning the previous model. j Only the client device 100 having a binding model including j,k Two or more client devices 100 may be configured to perform learning of the converter. j,k It is best to have the students learn the following.

[0129] In the example of FIG. 9, each of the client devices 100-1 to 100-n includes a converter, similar to that described with reference to FIG. j,k Then, the learning result transmission processing unit 197 of each of the client devices 100-1 to 100-n transmits the converter obtained by the learning via the client-side communication unit 110. j,k The parameter values ​​are transmitted to the server device 200.

[0130] In the example of FIG. 9, a converter obtained by learning in the client device 100-i j,k parameter value of j,k <i> Here, i is an integer such that 1≦i≦n. As mentioned above, a superscript number enclosed in <> indicates the index of the training data used. In the case of a parameter value, a superscript number enclosed in <> indicates that the parameter value is a parameter value obtained by learning using the training data identified by the index of that number.

[0131] In the server device 200, the integrating unit 292 integrates the parameter values ​​transmitted from each client device 100. As in the case of the associative learning of the subsequent model, in the associative learning of the converter, the integrating unit 292 may also integrate the parameter values ​​by weighting them, such as by taking a weighted average of the parameter values ​​at the same position in the model structure. For example, the integrating unit 292 may set a larger weight for a parameter value that has been used in the learning process in a larger number of training data sets.

[0132] The integrated result transmission processing unit 293 of the server device 200 transmits the integrated parameter values ​​to the client device 100 via the server side communication unit 210. In the example of FIG. j,k parameter value of j,k In each of the client devices 100, the parameter value update unit 198 updates the parameter value j,k ' to converter j,k Set to.

[0133] When training the subsequent model, the conversion unit 194 may calculate the estimated data by connecting a plurality of converters. Fig. 10 is a diagram showing an example of the combination of converters. Fig. 10 shows an example in which the conversion unit 194 of the client device 100 uses the converters j,l and converter l,k Combined with the converter j,k This shows an example of using it as:

[0134] Here, i is an integer such that 1≦i≦n. Here, n is an integer such that n≧2, which indicates the number of client devices 100 included in the learning system 1. Also, here, j, k, and l are integers such that 1≦j, k, and l≦p, and j≠k≠l. Here, p is an integer such that p≧2, which indicates the number of previous models included in the learning system 1 when the same previous model is counted as one. Also, in the client device 100-i, the intermediate data calculated by the previous model calculation unit 193 is referred to as intermediate data j <i> Let's say.

[0135] The conversion unit 194 converts the intermediate data j <i> The converterj,l Enter the estimated data l <i> Furthermore, the conversion unit 194 calculates the intermediate data l <i> The converter l,k Enter the estimated data k <i> In this way, the estimated data output by the converter is input to another converter to calculate the estimated data, which is also called the combination of the converters.

[0136] The conversion unit 194 is a converter j,l and converter l,k Combined with the converter j,k By using it as a converter j,k In this respect, the load on the learning system 1 during learning of the converter can be reduced.

[0137] Intermediate data j <i> is the previous model j In the example of FIG. 10, the output data of the previous model j corresponds to an example of the second pre-model. l <i> is the previous model l In the example of FIG. 10, the output data of the previous model l corresponds to an example of the third pre-stage model. Estimation data k <i> is the previous model k In the example of FIG. 10, the output data of the previous model k corresponds to an example of the first pre-stage model.

[0138] Fig. 11 is a diagram illustrating an example of the configuration of an estimation device according to at least one embodiment. In the configuration illustrated in Fig. 11, the estimation device 300 includes an estimation-side communication unit 310, an estimation-side display unit 320, an estimation-side operation input unit 330, an estimation-side storage unit 380, and an estimation-side processing unit 390. The estimation-side processing unit 390 includes an estimation target data acquisition unit 391, a model calculation unit 392, and an estimation result output processing unit 395. The model calculation unit 392 includes a front-stage model calculation unit 393 and a rear-stage model calculation unit 394.

[0139] The estimation device 300 is a device that performs estimation using the combined model trained by the learning system 1. The estimation device 300 may be configured using a computer. Furthermore, the estimation device 300 may be configured as a single device or as a combination of multiple devices.

[0140] The estimation-side communication unit 310 communicates with other devices. For example, the estimation-side communication unit 310 may receive data to be estimated, such as sensing data obtained by a sensor.

[0141] The estimation-side display unit 320 has a display screen such as a liquid crystal panel or an LED panel, and displays various images. For example, the estimation-side display unit 320 may display the estimation results obtained by the estimation device 300.

[0142] The estimation-side operation input unit 330 includes input devices such as a keyboard and a mouse, and receives user operations. For example, the estimation-side operation input unit 330 may receive a user operation to instruct the execution of estimation.

[0143] The estimation storage unit 380 stores various data. For example, the estimation storage unit 380 may store a trained combined model. The estimation storage unit 380 is configured using a storage device included in the estimation device 300.

[0144] The estimating-side processing unit 390 performs various processes by controlling each unit of the estimating device 300. The functions of the estimating-side processing unit 390 may be performed by a CPU included in the estimating device 300 reading and executing a program from the estimating-side storage unit 380.

[0145] The estimation target data acquisition unit 391 acquires estimation target data. The estimation target data here refers to data that is to be estimated by the estimation device 300. For example, the estimation target data acquisition unit 391 may extract the estimation target data from data received by the estimation-side communication unit 310.

[0146] The model calculation unit 392 performs calculations using the trained binding model. The calculations performed by the model calculation unit 392 using the binding model correspond to estimation by the estimation device 300. The previous-stage model calculation unit 393 inputs the estimation target data to a previous-stage model among the trained binding models and calculates intermediate data. The subsequent-stage model calculation unit 394 inputs the intermediate data to a subsequent-stage model among the trained binding models and calculates output data of the subsequent-stage model. The output data of the subsequent-stage model calculated by the subsequent-stage model calculation unit 394 corresponds to the estimation result by the estimation device 300.

[0147] The estimation result output processing unit 395 outputs the estimation result by the estimation device 300 (output data of the subsequent-stage model calculated by the subsequent-stage model calculation unit 394). The method by which the estimation result output processing unit 395 outputs the estimation result is not limited to a specific method. For example, the estimation result output processing unit 395 may display the estimation result on the estimation-side display unit 320. Alternatively, the estimation result output processing unit 395 may transmit the estimation result to another device via the estimation-side communication unit 310.

[0148] In the learning system 1, a trained combined model is obtained for each client device 100. It is possible to configure n estimating devices 300 corresponding to the n client devices 100 provided in the learning system 1. When distinguishing between the individual estimating devices 300, they are also referred to as estimating device 300-1, estimating device 300-2, ..., estimating device 300-n.

[0149] The estimation device 300 may be configured as one of the functions of the client device 100. In this case, the client-side communication unit 110 may be caused to execute the function of the estimation-side communication unit 310, the client-side display unit 120 may be caused to execute the function of the estimation-side display unit 320, the client-side operation input unit 130 may be caused to execute the function of the estimation-side operation input unit 330, the client-side storage unit 180 may be caused to execute the function of the estimation-side storage unit 380, and the client-side processing unit 190 may be caused to execute the function of the estimation-side processing unit 390. Alternatively, the training data acquisition unit 191 may be caused to execute the function of the estimation target data acquisition unit 391, the model calculation unit 192 may be caused to execute the function of the model calculation unit 392, the previous-stage model calculation unit 193 may be caused to execute the function of the previous-stage model calculation unit 393, the subsequent-stage model calculation unit 195 may be caused to execute the function of the subsequent-stage model calculation unit 394, and the learning result transmission processing unit 197 may be caused to execute the function of the estimation result output processing unit 395. Alternatively, the estimation device 300 may be configured as a device separate from the client device 100 .

[0150] Fig. 12 is a diagram showing an example of input and output of data in a model when estimation is performed by the estimation device 300. Fig. 12 shows an example when estimation is performed by an estimation device 300-i. Here, i is an integer satisfying 1≦i≦n. Here, n is an integer satisfying n≧2, indicating the number of estimation devices.

[0151] The estimation device 300-i is a device that performs estimation using a trained combined model acquired by the client device 100-i. The combined model possessed by the estimation device 300-i is a pre-model j and the later model j Here, j is an integer in the range of 1≦j≦p, and p is an integer in the range of 2≦p≦n, which indicates the number of previous models included in the learning system 1 when the same previous model is counted as one.

[0152] The estimation target data acquisition unit 391 acquires estimation target data and outputs the acquired estimation target data to the previous-stage model calculation unit 393. The estimation target data acquired by the estimation target data acquisition unit 391 of the estimation device 300-i is referred to as estimation target data jThe previous model calculation unit 393 calculates the previous model j Estimation target data j Enter the intermediate data j <i> Here, we calculate the previous model j Estimation target data i The output data obtained by inputting is intermediate data. j <i> It is written as follows.

[0153] The latter model calculation unit 394 calculates the latter model j Intermediate data j <i> The estimation results are calculated by inputting the following: j Intermediate data j <i> The output data obtained by inputting the estimation results j <i> The subsequent model calculation unit 394 calculates the estimation result as j <i> to the estimation result output processing unit 395. The estimation result output processing unit 395 outputs the estimation result j <i> As described above, the method by which the estimation result output processing unit 395 outputs the estimation result is not limited to a specific method.

[0154] 13 is a diagram showing an example of a processing procedure when the client device 100 learns a converter. The client device 100 performs the processing shown in FIG. 13 for each converter to be learned. In the processing shown in FIG. 13, the training data acquisition unit 191 acquires training data (step S101).

[0155] Next, the previous-stage model calculation unit 193 inputs training data to the previous-stage model and calculates intermediate data that is output data of the previous-stage model (step S102). Here, the previous-stage model calculation unit 193 calculates intermediate data that will be input data to the converter and intermediate data that will be used as a correct answer to the output data of the converter. In the example of FIG. 8, the previous-stage model calculation unit 193 calculates the intermediate data that will be input data to the converter and intermediate data that will be used as a correct answer to the output data of the converter. j,k Intermediate data that will be input to j <i> and the converter j,kThe intermediate data used as the correct answer for the output data k <i> The following is calculated:

[0156] Next, the learning execution unit 196 performs learning of the converter using the intermediate data calculated by the previous-stage model calculation unit 193 (step S103). Specifically, the learning execution unit 196 performs learning of the converter so that when the intermediate data calculated as input data to the converter is input to the converter, the output data of the converter becomes the intermediate data calculated as the correct answer to the output data of the converter. After step S103, the client device 100 ends the processing of FIG. 13.

[0157] 14 is a diagram showing an example of a processing procedure in which the client device 100 acquires input data for the subsequent model. Each client device 100 performs the processing shown in FIG. 14 as preprocessing before learning the subsequent model. In the processing shown in FIG. 14, the training data acquisition unit 191 acquires training data for the combined model (step S201).

[0158] Next, the previous-stage model calculation unit 193 inputs the training data for the combined model to the previous-stage model and calculates intermediate data, which is output data of the previous-stage model (step S202). As described above, inputting input data to the model from the training data to the model is referred to as inputting training data to the model.

[0159] Next, the conversion unit 194 converts the intermediate data, which is the output data of the previous model, into estimated data, which is an estimated value of the output data of another previous model (step S203). j,1 From the converter j,j-1 Up to, and converter j,j+1 From the converter j,p Intermediate data for each j <i> Enter the estimated data 1 <i> Estimated data from j-1 <i> Up to, and estimated data j+1 <i> Estimated data from p <i>It is calculated up to.

[0160] After step S203, the client device 100 ends the processing in Fig. 14. The client device 100 uses the training data or estimated data obtained in the processing in Fig. 14 as input data to the subsequent model, and can use data obtained by combining this input data with parts of the training data for the combined model other than the input data to the combined model, such as the correct answer label, as training data for the subsequent model.

[0161] 15 is a diagram showing an example of the procedure of processing performed by the client device 100 in federated learning of subsequent models by the learning system 1. Each client device 100 performs the processing of FIG. 15 for each subsequent model included in that client device 100.

[0162] In the processing of Fig. 15, the learning execution unit 196 executes learning of the subsequent model using training data including the intermediate data or estimated data obtained in the processing of Fig. 14 (step S211). The learning execution unit 196 executes learning of the subsequent model included in the combined model using the intermediate data, and executes learning of the other subsequent models using the estimated data.

[0163] In the example of FIG. 7, the learning execution unit 196 j The latter model that forms a combined model in combination with j <i> The learning of intermediate data j <i> The learning execution unit 196 uses the following model: 1 <i> From the later model j-1 <i> Up to and after the model j+1 <i> From the later model p <i> The learning up to the estimation data 1 <i> Estimated data from j-1 <i> Up to, and estimated data j+1 <i> Estimated data from p <i> This is done using up to.

[0164] Next, the learning result transmission processing unit 197 transmits the parameter values ​​obtained by learning to the server device 200 via the client-side communication unit 110 (step S212). Next, the client-side communication unit 110 receives the integrated parameter values ​​transmitted from the server device 200 (step S213). Then, the parameter value update unit 198 rewrites the parameter values ​​of each subsequent-stage model to the integrated parameter values ​​(step S214).

[0165] Next, the learning execution unit 196 determines whether the termination condition of the subsequent-stage model federated learning is satisfied (step S215). The termination condition here is not limited to a specific condition. Alternatively, only one of the multiple client devices 100 and the server device 200 may determine whether the termination condition is satisfied using the termination condition judgment criteria. In this case, the device that made the judgment using the judgment criteria notifies the other devices of the judgment result. The device that received the notification determines whether the termination condition is satisfied according to the notification.

[0166] For example, one client device 100 having a combined model including the latter-stage model that is the target of the processing in Fig. 15 may calculate the estimation accuracy of the latter-stage model using training data, and determine that the termination condition is met if the calculated estimation accuracy is equal to or greater than a predetermined threshold. Furthermore, each time the client device 100 determines whether the termination condition is met, it may notify the other client devices and the server device of the determination result.

[0167] Alternatively, a determination criterion may be used that enables each of the multiple client devices 100 and the server device 200 to make the same determination. For example, the termination condition here may be a condition that the number of times the server device 200 has executed the integration of parameter values ​​for the group of subsequent-stage models that is the subject of the processing in Fig. 15 reaches a threshold value common to each client device 100 and the server device 200. Each client device 100 may then determine for itself whether the termination condition is met, and the server device 200 may also determine for itself whether the termination condition is met.

[0168] If the learning execution unit 196 determines that the termination condition is not met (step S215: NO), the process returns to step 211. On the other hand, if the learning execution unit 196 determines that the termination condition is met (step S215: YES), the client device 100 ends the process of FIG.

[0169] 16 is a diagram showing an example of the procedure of processing performed by the server device 200 in federated learning of the subsequent model by the learning system 1. In the processing of FIG. 16, the server-side communication unit 210 receives parameter values ​​of the subsequent model transmitted from each of the client devices 100 (step S301).

[0170] Next, the integration unit 292 integrates the parameter values ​​of the subsequent model (step S302). In the example of FIG. 6, the integration unit 292 integrates the parameter values ​​of the subsequent model. i <1> to parameter value i <1> The parameter value after integration is i Next, the integrated result transmission processing unit 293 transmits the integrated parameter values ​​to each of the client devices 100 via the server-side communication unit 210 (step S303).

[0171] Next, the integration unit 292 determines whether the termination condition of the subsequent-stage model federated learning is satisfied (step S304). As described above, the termination condition here is not limited to a specific condition. Alternatively, only one of the multiple client devices 100 and the server device 200 may determine whether the termination condition is satisfied using the termination condition determination criteria. Alternatively, the multiple client devices 100 and the server device 200 may use determination criteria that allow each of them to make the same determination, and each client device 100 may determine whether the termination condition is satisfied on its own, and the server device 200 may also determine whether the termination condition is satisfied on its own.

[0172] If the integrating unit 292 determines that the termination condition is not met (step S304: NO), the process returns to step 301. On the other hand, if the integrating unit 292 determines that the termination condition is met (step S304: YES), the server device 200 ends the process of FIG.

[0173] As described above, the conversion unit 194 converts the training data i The previous model for the input k The estimated data is the estimated value of the output data k <i> the previous model for the training data input. j The intermediate data is the output data of j <i> The learning execution unit 196 calculates the estimated data k <i> Using the latter model k <i> The previous model is trained. k corresponds to an example of the first pre-stage model. j corresponds to an example of the second pre-stage model.

[0174] According to the client device 100, the previous model j The intermediate data is the output data of j <i> Based on the pre-stage model k The estimated data is the estimated value of the output data k <i> Therefore, the previous model k There is no need to separately obtain output data for the subsequent model. k <i> The latter model can be trained. k The parameter values ​​of can be integrated.

[0175] In this regard, according to the client device 100, each of the multiple client devices 100 has a combined model that is a combination of a trained front-end model and a rear-end model to be trained, and even if the front-end models differ depending on the client device 100, the parameter values ​​of the rear-end model can be integrated relatively efficiently.

[0176] The conversion unit 194 also converts the training datai The previous model for the input j The intermediate data is the output data of j <i> the training data i The previous model for the input l The estimated data is the estimated value of the output data l <i> and estimate the data l <i> Estimate the data k <i> Convert to.

[0177] In the client device 100, the latter model k <i> Estimation data for training k <i> A converter j,l and converter l,k and intermediate data j <i> Estimated data from k <i> Converter for directly calculating j,k In this respect, the client device 100 can reduce the load of learning on the converter.

[0178] The conversion unit 194 also converts the previous model j The intermediate data is the output data of j <i> A converter constructed using a machine learning model j,k The input data to the previous model k The training data is the input data to i The previous model for the same input data k The output data is converted j,k A converter trained using training data that is the correct answer for the output data of j,k The client device 100 can generate a converter through learning. In this respect, the client device 100 can reduce the burden of designing a converter.

[0179] The conversion unit 194 also converts the training data i Converter using j,kThe parameter values ​​obtained by learning and the training data h Converter using j,k A converter with parameter values ​​calculated based on the parameter values ​​obtained by learning j,k The data is converted using

[0180] Here, i and h are integers such that 1≦i, h≦p and i≠h. Here, p is an integer p≧2 that indicates the number of previous models included in the learning system 1 when the same previous model is counted as one. It is expected that the client device 100 will be able to integrate the parameter values ​​of multiple converters to improve the accuracy of the converters.

[0181] Second Embodiment Fig. 17 is a diagram showing an example of the configuration of a learning device according to at least one embodiment. In the configuration shown in Fig. 17, a learning device 610 includes a conversion unit 611 and a learning unit 612.

[0182] With this configuration, the conversion unit 611 calculates estimated data, which is an estimate of the output data of the first pre-stage model in response to the input of training data, from the output data of the second pre-stage model in response to the input of that training data. The learning unit 612 uses the estimated data to train the post-stage model. The conversion unit 611 is an example of a conversion means. The learning unit 612 is an example of a learning means.

[0183] According to the learning device 610, an estimated value of the output data of the first front-stage model can be obtained based on the output data of the second front-stage model, so that the subsequent model can be learned without the need to separately obtain the output data of the first front-stage model, and the parameter values ​​of the subsequent model can be integrated.

[0184] In this regard, according to the learning device 610, each of the multiple learning devices has a combined model that combines a trained upstream model and a downstream model to be trained, and even if the upstream models differ depending on the learning device, the parameter values ​​of the downstream model can be integrated relatively efficiently.

[0185] The conversion unit 611 can be realized, for example, by using the function of the conversion unit 194 in Fig. 3. The learning unit 612 can be realized, for example, by using the function of the learning execution unit 196 in Fig. 3.

[0186] 18 is a diagram showing an example of the configuration of a learning system according to at least one embodiment. In the configuration shown in Fig. 18, learning system 620 includes server device 630 and multiple client devices including first client device 640 and second client device 650.

[0187] The server device 630 includes an integration unit 631 and a server-side transmission unit 632. The first client device 640 includes a first learning unit 641, a first transmission unit 642, and a first setting unit 643. The second client device 650 includes a conversion unit 651, a second learning unit 652, a second transmission unit 653, and a second setting unit 654.

[0188] With this configuration, the first learning unit 641 learns the first subsequent model using output data of the first previous model in response to the input of the first training data. The first transmission unit 642 transmits parameter values ​​of the first subsequent model obtained by the learning to the server device 630. The first setting unit 643 sets the parameter values ​​transmitted from the server device 630 in the first subsequent model.

[0189] The conversion unit 651 converts the output data of the second front-end model in response to the input of second training data into estimated data, which is an estimate of the output data of the first front-end model in response to the input of the second training data. The second learning unit 652 uses the estimated data to learn the second back-end model. The second transmission unit 653 transmits parameter values ​​of the second back-end model obtained by learning to the server device 630. The second setting unit 654 sets the parameter values ​​transmitted from the server device 630 in the second back-end model.

[0190] The integration unit 631 calculates a parameter value by integrating parameter values ​​transmitted from a plurality of client devices including the first client device 640 and the second client device 650. The server-side transmission unit 632 transmits the calculated parameter value to each of the plurality of client devices including the first client device 640 and the second client device 650.

[0191] The first learning unit 641 corresponds to an example of a first learning means. The first transmitting unit 642 corresponds to an example of a first transmitting means. The first setting unit 643 corresponds to an example of a first setting means. The converting unit 651 corresponds to an example of a converting means. The second learning unit 652 corresponds to an example of a second learning means. The second transmitting unit 653 corresponds to an example of a second transmitting means. The second setting unit 654 corresponds to an example of a second setting means. The integrating unit 631 corresponds to an example of an integrating means. The server-side transmitting unit 632 corresponds to an example of a server-side transmitting means.

[0192] According to the learning system 620, an estimate of the output data of the first front-stage model can be obtained based on the output data of the second front-stage model, so that the subsequent-stage model can be trained without the need to separately obtain output data of the first front-stage model, and the parameter values ​​of the subsequent-stage model can be integrated. In this respect, according to the learning system 620, each of the multiple client devices has a combined model that combines a trained previous-stage model with a subsequent-stage model to be trained, so that even if the previous-stage models differ depending on the client device, the parameter values ​​of the subsequent-stage models can be integrated relatively efficiently.

[0193] 19 is a diagram illustrating an example of the configuration of an estimation device according to at least one embodiment. In the configuration illustrated in FIG. 19, an estimation device 660 includes a first-stage model calculation unit 661 and a second-stage model calculation unit 662.

[0194] In this configuration, the first-stage model calculation unit 661 inputs estimation target data to the first first-stage model and calculates the output of the first first-stage model. The second-stage model calculation unit 662 inputs the output data of the first first-stage model to the first second-stage model, in which parameter values ​​are set that combine parameter values ​​of the first second-stage model obtained by learning using output data of the first first-stage model in response to input of first training data and parameter values ​​of the second second-stage model obtained by learning using estimation data, which is data obtained by converting output data of the second first-stage model in response to input of second training data into estimates of the output data of the first first-stage model, and calculates an estimation result to be output by the first second-stage model. The first-stage model calculation unit 661 corresponds to an example of a first-stage model calculation means. The second-stage model calculation unit 662 corresponds to an example of a second-stage model calculation means.

[0195] The estimation device 660 can perform estimation using a subsequent model in which integrated parameter values ​​are set, and in this respect, high estimation accuracy is expected. Furthermore, the estimation device 660 can obtain estimated values ​​of the output data of the first previous model based on the output data of the second previous model during joint learning of the parameter values ​​of the subsequent model, so that the subsequent model can be trained without having to separately obtain output data of the first previous model. In this respect, the estimation device 660 can perform estimation with relatively high accuracy using parameter values ​​of the subsequent model that are relatively efficiently integrated.

[0196] Fifth Embodiment Fig. 20 is a diagram showing an example of a processing procedure in a learning method according to at least one embodiment. The learning method shown in Fig. 20 includes performing conversion (step S611) and performing learning (step S612).

[0197] In the conversion step (step S611), the computer calculates estimated data, which is an estimate of the output data of the first front-end model in response to the input of training data, from the output data of the second front-end model in response to the input of the training data. In the training step (step S612), the computer trains the rear-end model using the estimated data.

[0198] According to the learning method shown in Figure 20, an estimated value of the output data of the first front-stage model can be obtained based on the output data of the second front-stage model, so that the subsequent model can be learned without the need to separately obtain the output data of the first front-stage model, and the parameter values ​​of the subsequent model can be integrated.

[0199] In this regard, according to the learning method shown in Figure 20, each of the multiple learning devices has a combined model that combines a trained previous model and a subsequent model to be learned, and even if the previous models differ depending on the learning device, the parameter values ​​of the subsequent model can be integrated relatively efficiently.

[0200] Sixth Embodiment Fig. 21 is a diagram showing an example of a processing procedure in a learning method according to at least one embodiment, in which a learning system including a plurality of client devices, including a first client device and a second client device, and a server device performs learning.

[0201] The learning method shown in FIG. 21 includes a first client device performing learning of a first subsequent model (step S621), transmitting parameter values ​​of the first subsequent model (step S622), and setting the parameter values ​​to the first subsequent model (step S623); a second client device performing conversion (step S631), performing learning of a second subsequent model (step S632), transmitting parameter values ​​of the second subsequent model (step S633), and setting the parameter values ​​to the second subsequent model (step S634); and a server device performing integration (step S641) and transmitting the integrated parameter values ​​(step S642).

[0202] In executing learning of a first subsequent model (step S621), the first client device learns the first subsequent model using output data of the first pre-model in response to input of first training data. In transmitting parameter values ​​of the first subsequent model (step S622), the first client device transmits the parameter values ​​of the first subsequent model obtained by learning to the server device. In setting parameter values ​​in the first subsequent model (step S623), the first client device sets the parameter values ​​transmitted from the server device to the first subsequent model.

[0203] In performing conversion (step S631), the second client device converts output data of the second front-end model in response to the input of the second training data into estimated data that is an estimate of the output data of the first front-end model in response to the input of the second training data. In performing training of the second back-end model (step S632), the second client device trains the second back-end model using the estimated data.

[0204] In transmitting parameter values ​​of the second subsequent model (step S633), the second client device transmits the parameter values ​​of the second subsequent model obtained by learning to the server device. In setting parameter values ​​in the second subsequent model (step S634), the second client device sets the parameter values ​​transmitted from the server device in the second subsequent model.

[0205] In the integration step (step S641), the server device calculates a parameter value that integrates parameter values ​​transmitted from a plurality of client devices, including the first client device and the second client device. In the transmission step (step S642), the server device transmits the calculated parameter value to each of the plurality of client devices, including the first client device and the second client device.

[0206] 21 , an estimate of the output data of the first front-stage model can be obtained based on the output data of the second front-stage model, so that the subsequent-stage model can be trained without the need to separately obtain output data of the first front-stage model, and the parameter values ​​of the subsequent-stage model can be integrated. In this respect, according to the learning method shown in Fig. 21 , each of the multiple client devices has a combined model that combines a trained previous-stage model and a subsequent-stage model to be trained, so that even if the previous-stage models differ depending on the client device, the parameter values ​​of the subsequent-stage models can be integrated relatively efficiently.

[0207] Seventh Embodiment Fig. 22 is a diagram showing an example of a processing procedure in an estimation method according to at least one embodiment. The estimation method shown in Fig. 22 includes calculating output data of a previous model (step S651) and calculating output data of a next model (step S652).

[0208] In calculating output data of the previous model (step S651), the computer inputs estimation target data to the previous model and calculates an output of the previous model. In calculating output data of the subsequent model (step S652), the computer inputs the output of the previous model to a subsequent model that has been trained using at least estimation data calculated from second output data that is the output of a second previous model in response to the input of training data, and calculates an output of the subsequent model.

[0209] According to the estimation method shown in Fig. 22, estimation can be performed using a subsequent model in which integrated parameter values ​​are set, and in this respect, high estimation accuracy is expected. Furthermore, according to the estimation method shown in Fig. 22, during joint learning of the parameter values ​​of the subsequent model, estimated values ​​of the output data of the first previous model can be obtained based on the output data of the second previous model, so that learning of the subsequent model can be performed without the need to separately obtain output data of the first previous model. In this respect, according to the estimation method shown in Fig. 22, estimation can be performed with relatively high accuracy using parameter values ​​of the subsequent model that are relatively efficiently integrated.

[0210] There is no particular limitation on the target to which learning system 1, estimation device 300, learning device 610, learning system 620, or estimation device 660 is applied. Below, an example of a target for learning by learning system 1 will be described, but the same applies to a target for estimation by estimation device 300, a target for learning by learning device 610, a target for learning by learning system 620, and a target for estimation by estimation device 660.

[0211] The learning system 1 may be configured to learn a combined model that receives input of information about a borrower, such as the borrower's debt status or business status, or a combination of these, regarding loans of money, etc., and outputs the magnitude of loan default risk. In this case, a client device 100 may be provided for each branch of a financial institution or each affiliate of a financial institution, and training data for each branch of the financial institution or each affiliate of a financial institution may be obtained.

[0212] The learning system 1 may be configured to learn a combined model that receives input of information about the insured, such as the insured's medical history, age, blood pressure, or genetic information, or a combination of these, and outputs insurance premiums. In this case, a client device 100 may be provided for each insurance company branch or each insurance company, and training data may be obtained for each insurance company branch or each insurance company.

[0213] The learning system 1 may receive input of information about a patient, such as information written in a medical record, and perform learning of a combined model that outputs the cause of a disease, a treatment method, etc. In this case, a client device 100 may be provided for each medical institution, such as a clinic, and training data for each medical institution may be acquired.

[0214] The learning system 1 may receive information on the chemical structure of a substance, such as information on the structure of a compound, information on the structure of a ligand, information on the structure of a protein, or a combination thereof, and perform learning of a binding model that outputs predicted information on the activity of the substance, such as a compound. In this case, the client device 100 may acquire training data for each pharmaceutical company, such as activity data for compound libraries held by each pharmaceutical company.

[0215] The learning system 1 may be configured to receive input of information about a factory, such as the manufacturing status of products in a machine factory or the like, the transportation status of products or materials in the factory, or a combination of these, and perform learning of a combined model that outputs control command values ​​for devices or systems used in the factory, such as industrial robots used in the machine factory. In this case, a client device 100 may be provided for each factory or each company, and training data for each factory or each company may be acquired.

[0216] The learning system 1 may receive information about a transportation device, such as an image of the surroundings taken by the transportation device, such as an automobile, a railroad vehicle, an aircraft, or a ship, location information about the transportation device, and the degree of congestion of a route that the transportation device can travel, or a combination of these, and may perform learning of a combined model that outputs a control command value for the transportation device or a control command value for a signal to be transmitted to the transportation device. In this case, a client device 100 may be provided for each transportation device, and training data for each transportation device may be acquired.

[0217] The learning system 1 may be configured to receive input of information about a case, such as the circumstances of the crime in the case, the circumstances of the evidence in the case, legal information related to the case, or information on precedents related to the case, or a combination of these, and then perform training of a combined model that outputs information to support the progress of the trial of the case, such as a calculated value for the sentence for the case or an estimated value for the suspect's recidivism rate. In this case, a client device 100 may be provided for each court, and training data for each court may be acquired.

[0218] The learning system 1, the learning device 610, or the learning system 620 may have a configuration for communicating parameter values ​​using homomorphic encryption or the like, so that the server device can integrate parameter values ​​while maintaining the confidentiality of the parameter values ​​of the subsequent model. Alternatively, the learning system 1, the learning device 610, or the learning system 620 may have a configuration for encrypting and decrypting parameter values ​​between the client device or the learning device and the server device.

[0219] To reduce the possibility of training data being leaked due to an attack, noise may be added to the training data used by the learning system 1, the learning device 610, or the learning system 620. Furthermore, noise may be added to the parameter values ​​of the subsequent model transmitted and received by the learning system 1, the learning device 610, or the learning system 620. Furthermore, noise may be added to the parameter values ​​of the subsequent model used by the learning system 1, the learning device 610, or the learning system 620.

[0220] The learning system 1, the learning device 610, or the learning system 620 may have a configuration for reducing the size of the transmitted data when transmitting parameter values ​​of the subsequent model. For example, the learning system 1, the learning device 610, or the learning system 620 may have a configuration for performing data compression.

[0221] Furthermore, learning system 1, learning device 610, or learning system 620 may transmit and receive only a portion of the parameter values ​​of the subsequent model, such as transmitting and receiving only 10 percent of the parameter values ​​of the subsequent model. In this case, the parameter values ​​transmitted and received by learning system 1, learning device 610, or learning system 620 may be selected randomly. Alternatively, learning system 1, learning device 610, or learning system 620 may transmit and receive predetermined parameter values.

[0222] A configuration may be provided for reducing the size of a model by reconstructing the model used by learning system 1, estimation device 300, learning device 610, learning system 620, or estimation device 660. For example, a configuration may be provided for reducing the size of a model using techniques such as a distilled model, a derived model, a pseudo model, or a higher-level model.

[0223] For example, if the processing power of a client device or learning device is relatively small, a smaller model size can reduce the load on the client device or learning device, and is expected to prevent the load on the client device or learning device from becoming excessive.

[0224] Furthermore, when the client device or learning device has limited power available, such as when the client device or learning device is configured as a wearable device, the reduced size of the model is expected to reduce the power consumption of the client device or learning device and avoid the inability to supply the required power.

[0225] Even when the amount of power available to a client device or a learning device is limited, the smaller model size can reduce the power consumption of the client device or learning device, and is expected to prevent the client device or learning device from malfunctioning due to a lack of electrical energy. The same applies to an estimation device that uses a model obtained by learning using a client device or a learning device.

[0226] The function for reconstructing the model may be provided as one of the functions of the client device or the learning device, or as one of the functions of the server device, or alternatively, the function for reconstructing the model may be provided as a function of a device external to the client device or the learning device and the server device.

[0227] In the configuration of learning system 1, a party that owns server device 200 (e.g., a business operator) may provide the functions of server device 200 to a party that uses client device 100 (e.g., a business operator), thereby allowing the processing of learning system 1 to be performed. Alternatively, a party that owns server device 200 may receive training data from a party that owns the training data, or a portion of such a party, and the processing of learning system 1 may be performed under a system that maintains the confidentiality of the training data. Here, the party that owns the training data is a party that uses client device 100, or an equivalent party.

[0228] Regarding learning system 620, a party that owns server device 630 (e.g., a business operator) may provide the functions of server device 630 to a party that owns a client device (e.g., a business operator), thereby allowing processing of learning system 620 to be performed. Alternatively, a party that owns server device 630 may receive training data from a party that owns the training data, or a portion of such a party, and processing of learning system 620 may be performed under a system that maintains the confidentiality of the training data. Here, the party that owns the training data is a person that uses a client device or an equivalent party.

[0229] 23 is a schematic block diagram illustrating the configuration of a computer according to at least one embodiment. In the configuration shown in FIG. 23, a computer 700 includes a CPU 710, a main memory device 720, an auxiliary memory device 730, and an interface 740.

[0230] One or more of the client device 100, server device 200, estimation device 300, learning device 610, server device 630, first client device 640, second client device 650, and estimation device 660, or a part thereof, may be implemented in the computer 700. In this case, the operation of each of the above-described processing units is stored in the auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program. The CPU 710 also allocates storage areas in the main storage device 720 corresponding to each of the above-described storage units in accordance with the program. Communication between each device and other devices is performed by the interface 740, which has a communication function, and performs communication under the control of the CPU 710.

[0231] When the client device 100 is implemented in a computer 700, the operations of the client-side processing unit 190 and each of its units are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0232] Furthermore, the CPU 710 allocates a storage area for the client-side storage unit 180 in the main storage device 720 in accordance with the program. Communication with other devices by the client-side communication unit 110 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the client-side display unit 120 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Acceptance of user operations by the client-side operation input unit 130 is performed by the interface 740 having an input device and accepting user operations under the control of the CPU 710.

[0233] When the server device 200 is implemented in the computer 700, the operations of the server-side processing unit 290 and each of its units are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0234] Furthermore, the CPU 710 allocates a storage area for the server-side storage unit 280 in the main storage device 720 in accordance with the program. Communication with other devices by the server-side communication unit 210 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the server-side display unit 220 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the server-side operation input unit 230 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0235] When the estimating device 300 is implemented in a computer 700, the operations of the estimation processing unit 390 and each unit thereof are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0236] Furthermore, the CPU 710 allocates a storage area for the inference-side storage unit 380 in the main storage unit 720 in accordance with the program. Communication with other devices by the inference-side communication unit 310 is performed by the interface 740, which has a communication function and operates under the control of the CPU 710. Display of images by the inference-side display unit 320 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the inference-side operation input unit 330 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0237] When the learning device 610 is implemented in the computer 700, the operations of the conversion unit 611 and the learning unit 612 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0238] Furthermore, CPU 710 allocates a storage area in main memory 720 for learning device 610 to perform processing in accordance with the program. Communication between learning device 610 and other devices is achieved by interface 740 having a communication function and performing communication under the control of CPU 710. Interaction between learning device 610 and a user is achieved by interface 740 having a display device and an input device, displaying various images under the control of CPU 710, and accepting user operations.

[0239] When the server device 630 is implemented in the computer 700, the operations of the integrating unit 631 and the server-side transmitting unit 632 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0240] Furthermore, the CPU 710, in accordance with the program, allocates a storage area in the main storage device 720 for the server device 630 to perform processing. Communication between the server device 630 and other devices, such as transmission performed by the server-side transmitting unit 632, is performed by the interface 740 having a communication function and performing communication under the control of the CPU 710. Interaction between the server device 630 and a user is performed by the interface 740 having a display device and an input device, displaying various images under the control of the CPU 710, and accepting user operations.

[0241] When the first client device 640 is implemented in the computer 700, the operations of the first learning unit 641, the first transmission unit 642, and the first setting unit 643 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0242] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the first client device 640 to perform processing in accordance with the program. Communication between the first client device 640 and other devices, such as transmission performed by the first transmission unit 642, is performed by the interface 740 having a communication function and performing communication under the control of the CPU 710. Interaction between the first client device 640 and a user is performed by the interface 740 having a display device and an input device, displaying various images under the control of the CPU 710, and accepting user operations.

[0243] When the second client device 650 is implemented in the computer 700, the operations of the conversion unit 651, the second learning unit 652, the second transmission unit 653, and the second setting unit 654 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0244] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the second client device 650 to perform processing in accordance with the program. Communication between the second client device 650 and other devices, such as transmission performed by the second transmission unit 653, is performed by the interface 740 having a communication function and performing communication under the control of the CPU 710. Interaction between the second client device 650 and a user is performed by the interface 740 having a display device and an input device, displaying various images under the control of the CPU 710, and accepting user operations.

[0245] When the estimation device 660 is implemented in the computer 700, the operations of the front-stage model calculation unit 661 and the rear-stage model calculation unit 662 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0246] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the processing of the estimation device 660 in accordance with the program. Communication between the estimation device 660 and other devices is performed by the interface 740, which has a communication function and performs communication under the control of the CPU 710. Interaction between the estimation device 660 and a user is performed by the interface 740, which has a display device and an input device, displaying various images under the control of the CPU 710 and accepting user operations.

[0247] One or more of the above-described programs may be recorded on nonvolatile recording medium 750. In this case, interface 740 may read the programs from nonvolatile recording medium 750. Then, CPU 710 may directly execute the programs read by interface 740, or may temporarily store the programs in main storage device 720 or auxiliary storage device 730 and then execute them.

[0248] Note that a program for executing all or part of the processing performed by the client device 100, server device 200, estimation device 300, learning device 610, server device 630, first client device 640, second client device 650, and estimation device 660 may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform the processing of each unit. Note that the term "computer system" herein includes hardware such as an operating system (OS) and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, read-only memories (ROMs), and compact disc read-only memories (CD-ROMs), as well as storage devices such as hard disks built into the computer system. The program may be for implementing part of the aforementioned functions, or may be capable of implementing the aforementioned functions in combination with a program already recorded on the computer system.

[0249] Although the embodiments have been described above, the specific configuration is not limited to these embodiments, and the present invention also includes designs that do not deviate from the gist of the present invention. Furthermore, the above-described embodiments can be combined with other embodiments as appropriate.

[0250] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.

[0251] (Supplementary Note 1) A learning device comprising: a conversion means for calculating estimated data, which is an estimated value of output data of a first front-end model in response to input training data, from output data of a second front-end model in response to input training data; and a learning means for learning a rear-end model using the estimated data.

[0252] (Supplementary Note 2) The conversion means converts output data of the second front-end model in response to the input of the training data into an estimated value of output data of a third front-end model in response to the input of the training data, and converts the estimated value of the output data of the third front-end model into the estimated data.

[0253] (Supplementary Note 3) The learning device according to Supplementary Note 1, wherein the conversion means converts data using a machine learning model trained using training data in which output data of the second pre-stage model is used as input data to a machine learning model, and output data of the first pre-stage model for input data that is the same as the input data to the second pre-stage model is used as a correct answer to output data of the machine learning model.

[0254] (Supplementary Note 4) The learning device according to any one of Supplementary Notes 1 to 3, wherein the conversion means converts data using a machine learning model in which parameter values ​​calculated based on parameter values ​​obtained by training the machine learning model using first training data and parameter values ​​obtained by training the machine learning model using second training data are set.

[0255] (Supplementary Note 5) A system comprising: a plurality of client devices including a first client device and a second client device; and a server device; wherein the first client device comprises: first learning means for learning a first subsequent model using output data of a first previous model in response to input of first training data; first transmission means for transmitting parameter values ​​of the first subsequent model obtained by learning to the server device; and first setting means for setting the parameter values ​​transmitted from the server device to the first subsequent model; wherein the second client device comprises: conversion means for converting output data of a second previous model in response to input of second training data into estimated data which is an estimated value of the output data of the first previous model in response to input of the second training data; second learning means for learning a second subsequent model using the estimated data; second transmission means for transmitting parameter values ​​of the second subsequent model obtained by learning to the server device; and second setting means for setting the parameter values ​​transmitted from the server device to the second subsequent model; and wherein the server device comprises: integration means for calculating a parameter value obtained by integrating parameter values ​​transmitted from a plurality of client devices including the first client device and the second client device; A learning system comprising: a server-side transmitting means for transmitting the calculated parameter value to each of a plurality of client devices including the first client device and the second client device.

[0256] (Supplementary Note 6) The learning system described in Supplementary Note 5, wherein the conversion means converts output data of the second front-end model in response to the input of the training data into an estimated value of output data of a third front-end model in response to the input of the training data, and converts the estimated value of the output data of the third front-end model into the estimated data.

[0257] (Supplementary Note 7) The learning system described in Supplementary Note 5, wherein the conversion means converts data using a machine learning model trained using training data in which output data of the second front-end model is used as input data to a machine learning model, and output data of the first front-end model for input data that is the same as the input data to the second front-end model is used as a correct answer to output data of the machine learning model.

[0258] (Supplementary Note 8) The learning system described in any one of Supplementary Notes 5 to 7, wherein the conversion means converts data using a machine learning model in which parameter values ​​calculated based on parameter values ​​obtained by training the machine learning model using first training data and parameter values ​​obtained by training the machine learning model using second training data are set.

[0259] (Supplementary Note 9) An estimation device comprising: a front-stage model calculation unit that inputs estimation target data to a front-stage model and calculates an output of the front-stage model; and a rear-stage model calculation unit that inputs an output of the front-stage model to a rear-stage model that is trained using at least estimation data, the estimation data being an estimate of an output of a first front-stage model in response to an input of training data, calculated from second output data being the output of a second front-stage model in response to the input of the training data, and calculates an output of the rear-stage model.

[0260] (Supplementary Note 10) A learning method including: a computer calculating estimated data, which is an estimate of output data of a first front-end model in response to input training data, from output data of a second front-end model in response to input training data; and learning a back-end model using the estimated data.

[0261] (Supplementary Note 11) The learning method according to Supplementary Note 10, wherein performing the conversion includes converting, by the computer, output data of the second front-end model in response to the input of the training data into an estimate of output data of a third front-end model in response to the input of the training data, and converting the estimate of the output data of the third front-end model into the estimated data.

[0262] (Supplementary Note 12) The learning method according to Supplementary Note 10, wherein performing the conversion includes converting data using a machine learning model trained using training data in which output data of the second pre-stage model is used as input data to a machine learning model, and output data of the first pre-stage model for input data that is the same as the input data to the second pre-stage model is used as a correct answer to output data of the machine learning model.

[0263] (Supplementary Note 13) The learning method according to any one of Supplementary Notes 10 to 12, wherein performing the conversion includes converting the data using a machine learning model in which parameter values ​​calculated based on parameter values ​​obtained by training the machine learning model using first training data and parameter values ​​obtained by training the machine learning model using second training data are set by the computer.

[0264] (Supplementary Note 14) A learning system including a plurality of client devices including a first client device and a second client device, and a server device, comprising: the first client device: learning a first subsequent model using output data of a first previous model in response to input of first training data; transmitting parameter values ​​of the first subsequent model obtained by learning to the server device; and setting the parameter values ​​transmitted from the server device to the first subsequent model; the second client device: converting output data of a second previous model in response to input of second training data into estimated data which is an estimate of the output data of the first previous model in response to input of the second training data, learning a second subsequent model using the estimated data; transmitting parameter values ​​of the second subsequent model obtained by learning to the server device; and setting the parameter values ​​transmitted from the server device to the second subsequent model; the server device: calculating a parameter value by integrating parameter values ​​transmitted from a plurality of client devices including the first client device and the second client device; transmitting the calculated parameter value to each of a plurality of client devices, including the first client device and the second client device.

[0265] (Supplementary Note 15) The learning method according to Supplementary Note 14, wherein performing the conversion includes converting, by the computer, output data of the second front-end model in response to the input of the training data into an estimate of output data of a third front-end model in response to the input of the training data, and converting the estimate of the output data of the third front-end model into the estimated data.

[0266] (Supplementary Note 16) The learning method according to Supplementary Note 14, wherein performing the conversion includes converting data using a machine learning model trained using training data in which output data of the second pre-stage model is used as input data to a machine learning model, and output data of the first pre-stage model for input data that is the same as the input data to the second pre-stage model is used as a correct answer to output data of the machine learning model.

[0267] (Supplementary Note 17) The learning method according to any one of Supplementary Notes 14 to 16, wherein performing the conversion includes converting the data using a machine learning model in which parameter values ​​calculated based on parameter values ​​obtained by training the machine learning model using first training data and parameter values ​​obtained by training the machine learning model using second training data are set by the computer.

[0268] (Supplementary Note 18) An estimation method including: a computer inputting data to be estimated into a front-stage model and calculating an output of the front-stage model; and a rear-stage model trained using at least estimated data, the estimated value of the output of a first front-stage model in response to input of training data, calculated from second output data, the output of a second front-stage model in response to input of the training data, to calculate an output of the rear-stage model.

[0269] (Supplementary Note 19) A recording medium having recorded thereon a program that causes a computer to execute the following: calculating estimated data, which is an estimate of output data of a first front-end model in response to input of training data, from output data of a second front-end model in response to input of the training data; and learning a back-end model using the estimated data.

[0270] (Supplementary Note 20) The recording medium described in Supplementary Note 19, wherein in performing the conversion, the program causes the computer to convert output data of the second front-end model in response to the input of the training data into an estimated value of output data of a third front-end model in response to the input of the training data, and convert the estimated value of the output data of the third front-end model into the estimated data.

[0271] (Supplementary Note 21) The recording medium described in Supplementary Note 19, wherein in performing the conversion, the program causes the computer to convert data using a machine learning model trained using training data in which output data of the second pre-stage model is used as input data to a machine learning model, and output data of the first pre-stage model for input data that is the same as the input data to the second pre-stage model is used as a correct answer to output data of the machine learning model.

[0272] (Supplementary Note 22) The recording medium described in any one of Supplementary Notes 19 to 21, wherein in performing the conversion, the program causes the computer to convert data using a machine learning model in which parameter values ​​calculated based on parameter values ​​obtained by training the machine learning model using first training data and parameter values ​​obtained by training the machine learning model using second training data are set.

[0273] (Supplementary Note 23) A recording medium having recorded thereon a program for causing a computer to execute the following steps: inputting data to be estimated into a front-stage model and calculating the output of the front-stage model; and inputting the output of the front-stage model into a back-stage model that has been trained using at least estimated data, the estimated value of the output of a first front-stage model in response to input of training data, calculated from second output data that is the output of a second front-stage model in response to input of that training data, and calculating the output of the back-stage model.

[0274] The present invention may be applied to a learning device, a learning system, a learning method, and a recording medium.

[0275] 1, 620 Learning system 100 Client device 110 Client-side communication unit 120 Client-side display unit 130 Client-side operation input unit 180 Client-side storage unit 190 Client-side processing unit 191 Training data acquisition unit 192, 392 Model calculation unit 193, 393, 661 Pre-stage model calculation unit 194, 611, 651 Conversion unit 195, 394, 662 Subsequent-stage model calculation unit 196 Learning execution unit 197 Learning result transmission processing unit 198 Parameter value update unit 200, 630 Server device 210 Server-side communication unit 220 Server-side display unit 230 Server-side operation input unit 280 Server-side storage unit 290 Server-side processing unit 291 Parameter value acquisition unit 292, 631 Integration unit 293 Integration result transmission processing unit 300, 660 Estimation device 310 Estimation-side communication unit 320 Estimation-side display unit 330 Estimation-side operation input unit 380 Estimation-side storage unit 390 Estimation-side processing unit 391 Estimation target data acquisition unit 395 Estimation result output processing unit 610 Learning device 612 Learning unit 632 Server-side transmission unit 640 First client device 641 First learning unit 642 First transmission unit 643 First setting unit 650 Second client device 652 Second learning unit 653 Second transmission unit 654 Second setting unit

Claims

1. A learning device comprising: a conversion means for calculating estimated data, which is an estimate of the output data of a first front-stage model in response to input training data, from the output data of a second front-stage model in response to the input training data; and a learning means for learning a rear-stage model using the estimated data.

2. The learning device according to claim 1, wherein the conversion means converts output data of the second front-stage model in response to the input of the training data into an estimate of output data of a third front-stage model in response to the input of the training data, and converts the estimate of the output data of the third front-stage model into the estimated data.

3. The learning device described in claim 1, wherein the conversion means converts data using a machine learning model trained using training data in which the output data of the second front-stage model is used as input data to a machine learning model, and the output data of the first front-stage model for input data that is the same as the input data to the second front-stage model is used as the correct answer for output data of the machine learning model.

4. A learning device as claimed in any one of claims 1 to 3, wherein the conversion means converts data using a machine learning model in which parameter values ​​are set that are calculated based on parameter values ​​obtained by training a machine learning model using first training data and parameter values ​​obtained by training a machine learning model using second training data.

5. A system comprising: a plurality of client devices including a first client device and a second client device; and a server device, wherein the first client device comprises: a first learning means for learning a first subsequent model using output data of a first subsequent model in response to input of first training data; a first transmission means for transmitting parameter values ​​of the first subsequent model obtained by learning to the server device; and a first setting means for setting the parameter values ​​transmitted from the server device to the first subsequent model; wherein the second client device comprises: a conversion means for converting output data of a second subsequent model in response to input of second training data into estimated data which is an estimated value of the output data of the first subsequent model in response to the input of the second training data; a second learning means for learning a second subsequent model using the estimated data; a second transmission means for transmitting parameter values ​​of the second subsequent model obtained by learning to the server device; and a second setting means for setting the parameter values ​​transmitted from the server device to the second subsequent model; wherein the server device comprises: an integration means for calculating a parameter value obtained by integrating parameter values ​​transmitted from a plurality of client devices including the first client device and the second client device; A learning system comprising: a server-side transmitting means for transmitting the calculated parameter value to each of a plurality of client devices including the first client device and the second client device.

6. A learning method comprising: a computer calculating estimated data, which is an estimate of the output data of a first front-stage model in response to input training data, from the output data of a second front-stage model in response to the input training data; and training a front-stage model using the estimated data.

7. A learning system including a plurality of client devices including a first client device and a second client device, and a server device, comprising: the first client device: learning a first subsequent stage model using output data of a first pre-stage model in response to input of first training data; transmitting parameter values ​​of the first subsequent stage model obtained by learning to the server device; and setting the parameter values ​​transmitted from the server device to the first subsequent stage model; the second client device: converting output data of a second pre-stage model in response to input of second training data into estimated data which is an estimate of the output data of the first pre-stage model in response to input of the second training data, learning a second subsequent stage model using the estimated data; transmitting parameter values ​​of the second subsequent stage model obtained by learning to the server device; and setting the parameter values ​​transmitted from the server device to the second subsequent stage model; the server device: calculating a parameter value by integrating parameter values ​​transmitted from the plurality of client devices including the first client device and the second client device; and transmitting the calculated parameter value to each of the plurality of client devices including the first client device and the second client device; Including, how to learn.

8. A recording medium having recorded thereon a program that causes a computer to execute the following: calculating estimation data, which is an estimate of the output data of a first front-stage model in response to input of training data, from the output data of a second front-stage model in response to input of the training data; and training a front-stage model using the estimation data.

Citation Information

Patent Citations

  • Bank capital satisfaction rate prediction method and device, electronic equipment and medium

    CN112613978A

  • Information processing system, information processor, method for processing information, and program

    JP2021085785A

Cited By

  • Information processing device, information processing method, and program

    JP7855772B1