Information processing device, information processing system, information processing method, and program

By integrating multiple models and performing a learning process, the information processing device enhances model accuracy in distributed data sets, even when correct labels are scarce, effectively addressing the challenges of federated learning.

WO2025104771A1PCT designated stage expired Publication Date: 2025-05-22NEC CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/040702
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

In machine learning for distributed data sets, it is often challenging to obtain correct labels for each client device, which hinders the improvement of model accuracy.

Method used

An information processing device and method that integrate multiple models from a model group, including models obtained directly or indirectly from a server device, and perform a learning process using the integrated model to enhance model accuracy.

Benefits of technology

The proposed solution effectively improves the accuracy of the model even when correct labels are difficult to obtain, by integrating multiple models and performing a learning process, thereby addressing the limitations of existing technologies in federated learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023040702_22052025_PF_FP_ABST
    Figure JP2023040702_22052025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises: an integration unit that integrates a plurality of models included in a model group that includes one or more models obtained by directly or indirectly using a model provided by a server device; and a learning unit that executes machine learning processing using the model integrated by the integration unit.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing system, information processing method, and program

[0001] The present disclosure relates to an information processing device, an information processing system, an information processing method, and a program.

[0002] A known machine learning method for distributed data sets is called federated learning, which involves distributing a model aggregated on a server device to each client device, updating the model on each client device, transmitting the updated model to the server device, and then re-aggregating the models on the server device, repeating this process (see, for example, Patent Literature 1).

[0003] Japanese Patent Application Publication No. 2022-169470

[0004] Generally, in machine learning for a distributed dataset, it may be difficult for each client device to obtain a correct label. The technology described in Patent Literature 1 has a problem in that it is difficult to improve the accuracy of the model in each client device under such circumstances. The present disclosure has been made in consideration of the above problem, and an exemplary purpose thereof is to provide a technology that can suitably improve the accuracy of a model even when it is difficult for each client device to obtain a correct label.

[0005] An information processing device according to an exemplary aspect of the present disclosure includes an integration means for integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device, and a learning means for executing a learning process using the model integrated by the integration means.

[0006] An information processing device according to an exemplary aspect of the present disclosure includes an acquisition means for acquiring input data, and an inference means for performing an inference process on the input data using a trained model obtained by a learning process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device.

[0007] An information processing system according to an exemplary aspect of the present disclosure is an information processing system including a server device and a plurality of client devices, wherein the server device includes an acquisition means for acquiring a model from each of the plurality of client devices, an aggregation means for aggregating the models from each of the client devices, and a provision means for providing the model aggregated by the aggregation means to each of the plurality of client devices, and each of the plurality of client devices includes an integration means for integrating a plurality of models included in a model group including one or more models obtained directly or indirectly using the model provided from the server device, and a learning means for executing a learning process using the model integrated by the integration means.

[0008] An information processing method according to an exemplary aspect of the present disclosure includes integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device, and performing a learning process using the integrated model.

[0009] An information processing method according to an exemplary aspect of the present disclosure includes acquiring input data and performing an inference process on the input data using a trained model obtained by a training process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device.

[0010] In addition, the information processing device according to each aspect may be realized by a computer. In this case, a program for realizing the information processing device on a computer by causing the computer to operate as each means provided in the information processing device, and a computer-readable recording medium on which the program is recorded, also fall within the scope of the present invention.

[0011] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that the accuracy of the model can be suitably improved even when it is difficult to obtain a correct label in each client device.

[0012] FIG. 1 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 2 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 3 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 4 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 5 is a block diagram showing a configuration of an information processing system according to the present disclosure. FIG. 6 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 7 is a block diagram showing a configuration of an information processing system according to the present disclosure. FIG. 8 is a diagram for explaining processing by an information processing system according to the present disclosure. FIG. 9 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 10 is a diagram for explaining processing by an information processing device according to the present disclosure. FIG. 11 is a block diagram showing the configuration of a computer functioning as an information processing device according to the present disclosure.

[0013] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.

[0014] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0015] (Overview of Information Processing Device 1) First, an overview of the information processing device 1 according to this exemplary embodiment will be described. As an example, the information processing device 1 according to this exemplary embodiment updates a model through a learning process based on a model provided from a server device, and provides the updated model to the server device. Here, obtaining the model from the server device, updating the model, and providing the updated model may or may not be repeated. Furthermore, a system including the server device and the information processing device 1 can be considered a system that performs so-called federated learning, but this term does not limit this exemplary embodiment. Furthermore, the information processing device 1 may also be referred to as a client device, but this term does not limit this exemplary embodiment.

[0016] As will be described below, the information processing device 1 according to this exemplary embodiment can, for example, suitably improve the accuracy of a model by performing the following processes: selecting and integrating multiple models from a model group consisting of one or more models obtained by directly or indirectly using a model provided by a server device; and executing a learning process based on the integrated model.

[0017] Furthermore, the information processing device 1 that executes the above-described processing has the effect of being able to suitably improve the accuracy of the model even when it is difficult to obtain a correct label.

[0018] Note that this exemplary embodiment can be applied to any model that is configured to be updatable (trainable), and specific examples of models include models based on machine learning algorithms such as CNN (Convolutional Neural Network) and RNN (Recurrent Neural Network), and combinations of these.

[0019] (Configuration of information processing device 1) The configuration of the information processing device 1 according to this exemplary embodiment will be described below with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing device 1 according to this exemplary embodiment. As shown in Fig. 1, the information processing device 1 includes an integration unit 11 and a learning unit 12.

[0020] (Integration Unit 11) The integration unit 11 integrates multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device. Here, a model obtained directly using a model provided from a server device (also referred to as a source model or a zeroth-generation model) refers, for example, to a model (also referred to as a first-generation model) obtained by training the source model through a training process using training data (sometimes simply referred to as input data). For example, the first-generation model is obtained by applying a series of training processes (parameter update processes within the model) consisting of a predetermined number of epochs (number of steps) to the source model.

[0021] Furthermore, a model obtained indirectly using a model (source model) provided by a server device refers, for example, to a model (second-generation model) obtained by training the above-mentioned first-generation model through a learning process using training data, or an n-th generation model (where n is an integer greater than or equal to 2) obtained by repeating similar steps multiple times.

[0022] Here, in this exemplary embodiment, the nth generation model may be obtained by applying a learning process to a model obtained by combining multiple models from the (n-1)th generation or earlier. In other words, in this exemplary embodiment, the nth generation model may be generated by the following process: selecting multiple models from a model group including multiple models from the (n-1)th generation or earlier; generating an integrated model by combining (integrating) the selected multiple models; and applying a learning process using training data (input data) to the generated integrated model to generate a trained model (nth generation model). Furthermore, the above-mentioned learning process may be performed by the learning unit 12, which will be described later.

[0023] In addition, integrating multiple models means, for example, integrating the parameters that define the m-th generation (m is an integer equal to or greater than 0) model (also called parameters included in the model) into a m When written as, for example, w' n = Σα m * w m The parameter w′ that defines the n-th generation model (where n is an integer equal to or greater than 1) is n Here, Σ indicates taking the sum from m=0 to m=n-1, the symbol "*" indicates multiplication, and α m is each w m represents the weighting coefficient multiplied by Σα m = 1. α m The specific value can be determined in advance as a suitable value depending on the given data set and the model configuration.

[0024] In the above notation, a primed w, such as w', indicates a parameter that defines the integrated and unlearned model, and an unprimed w indicates a parameter that defines the integrated and learned model, but this notation does not limit the present exemplary embodiment. In general, a model according to the present exemplary embodiment may have multiple parameters (a large number of parameters), and these parameters are collectively referred to as w mIt is expressed as follows:

[0025] In the above integration process, α for each m m The specific value can be set appropriately. For example, when integrating the most recent two generation models, the following value is set: n = α * w n-1 + (1-α) * w n-2 The parameter w' that defines the n-th generation model is n This process can also be expressed as a process in which the integration unit 11 performs integration processing by taking a weighted sum of a plurality of parameters that define the model obtained in the (n-2)th learning step and a plurality of parameters that define the model obtained in the (n-1)th learning step, and derives a post-integration model (n-th generation model).

[0026] In the above description, the integration process is performed by linearly combining parameters, but this does not limit the present exemplary embodiment. In general, the integration unit 11 calculates w' by using a function f with multiple variables as arguments. n = f (w m , w mー1 , ..., w 0 ) parameter w' n Here, the function f may be a nonlinear function.

[0027] In the above example of linear combination, the parameters of the model may be classified into a plurality of categories, and the weighting coefficients may be determined depending on the categories. In other words, the integrating unit 11 may set weights in the weighted sum depending on the category to which each of the plurality of parameters belongs. For example, the integrating unit 11 may set the weights in the weighted sum depending on the category to which each of the plurality of parameters belongs. n,c = Σα m,c * w m,c By the parameter w' n Here, w m,c , and w' n,c indicates a parameter belonging to a certain category c, and α m,c is the parameter w m,cIn addition, Σ represents the sum from m=0 to m=n−1, for example.

[0028] (Learning unit 12) The learning unit 12 executes a learning process using the model (the integrated model) integrated by the integration unit 11. More specifically, the learning unit 12 applies a learning process using learning data to the model integrated by the integration unit 11.

[0029] As an example, the learning unit 12 refers to the learning data and learns the above-described n-th generation model (pre-learning model), thereby obtaining the pre-learning parameter w′ n Update the learned parameters (parameters that define the learned model) w n get.

[0030] The learning unit 12 may be configured to perform the learning process by referring to learning data (input data) and labels (which may represent pseudo-labels) generated by referring to the learning data. However, this example does not limit the present exemplary embodiment. Also, the input data and the labels associated with the input data may be collectively referred to as learning data.

[0031] (Effects of information processing device 1) As described above, the information processing device 1 according to this exemplary embodiment is configured to: integrate multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device; and execute a learning process using the integrated model.

[0032] In this way, the information processing device 1 according to this exemplary embodiment integrates multiple models and executes a learning process using the integrated model, thereby suitably improving the accuracy of the model. Furthermore, the information processing device 1 according to this exemplary embodiment can be applied even when it is difficult to obtain a correct label. Furthermore, the information processing device 1 according to this exemplary embodiment can also be suitably applied to a mode in which a model aggregated in a server device is distributed to each client device, and the model is updated in each client device (sometimes called federated learning).

[0033] Therefore, the information processing device 1 according to this exemplary embodiment can, as an example, suitably improve the accuracy of the model even when it is difficult to obtain the correct label in each client device in federated learning.

[0034] (Flow of Information Processing Method S1) Next, the flow of the information processing method S1 according to this exemplary embodiment will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of the information processing method S1. As shown in Fig. 2, the information processing method S1 includes a process (step, process) S11 of integrating models and a process (step, process) S12 of executing a learning process using the integrated model.

[0035] (Step S11) In step S11, the integration unit 11 integrates multiple models included in a model group including one or more models obtained directly or indirectly using a model provided by the server device. The specific processing by the integration unit 11 has been described above, and therefore will not be described here.

[0036] (Step S12) In step S12, the learning unit 12 executes a learning process using the model (integrated model) integrated by the integration unit 11. The specific process by the learning unit 12 has been described above, and therefore will not be described here.

[0037] (Effects of Information Processing Method S1) As described above, the information processing method S1 according to this exemplary embodiment performs the following processes: integrates multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device, and performs a learning process using the integrated model.

[0038] According to the information processing method S1 including the above-described processes, the same effects as those of the information processing device 1 according to this exemplary embodiment can be achieved.

[0039] (Configuration of information processing device 2) Next, the configuration of the information processing device 2 according to this exemplary embodiment will be described with reference to Fig. 1. Fig. 3 is a block diagram showing the configuration of the information processing device 2 according to this exemplary embodiment. As shown in Fig. 3, the information processing device 2 includes an acquisition unit 21 and an inference unit 22.

[0040] (Acquisition unit 21) The acquisition unit 21 acquires input data. Here, the input data may be, for example, image data, text data, or other data, or a combination thereof. The input data acquired by the acquisition unit 21 is, for example, treated as a target for inference processing by the information processing device 2 in the inference phase.

[0041] (Inference unit 22) The inference unit 22 performs inference processing on the input data using a trained model. Here, the trained model is, for example, a trained model obtained by a learning process using a model obtained by integrating multiple models included in a model group including one or more models obtained by directly or indirectly using a model provided by the server device. In other words, the trained model is a model obtained by the following processes: selecting multiple models from a model group including one or more models obtained by directly or indirectly using a model provided by the server device, integrating the selected multiple models, and applying a learning process to the integrated model.

[0042] As an example, such a trained model can be obtained as a model generated by the information processing device 1 according to this exemplary embodiment (more specifically, a trained model generated by the learning unit 12).

[0043] (Effects of information processing device 2) As described above, the information processing device 2 according to this exemplary embodiment is configured to: acquire input data; and perform inference processing on the input data using a trained model obtained by a training process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device.

[0044] In this way, the information processing device 2 according to this exemplary embodiment performs an inference process using a trained model obtained by a training process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided by a server device, thereby enabling highly accurate inference processing. Furthermore, the trained model used by the information processing device 2 according to this exemplary embodiment can be generated even when it is difficult to obtain a correct label. Furthermore, the information processing device 2 according to this exemplary embodiment can also be suitably applied to a mode (sometimes called federated learning) in which models aggregated in a server device are distributed to each client device and the model is updated on each client device.

[0045] Therefore, the information processing device 2 according to this exemplary embodiment can perform highly accurate inference processing, even when it is difficult to obtain the correct label in each client device in federated learning, for example.

[0046] (Flow of Information Processing Method S2) Next, the flow of the information processing method S2 according to this exemplary embodiment will be described with reference to Fig. 4. Fig. 4 is a flow diagram showing the flow of the information processing method S2. As shown in Fig. 4, the information processing method S2 includes a process (step, process) S21 of acquiring input data and a process (step) S22 of executing an inference process.

[0047] (Step S21) In step S21, the acquisition unit 21 acquires input data. Specific examples of the input data have been described above, so a description thereof will be omitted here.

[0048] (Step S22) In step S22, the inference unit 22 executes an inference process on the input data using a trained model. Here, the trained model is, for example, a trained model obtained by a learning process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided by a server device. The process by the inference unit 22 has been described above, and therefore will not be described here.

[0049] (Effects of Information Processing Method S2) As described above, the information processing method S2 according to this exemplary embodiment performs the following processes: acquire input data; and perform inference processing on the input data using a trained model obtained by a training process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from the server device.

[0050] According to the information processing method S2 including the above-described processes, the same effects as those of the information processing device 2 according to this exemplary embodiment can be achieved.

[0051] (Configuration of Information Processing System) Next, the configuration of an information processing system according to this exemplary embodiment will be described with reference to FIG. 5. FIG. 5 is a block diagram showing the configuration of an information processing system according to this exemplary embodiment. As shown in FIG. 5, the information processing system according to this exemplary embodiment includes a server device 3 and multiple client devices 1-1, 1-2, .... As an example, the information processing system according to this exemplary embodiment can be suitably applied to a system that performs federated learning. However, this term does not limit this exemplary embodiment.

[0052] (Server Device 3) As shown in FIG. 5, the server device 3 includes an acquisition unit 31, an aggregation unit 32, and a provision unit 33.

[0053] (Acquisition Unit 31) The acquisition unit 31 acquires a model from each of the multiple client devices 1-1, 1-2, .... Here, the model acquired from each client device may, for example, be a trained model obtained by applying a training process to a source model provided in advance from the server device 3. Note that in this exemplary embodiment, "acquiring a model" includes, for example, acquiring one or more parameters included in the model (that define the model), or acquiring information about one or more parameters. For example, the acquisition unit 31 may be configured to acquire the values ​​of these parameters themselves, or may be configured to acquire the amount of change in these parameters (for example, the difference from the previous step when performing repeated processing).

[0054] (Aggregation unit 32) The aggregation unit 32 aggregates models from each of the client devices 1-1, 1-2, .... Details of the aggregation process by the aggregation unit 32 do not limit this exemplary embodiment, but as an example, parameters that define the aggregated model may be derived by taking a weighted average of parameters acquired from each of the client devices 1-1, 1-2, .... The aggregation process performed by the aggregation unit 32 can also be expressed as an "integration process." However, the integration process performed by the aggregation unit 32 may differ from the integration process performed by the integration unit 11 described above.

[0055] (Providing Unit 33) The providing unit 33 provides the model aggregated by the aggregating unit 32 to each of the multiple client devices 1-1, 1-2, .... Here, in this exemplary embodiment, "providing a model" includes, for example, providing one or more parameters included in the model (that define the model), or providing information about one or more parameters. For example, the providing unit 33 may be configured to provide the values ​​of these parameters themselves, or may be configured to provide the amount of change in these parameters (for example, the difference from the previous step when performing repeated processing).

[0056] (Client devices 1-1, 1-2, ...) Each of the client devices 1-1, 1-2, ... has, as an example, the same configuration as the information processing device 1 described in this exemplary embodiment. Furthermore, each of the client devices 1-1, 1-2, ... may further have the same configuration as the information processing device 2 described in this exemplary embodiment.

[0057] 5, the client device 1-1 includes an integration unit 11-1 and a learning unit 12-1. Here, the integration unit 11-1 and the learning unit 12-1 are included in the information processing device 1. Since they have the same configuration as the integration unit 11 and the learning unit 12, their explanation will be omitted here.

[0058] 5, the client device 1-2 includes an integration unit 11-2 and a learning unit 12-2. Here, the integration unit 11-2 and the learning unit 12-2 are included in the information processing device 1. Since they have the same configuration as the integration unit 11 and the learning unit 12, their explanation will be omitted here.

[0059] (Processing Flow in Information Processing System) Next, a processing flow by the information processing system according to this exemplary embodiment will be described with reference to Fig. 6. Fig. 6 is a flow diagram showing a processing flow by the information processing system according to this exemplary embodiment.

[0060] (Step S33-0) In step S33-0, the providing unit 33 included in the server device 3 provides a model to each of the client devices 1-1 and 1-2. The model provided in this step is also called a source model. The source model may or may not be a model obtained by aggregating (integrating) models that have undergone learning processing (update processing) in each client device.

[0061] (Step S11-1-0, Step S11-2-0) In step S11-1-0, the integration unit 11-1 included in the client device 1-1 executes integration processing. The integration processing by the integration unit 11-1 is similar to the integration processing by the integration unit 11 included in the information processing device 1, so a description thereof will be omitted here.

[0062] Similarly, in step S11-2-0, the integration unit 11-2 of the client device 1-2 executes integration processing. The integration processing by the integration unit 11-2 is similar to the integration processing by the integration unit 11 of the information processing device 1, and therefore a description thereof will be omitted here.

[0063] (Steps S11-1-0, S11-2-0) Next, in step S12-1-0, the learning unit 12-1 included in the client device 1-1 executes a learning process. The learning process by the learning unit 12-1 is similar to the learning process by the learning unit 12 included in the information processing device 1, so a description thereof will be omitted here. The client device 1-1 provides the model obtained by the learning process (the learned model) to the server device 3.

[0064] Similarly, in step S12-2-0, the learning unit 12-2 included in the client device 1-2 executes a learning process. The learning process by the learning unit 12-2 is similar to the learning process by the learning unit 12 included in the information processing device 1, and therefore a description thereof will be omitted here. The client device 1-2 provides the model obtained by the learning process (the learned model) to the server device 3.

[0065] (Steps S31-1, S32-1, S33-1) Subsequently, in step S31-1, the acquisition unit 31 included in the server device 3 acquires models from each of the client devices 1-1 and 1-2. Then, in step S32-1, the aggregation unit 32 included in the server device 3 aggregates the models from each of the client devices 1-1 and 1-2. In step S33-1, the provision unit 33 included in the server device 3 provides the models aggregated by the aggregation unit 32 to each of the client devices 1-1 and 1-2. Thereafter, as shown in FIG. 6 , integration processing and learning processing are performed in each of the client devices 1-1 and 1-2, and the learned models are again provided to the server device 3.

[0066] (Effects of the Information Processing System) As described above, the information processing system according to this exemplary embodiment is an information processing system including a server device and a plurality of client devices, and is configured as follows: the server device acquires a model from each of the plurality of client devices, aggregates the models from each client device, and provides the model aggregated by the aggregation means to each of the plurality of client devices; and each of the plurality of client devices integrates a plurality of models included in a model group including one or more models obtained directly or indirectly using the model provided from the server device, and executes a learning process using the model integrated by the integration means.

[0067] According to the information processing system configured as above, the same effects as those of the information processing device 1 according to this exemplary embodiment can be achieved.

[0068] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0069] (Configuration of Information Processing System 100A) The configuration of the information processing system 100A according to this exemplary embodiment will be described with reference to FIG. 7. FIG. 7 is a block diagram showing the configuration of the information processing system 100A. As shown in FIG. 7, the information processing system 100A includes a server device 3A and a client device (information processing device) 1A. Also, as shown in FIG. 7, the server device 3A and the client device 1A are communicatively connected via a network N. Here, the specific configuration of the network N does not limit this exemplary embodiment, but as an example, a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination of these networks can be used.

[0070] 7 shows only one client device 1A for presentation purposes, but this does not limit the present exemplary embodiment, and the information processing system 100A may include multiple client devices (information processing devices) having similar functions to the client device 1A. These client devices may also be referred to as individual client devices 1A. Furthermore, when each client device is explicitly described, they may be referred to using a subnumber, such as client device 1A-1, client device 1A-2, .... Alternatively, they may be referred to as client device 1A, client device 2A, ....

[0071] 7, the server device 3A includes a control unit 30A, a storage unit 37A, a communication unit 38A, and an input / output unit 39A. The control unit 30A controls each unit included in the server device 3A.

[0072] The communication unit 38A communicates with devices external to the server device 3A. As an example, the communication unit 38A communicates with the client device 1A. The communication unit 38A transmits data supplied from the control unit 30A to the client device 1A, and supplies data received from the client device 1A to the control unit 30A. Note that the data provided by the communication unit 38A to the client device 1A includes, for example, an initial model generated by the generation unit 34 described below, or an aggregated model generated by the aggregation unit 32. Note that the model provided from the server device 3A to each client device 1A may also be referred to as a source model.

[0073] Furthermore, in this exemplary embodiment, the term “model” may also include the meaning of “one or more parameters that define the model.” Therefore, the data provided by the communication unit 38A to the client device 1A may be expressed as including, for example, one or more parameters that define the initial model generated by the generation unit 34 described below, or one or more parameters that define the aggregated model generated by the aggregation unit 32.

[0074] Furthermore, the data that the communication unit 38A receives from the client device 1A includes, for example, a model updated by each client device 1A. In other words, the data that the communication unit 38A receives from the client device 1A includes, for example, one or more parameters that define the model updated by each client device 1A.

[0075] The input / output unit 39A is configured to include at least one of an input / output device such as a keyboard, a mouse, a display, a printer, or a touch panel. Alternatively, the input / output unit 39A may be configured to have an input / output device such as a keyboard, a mouse, a display, a printer, or a touch panel connected to it. In this configuration, the input / output unit 39A accepts various types of information input to the server device 3A from the connected input device. Furthermore, under the control of the control unit 30A, the input / output unit 39A outputs various types of information to the connected output device. An example of the input / output unit 39A is an interface such as a USB (Universal Serial Bus).

[0076] (Control Unit 30A) As shown in FIG. 7, the control unit 30A of the server device 3A includes an acquisition unit 31, an aggregation unit 32, a provision unit 33, and a generation unit .

[0077] (Acquisition Unit 31) The acquisition unit 31 acquires a model from each of the multiple client devices 1-1, 1-2, .... The model acquired by the acquisition unit 31 is stored in the storage unit 37A, for example. Here, as described in the exemplary embodiment 1, the model acquired from each client device may be, for example, a trained model obtained by applying a training process to a source model provided in advance from the server device 3A. Note that in this exemplary embodiment, "acquiring a model" also includes, for example, acquiring one or more parameters included in the model (that define the model), or acquiring information about one or more parameters. For example, the acquisition unit 31 may be configured to acquire the values ​​of these parameters themselves, or may be configured to acquire the amount of change in these parameters (e.g., the difference from the previous step in the case of repeated processing).

[0078] The acquiring unit 31 may also be configured to acquire pre-learning data that the generating unit 34 (described later) refers to in order to generate a pre-learning model. Here, the pre-learning data may include, for example, target data to be input into a model to be pre-learned and a correct label associated with the target data.

[0079] (Aggregation unit 32) The aggregation unit 32 aggregates models from each client device 1A. The aggregated model generated by the aggregation process by the aggregation unit 32 is stored in the storage unit 37A, for example. Details of the aggregation process by the aggregation unit 32 do not limit the present exemplary embodiment. For example, parameters defining the aggregated model may be derived by taking a weighted average of parameters acquired from each client device 1A. Furthermore, the aggregation unit 32 may be configured to generate the aggregated model by further referring to the feature information acquired from each client device 1A.

[0080] The aggregating unit 32 may also be configured to cluster the models from each client device 1A and aggregate the models for each cluster obtained by the clustering process.

[0081] As an example, the acquisition unit 31 may be configured to acquire, together with the model from each client device 1A, information (type information) indicating at least one of: the type of the client device; the type of the model; and the type of data to which the model (client device) is applied, and the aggregation unit 32 may be configured to cluster the models from each client device 1A according to the type information.

[0082] For example, if the information processing system 100A includes four client devices 1A-1, 1A-2, 1A-3, and 1A-4, and the client devices 1A-1 and 1A-2 are applied to image data captured from a car traveling on a road, and the client devices 1A-3 and 1A-4 are applied to image data captured from a transport vehicle traveling within a factory, the aggregation unit 32 may perform clustering in such a manner that the model from the client device 1A-1 (model 1A-1) and the model from the client device 1A-2 (model 1A-2) are classified into cluster A, and the model from the client device 1A-3 (model 1A-3) and the model from the client device 1A-4 (model 1A-4) are classified into cluster B. The aggregation unit 32 may be configured to generate an aggregated model for cluster A by aggregating the models 1A-1 and 1A-2 belonging to cluster A, and to generate an aggregated model for cluster B by aggregating the models 1A-3 and 1A-4 belonging to cluster B.

[0083] The aggregation process performed by the aggregation unit 32 can also be expressed as an “integration process.” However, the integration process performed by the aggregation unit 32 may differ from the integration process performed by the integration unit 11, which will be described separately.

[0084] (Providing Unit 33) The providing unit 33 provides the model aggregated by the aggregating unit 32 to each client device 1A. Here, in the present exemplary embodiment, "providing a model" includes, for example, providing one or more parameters included in the model (that define the model) or providing information on one or more parameters. For example, the providing unit 33 may be configured to provide the values ​​of these parameters themselves, or may be configured to provide the amount of change in these parameters (for example, the difference from the previous step when performing repeated processing).

[0085] (Generation unit 34) The generation unit 34 generates a pre-trained model to be provided to each client device 1A. As an example, the generation unit 34 generates the pre-trained model through a learning process that references pre-training data. Here, as an example, the pre-training data may be acquired by the acquisition unit 31. The pre-training data may also include target data and a correct label associated with the target data. The generation unit 34 can train the model by inputting the target data into the pre-trained model and updating the model so that the difference between the output of the model and the correct label is reduced. Note that the type of pre-trained model is not limited to this exemplary embodiment, and examples include a convolutional neural network (CNN), a recurrent neural network (RNN), and a combination thereof.

[0086] (Client Device (Information Processing Device) 1A) As shown in FIG. 7, the client device (information processing device) 1A includes a control unit 10A, a storage unit 17A, a communication unit 18A, and an input / output unit 19A.

[0087] The communication unit 18A communicates with devices external to the client device 1A. As an example, the communication unit 18A communicates with the server device 3A. The communication unit 18A transmits data supplied from the control unit 10A to the server device 3A, and supplies data received from the server device 3A to the control unit 10A. As an example, the data received by the communication unit 18A from the server device 3A may include at least one of the pre-trained model generated by the generation unit 34 and the aggregated model generated by the aggregation unit 32.

[0088] In addition, the data provided by the communication unit 18A to the server device 3A may include, as an example, a trained model obtained by applying an integration process by the integration unit 11 and a learning process by the learning unit 12 to the above model from the server device 3A.

[0089] The input / output unit 19A is configured to include at least one of an input / output device such as a keyboard, a mouse, a display, a printer, a touch panel, etc. Alternatively, the input / output unit 19A may be configured to have an input / output device such as a keyboard, a mouse, a display, a printer, a touch panel, etc. connected to it. In this configuration, the input / output unit 19A accepts various types of information input to the client device 1A from the connected input device. Furthermore, under the control of the control unit 10A, the input / output unit 19A outputs various types of information to the connected output device. An example of the input / output unit 19A is an interface such as a USB (Universal Serial Bus).

[0090] (Storage unit 17A) The storage unit 17A stores various types of data referenced by the control unit 10A and various types of data generated by the control unit 10A. As an example, the storage unit 17A stores each model generated or referenced in each step (each generation) of the iterative processing in the integration unit 11.

[0091] As an example, the memory unit 17A stores model n-1, which is a model of the n-1th generation (where n is a natural number), model n, which is a model of the nth generation, and model n+1, which is a model of the n+1th generation.

[0092] (Control Unit 10A) As shown in FIG. 7, the control unit 10A includes an integration unit 11, a learning unit 12, an acquisition unit 13, a pseudo label generation unit 14, a presentation unit 15, a repetition control unit 16, a provision unit 17, and an inference unit 22.

[0093] (Integration Unit 11) As in the exemplary embodiment 1, the integration unit 11 integrates multiple models included in a model group including one or more models obtained directly or indirectly using a model provided by the server device 3A. Here, a model obtained directly using a model provided by the server device (also referred to as a source model or a zeroth-generation model) refers, for example, to a model (also referred to as a first-generation model) obtained by training the source model through a training process using training data (sometimes simply referred to as input data), as described in the exemplary embodiment 1. For example, the first-generation model is obtained by applying a series of training processes (parameter update processes within the model) consisting of a predetermined number of epochs (number of steps) to the source model.

[0094] Furthermore, the model obtained indirectly using the model (source model) provided by the server device 3A refers to, as an example, a model (second-generation model) obtained by training the above-mentioned first-generation model through a learning process using training data, as described in exemplary embodiment 1, or an n-th generation model (where n is an integer greater than or equal to 2) obtained by repeating similar steps multiple times.

[0095] Here, in this exemplary embodiment, the nth generation model may be obtained by applying a learning process to a model obtained by combining multiple models from the (n-1)th generation or earlier. In other words, in this exemplary embodiment, the nth generation model may be generated by the following process: selecting multiple models from a model group including multiple models from the (n-1)th generation or earlier; generating an integrated model by combining (integrating) the selected multiple models; and applying a learning process using training data (input data) to the generated integrated model to generate a trained model (nth generation model). Furthermore, the above-mentioned learning process may be performed by the learning unit 12.

[0096] In addition, integrating multiple models means, for example, integrating the parameters that define the m-th generation (m is an integer equal to or greater than 0) model (also called parameters included in the model) into a m When written as, for example, w' n = Σα m * w m The parameter w′ that defines the n-th generation model (where n is an integer equal to or greater than 1) is n Here, Σ indicates taking the sum from m=0 to m=n-1, the symbol "*" indicates multiplication, and α m is each w m represents the weighting coefficient multiplied by Σα m = 1. α m The specific value can be determined in advance as a suitable value depending on the given data set and the model configuration.

[0097] In the above notation, a primed w, such as w', indicates a parameter that defines the integrated and unlearned model, and an unprimed w indicates a parameter that defines the integrated and learned model, but this notation does not limit the present exemplary embodiment. In general, a model according to the present exemplary embodiment may have multiple parameters (a large number of parameters), and these parameters are collectively referred to as w m It is expressed as follows:

[0098] In the above integration process, α for each m m The specific value can be set appropriately. For example, when integrating the most recent two generation models, the following value is set: n = α * w n-1 + (1-α) * w n-2 The parameter w' that defines the n-th generation model is n This process can also be expressed as a process in which the integration unit 11 performs integration processing by taking a weighted sum of a plurality of parameters that define the model obtained in the (n-2)th learning step and a plurality of parameters that define the model obtained in the (n-1)th learning step, and derives a post-integration model (n-th generation model).

[0099] In the above description, the integration process is performed by linearly combining parameters, but this does not limit the present exemplary embodiment. As mentioned in the exemplary embodiment 1, the integration unit 11 calculates w' using a function f with multiple variables as arguments. n = f (w m , w mー1 , ..., w 0 ) parameter w' n Here, the function f may be a nonlinear function.

[0100] In the above example of linear combination, the parameters of the model may be classified into a plurality of categories, and the weighting coefficients may be determined depending on the categories. In other words, the integrating unit 11 may set weights in the weighted sum depending on the category to which each of the plurality of parameters belongs. For example, the integrating unit 11 may set the weights in the weighted sum depending on the category to which each of the plurality of parameters belongs. n,c = Σα m,c * w m,c By the parameter w' n Here, w m,c , and w' n,c indicates a parameter belonging to a certain category c, and α m,c is the parameter w m,cIn addition, Σ represents the sum from m=0 to m=n−1, for example.

[0101] For example, if the model is configured to include multiple layers, such as a neural network, the integrating unit 11 may classify each of the multiple layers into one of multiple categories and perform the integrating process using a weighting factor set for each category. As an example, if the model is configured to include a total of four layers, namely, layer 1 (input layer), layer 2 (intermediate layer), layer 3 (intermediate layer), and layer 4 (output layer), the integrating unit 11 may classify layer 1 and layer 2 into a first category and layer 3 and layer 4 into a second category. Then, the integrating unit 11 calculates w' n,1 = Σα m,1 * w m,1 w' n,2 = Σα m,2 * w m,2 By the parameter w' n,1 and w' n,2 Here, w m,1 , and w' n,1 indicates a parameter belonging to a certain first category, and w m,2 , and w' n,2 indicates a parameter belonging to a second category, and α m,1 , and α m,2 are weighting coefficients multiplied by the parameters belonging to the first and second categories, respectively, and are generally m,1 ≠α m,2 can be set to.

[0102] The integrating unit 11 may be configured to determine the weighting coefficient depending on the type of data set (input data) to be processed by the client device 1 A. In other words, the integrating unit 11 may be configured to set the weight in the weighted sum depending on the type of input data.

[0103] For example, the integrating unit 11 may be configured to set weights in the weighted sum depending on whether the input data has been captured from a car traveling on the ground or a drone flying in the sky. As an example, the integrating unit 11 may classify input data captured from a car traveling on the ground into type 1, and input data captured by a drone flying in the sky into type 2. Then, the integrating unit 11 may calculate w' n,1 = Σα m,1 * w m,1 w' n,2 = Σα m,2 * w m,2 By the parameter w' n,1 and w' n,2 Here, w m,1 , and w' n,1 indicates a parameter belonging to a certain type 1, and w m,2 , and w' n,2 indicates a parameter belonging to type 2, and α m,1 , and α m,2 are weighting coefficients multiplied by the parameters belonging to type 1 and type 2, respectively. Generally, α m,1 ≠α m,2 can be set to.

[0104] The integrating unit 11 may also be configured to set the weights in the weighted sum according to the learning status of the learning unit 12 included in the client device 1A. For example, the weights in the weighted sum may be set according to the learning accuracy of the learning process in the learning unit 12. The weights in the weighted sum may also be set according to the degree of increase in the learning accuracy of the learning process in the learning unit 12. If the degree of increase in the learning accuracy changes significantly, it is possible that overlearning has occurred, and therefore it is preferable to set the weights in the weighted sum according to the degree of increase in the learning accuracy, as described above. As an example, if the learning accuracy of the learning process performed by the learning unit 12 on an n-generation model obtained by integrating an n-1th generation model and an n-2th generation model is higher by a predetermined percentage or more than the learning accuracy of the learning process performed by the learning unit 12 on the n-1th generation model, there is a possibility that overlearning has occurred in the nth generation or the n-1th generation. In such a case, the integration unit 11 may generate the nth generation model again by setting a larger weight for the n-2th generation model, or may be configured to generate the n+1th generation model by an integration process that leaves the contribution of the n-2th generation model.

[0105] As described above, the integrating unit 11 integrates multiple models included in a model group including one or more models obtained by directly or indirectly using a model provided by the server device 3A, and applies a learning process by the learning unit 12 to the integrated model, thereby suitably suppressing overlearning of the model (also referred to as a local model) in the client device 1A. This allows the client device 1A to suitably suppress deterioration in the accuracy of the model (local model). Note that more specific processing by the integrating unit 11 in each step of the repetitive processing will be described later with reference to different drawings.

[0106] (Learning unit 12) The learning unit 12 executes a learning process using the model (the integrated model) integrated by the integration unit 11. More specifically, the learning unit 12 applies a learning process using learning data to the model integrated by the integration unit 11.

[0107] As an example, the learning unit 12 refers to the learning data and learns the above-described n-th generation model (pre-learning model), thereby obtaining the pre-learning parameter w′ n Update the learned parameters (parameters that define the learned model) w n get.

[0108] The learning unit 12 may be configured to perform the learning process by referring to learning data (input data) and labels (which may represent pseudo labels) generated by referring to the learning data. As an example, the learning unit 12 inputs learning data (input data) into the model (integrated model) integrated by the integration unit 11, and updates the integrated model so that the difference between the output of the model and the pseudo labels becomes smaller, thereby training the integrated model. As an example, the pseudo labels referred to by the learning unit 12 can be generated in advance by a pseudo label generation unit 14 described later.

[0109] Note that more specific processing by the integration unit 11 in each step of the repetitive processing will be described later with reference to different drawings.

[0110] (Acquisition unit 13) The acquisition unit 13 acquires input data to be processed by the client device 1A. Here, the input data acquired by the acquisition unit 13 is, for example, data that does not include a correct answer label. The input data includes data (learning data) that is referenced in the learning process executed by the learning unit 12 in the learning phase.

[0111] The acquisition unit 13 may also be configured to further acquire inference data to be referenced by the inference unit 22 (described later) in the inference phase. The inference data is also data that does not include a correct answer label. In this way, the data acquired by the acquisition unit 13 and referenced by the learning unit 12 or the inference unit 22 is, for example, data that does not include a correct answer label, and therefore may be simply referred to as input data without any particular distinction between the learning data and the inference data.

[0112] (Pseudo Label Generation Unit 14) The pseudo label generation unit 14 generates pseudo labels by referring to the input data acquired by the acquisition unit 13. As an example, the pseudo label generation unit 14 may be configured to generate the pseudo labels according to the results of a clustering process on the input data. More specific processing by the pseudo label generation unit 14 will be described later with reference to different drawings.

[0113] (Presentation Unit 15) The presentation unit 15 presents various types of information acquired or generated by the client device 1 A or the server device 3 A. As an example, the presentation unit 15 visually presents the information via a display provided in the input / output unit 19 A.

[0114] As an example, the presentation unit 15 may be configured to present information on at least any of the following: at least any of the models included in the group of models referenced by the integration unit 11 (in other words, a group of models including one or more models obtained by directly or indirectly using a model provided by the server device 3A); a model integrated by the integration unit 11 (the integrated model); and a trained model obtained by applying a learning process by the learning unit 12 to the integrated model; and information obtained using at least any of the above models, the integrated model, and the trained model, which assists the user in making decisions.

[0115] As an example, the presentation unit 15 may be configured to display the values ​​of one or more parameters that define the model after the learning, or may be configured to display the amount of change in these parameters (e.g., the difference from the previous step when performing repeated processing).

[0116] According to the above configuration, the user can easily check how the integration process by the integration unit 11 and the learning process by the learning unit 12 are being executed.

[0117] (Repetition control unit 16) The repetition control unit 16 controls the repetition of the integration process by the integration unit 11 and the learning process by the learning unit 12. As an example, the repetition control unit 16 refers to a predetermined condition, and repeats the integration process by the integration unit 11 and the learning process by the learning unit 12 until the predetermined condition is satisfied. More specific processing by the repetition control unit 16 will be described later with reference to different drawings.

[0118] (Providing unit 17) The providing unit 17 provides, to the server device 3A, information about the trained model obtained by the integration process by the integrating unit 11 and the learning process by the learning unit 12. As an example, the providing unit 17 provides, to the server device 3A, information about the trained model via the communication unit 18A.

[0119] More specifically, the providing unit 17 may be configured to provide the server device 3A with the values ​​of one or more parameters that define the trained model, or may be configured to provide the server device 3A with the amount of change in these parameters (for example, the difference from the previous step when performing repeated processing).

[0120] The providing unit 17 may also be configured to provide the server device 3A with feature information obtained by inputting the input data acquired by the acquiring unit 13 into the model.

[0121] (Inference unit 22) The inference unit 22 performs inference processing on the input data (data for inference) acquired by the acquisition unit 13 using the trained model obtained by the integration processing by the integration unit 11 and the learning processing by the learning unit 12.

[0122] Here, as described above, the trained model is a model obtained by: - ​​the integration unit 11 selecting multiple models from a model group including one or more models obtained directly or indirectly using a model provided by the server device 3A; - the integration unit 11 integrating the selected multiple models to generate an integrated model; and - applying a learning process by the learning unit 12 to the generated integrated model.

[0123] In addition, the inference unit 22 may be configured to perform inference processing on the input data using an aggregated model obtained by the server device 3A through aggregation processing that references the trained model.

[0124] According to the above configuration, for example, even if it is difficult to obtain a correct label in each client device, it is possible to execute highly accurate inference processing.

[0125] (Explanation of aspects as a federated learning system) Next, with reference to FIG. 8, an aspect of the information processing system 100A as a federated learning system will be described. FIG. 8 is a diagram for explaining an aspect of the information processing system 100A as a federated learning system. As described above, the information processing device 100A includes a server device 3A and one or more client devices (information processing device 1A). In the example shown in FIG. 8, each of the multiple client devices is shown as client device 1A and client device 2A. Here, client device 2A has a configuration similar to that of client device 1A.

[0126] In the example shown in FIG. 8 , domain 1, domain 2, and domain 3 exist as data domains. Here, domain 1, as an example, is composed of data including target data (also referred to as features in FIG. 8 ) and correct labels associated with the target data (simply referred to as labels in FIG. 8 ). Domain 1 is also referred to as a source domain or source domain data. Domain 1 (source domain data) is referenced by the server device 3A as shown in FIG. 8 . More specifically, the source domain data is referenced by the generation unit 34 as pre-training data and used to generate a pre-trained model. The generated pre-trained model is provided to each client device 1A, 2A as a source model.

[0127] On the other hand, as shown in Figure 8, domain 2 and domain 3 are both composed of data that do not include a correct label. In the example shown in Figure 8, client device 1A refers to data (input data) that does not include a correct label, updates the model (local model) targeted by client device 1A through the integration process and learning process described above, and provides the updated model to server device 3A. The same is true for client device 2A. Note that domain 2 and domain 3 may also be referred to as target domains, but this term does not limit this exemplary embodiment.

[0128] In this way, the information processing system 100A according to this exemplary embodiment has aspects as a system that performs so-called federated learning, in which: - the server device 3A provides a model to each of the client devices 1A, 2A; - the model provided from the server device 3A is updated in each of the client devices 1A, 2A, and the updated model is provided to the server device 3A; and - aggregating multiple models provided from each of the client devices 1A, 2A to generate an aggregated model, and providing the generated aggregated model to each of the client devices 1A, 2A.

[0129] As described above, the information processing system 100A according to this exemplary embodiment also has an aspect as a system that performs domain adaptation, in which: a source model generated by referring to a source domain is updated by applying it to a target domain; the updated model is aggregated in the server device 3A to generate an aggregated model, and the generated aggregated model is provided to each of the client devices 1A and 2A.

[0130] 8, the information processing system 100A according to this exemplary embodiment can be applied to a problem scenario in which none of the multiple client devices 1A, 2A can access source domain data. In other words, in the example shown in FIG. 8, the information processing system 100A can perform the above-described processing without accessing source domain data. Performing domain adaptation without accessing source domain data in this manner is sometimes referred to as source-free domain adaptation.

[0131] Therefore, the information processing system 100A according to this exemplary embodiment is a system that can be suitably applied to a federated learning setting for source-free domain adaptation.

[0132] (Application Example) As an application example of the information processing system 100A to a federated learning setting of source-free domain adaptation, for example, there is a case where several companies in different locations jointly create a model that supports autonomous driving based on images from an in-vehicle camera.

[0133] In such an application example, a source model trained using "in-vehicle camera images in urban areas" (domain 1) is stored in server device 3A, and the data used for training (source domain data) is not made public.

[0134] For example, several companies in different locations each own "in-vehicle camera images in coastal areas" (Domain 2) and "in-vehicle camera images in mountainous areas" (Domain 3). Although the domains (e.g., background information) of each image are different, the task they want to solve (the model they want to create) is common to all of them: "domain-independent autonomous driving assistance."

[0135] In addition, the images from the in-vehicle camera are not shared among the client devices, and the model is created using a federated learning setting. Furthermore, since the images from the in-vehicle cameras held by each company do not have labels indicating the situation of each image, each client device is assumed to hold only the features (image data, the input data described above).

[0136] In such a situation, the information processing device 1A according to this exemplary embodiment performs the following processes: - acquire a source model from the server device 3A; - use the acquired source model to update a local model by referring to "in-vehicle camera images of a coastal area" (domain 2) as input data; and - provide the updated model to the server device 3A. These processes may be repeated. Similarly, an information processing system 2A having a similar configuration to the information processing device 1A performs the following processes: - acquire a source model from the server device 3A; - use the acquired source model to update a local model by referring to "in-vehicle camera images of a mountainous area" (domain 3) as input data; and - provide the updated model to the server device 3A. These processes may be repeated.

[0137] In this way, the information processing system 100A according to this exemplary embodiment can preferably update the local model and the source model even when the source domain data is not publicly disclosed.

[0138] (Processing Flow in Information Processing System 100A) Next, the processing flow by the information processing system 100A according to this exemplary embodiment will be described with reference to Fig. 6. Fig. 6 is a flow diagram showing the processing flow S100A by the information processing system 100A according to this exemplary embodiment.

[0139] (Step S33) In step S33, the server device 3A distributes (provides) the model to each client device 1A. More specifically, the providing unit 33 of the server device 3A provides the model to each client device 1A via the communication unit 38A. The model provided to each client device 1A in this step is also called a source model.

[0140] (Step S14) Subsequently, in step S14, the pseudo label generating unit 14 of the client device 1A refers to the input data acquired by the acquiring unit 13 and generates a pseudo label.

[0141] (Step S161) Subsequently, in step S161, the repeat control unit 16 of the client device 1A determines whether or not this is the first learning since the source model was acquired. If this is the first learning since the source model was acquired (YES in step S161), the process proceeds to step S121; if not (NO in step S161), the process proceeds to step S11.

[0142] (Step S121) If the result of step S161 is YES, the learning unit 12 of the client device 1A executes a learning process based on the source model in step S121.

[0143] (Step S11) On the other hand, if the result of step S161 is NO, the integration unit 11 of the client device 1A executes a model integration process in step S11.

[0144] (Step S122) Subsequently, in step S122, the learning unit 12 of the client device 1A executes a learning process based on the integrated model integrated in step S11.

[0145] In step S162, the repeat control unit 16 of the client device 1A determines whether learning has been performed a specified number of times (also referred to as the first specified number of times). If learning has been performed a specified number of times (YES in step S162), the process proceeds to step S17. If not (NO in step S62), the process proceeds to step S123.

[0146] (Step S123) If the result of step S162 is NO, the control unit 10A of the client device 1A stores the model learned in step S122 in the storage unit 17A as one element of the model group.

[0147] (Step S17) On the other hand, if the result of step S162 is YES, the providing unit 17 of the client device 1A sends (provides) the trained model trained by the above-mentioned specified number of times of training processes to the server device 3A.

[0148] (Step S321) Then, in step S321, the aggregation unit 32 of the server device 3A clusters the multiple models acquired from each client device 1A.

[0149] (Step S322) Then, in step S322, the aggregation unit 32 of the server device 3A aggregates models based on the clustering result.

[0150] In step S323, the aggregation unit 32 of the server device 3A determines whether aggregation has been performed a specified number of times (also referred to as a second specified number of times). If aggregation has been performed a specified number of times (YES in step S323), the processing ends. If not (NO in step S323), the processing returns to step S33.

[0151] (Specific Example of Integration Processing and Learning Processing) Next, a specific example of the integration processing by the integration unit 11 of the client device 1A and the learning processing by the learning unit 12 will be described with reference to Fig. 10. Fig. 10 is a diagram for explaining a specific example of the integration processing and the learning processing.

[0152] 10 , the server device 3A distributes (provides) the source model to the client device 1A (corresponding to step S33 described above). Then, the pseudo label generation unit 14 of the client device 1A generates pseudo labels by referring to the input data acquired by the acquisition unit 13 (corresponding to step S14 described above). The generated pseudo labels are referenced in the learning process by the learning unit 12.

[0153] If this is the first learning after acquiring the source model (corresponding to YES in step S161), the learning unit 12 0 As an example, the learning unit 12 applies the learning process to the source model w 0 and input the source model w 0 The source model w is calculated so that the difference between the output of 0 By updating the updated model w 1 Generate.

[0154] In the above description, the source model is considered to be a model of the 0th generation (a model relating to the 0th step), and the source model (more specifically, the parameters that define the source model) is defined as w 0 and the model obtained by applying the learning process to the 0th generation model (more specifically, the parameters that define the model) is denoted as w 1 In the following explanation, the mth generation model will be referred to as w m It is sometimes written as:

[0155] On the other hand, if this is the second learning after acquiring the source model (corresponding to NO in step S161), the integration unit 11 uses the 0th generation model (source model) w 0 And the first generation model 1 By integrating these, the integrated model (more specifically, the parameters that define the integrated model) w' 2 As an example, as shown in FIG. 10, w' is generated using a weighting coefficient α (corresponding to step S121 described above). 2 = α * w 1 + (1-α) * w 0 Therefore, the integrated model w' 2 In the above process, the integrating unit 11 may be configured to classify the parameters of the model into a plurality of categories and determine the weighting coefficients depending on the categories. In other words, the integrating unit 11 may be configured to set weights in the weighted sum depending on the category to which each of the plurality of parameters belongs. For example, the integrating unit 11 may generate w' 2,c = α (c) * w 1 + (1-α (c) ) * w 0 As shown above, the weighting coefficient α set according to the category c is (c) Using the integrated model w' 2 A more specific example of the process of setting weighting factors according to categories has been described above, and therefore will not be described here.

[0156] Next, the learning unit 12 calculates the integrated model w′ generated as described above. 2 By applying the learning process to the second-generation model w (corresponding to step S122 described above), the second-generation model w 2 As an example, the learning unit 12 generates a model w′ after integrating the input data acquired by the acquisition unit 13. 2 and the integrated model w' 2 The integrated model w' is then calculated so that the difference between the output of the model w' and the pseudo-label generated for the input data is small. 2 By updating the above second generation model w 2 Generate.

[0157] As shown in FIG. 10, this process is repeated a predetermined number of times (corresponding to the processes from steps S161 to S162 described above).

[0158] In other words, the integration unit 11 integrates a plurality of parameters w that define the model obtained by the n-th learning step (the n-th generation model). n and a plurality of parameters w that define the model obtained by the (n-1)th learning step (the (n-1)th generation model). n-1 Then, the learning unit 12 applies a learning process to the integrated model to obtain the (n+1)th generation model w n+1 Generate.

[0159] In the example shown in Fig. 10, the specified number of times is expressed by a natural number N. The example shown in Fig. 10 includes: one learning process with reference to the 0th generation model; and N-1 integration and learning processes with reference to the (n-1)th generation model and the nth generation model, which corresponds to the above-mentioned predetermined condition referred to by the repetition control unit 16 indicating that the learning process, or the learning and integration process, is to be repeated a total of N times.

[0160] The Nth generation model obtained by repeating the learning process or the learning and integration process a total of N times (the model obtained in YES in step S162 described above) w Nis provided to the server device 3A by the providing unit 17 (corresponding to step S17 described above).

[0161] In this way, in the client device 1A, the integration unit 11 can be described as performing the following processing: - referring to a model group including one or more models obtained by each of one or more learning processes (learning steps); - integrating the model obtained by the nth learning step with one or more models obtained by learning steps prior to the nth; and the learning unit 12 - generating the (n+1)th model by executing a learning process using the models integrated by the integration unit 11.

[0162] According to the client device (information processing device) 1A having the integration unit 11 and learning unit 12 configured as described above, a model related to a past step and a model related to the immediately preceding step are integrated before performing the learning process, which makes it less likely that the problem of excessive adaptation to the input data acquired by the client device 1A (overlearning of the input data) will occur. Therefore, according to the client device (information processing device) 1A configured as described above, in a configuration in which input data is acquired for each client device 1A and a local model is updated using the input data, the accuracy of the local model can be improved. Furthermore, the accuracy of the aggregated model obtained by aggregating such multiple highly accurate local models in the server device 3A can be improved.

[0163] (Another Processing Example by the Iteration Control Unit 16) In the above example, a configuration has been described in which the iteration control unit 16 controls the integration process and the learning process so that they are repeated a predetermined number of times, but this does not limit the present exemplary embodiment. For example, the client device 1A may be configured to execute the integration process and the learning process with reference to a verification dataset, and the iteration control unit 16 may be configured to determine the value of N as the number of times until the processing on the verification dataset is completed (converged).

[0164] As an example, a validation dataset may be input to the client device 1A, and an output for the validation dataset may be output using the above-mentioned models of each generation. The number of iterations at which the error related to the output becomes equal to or less than a predetermined value (in other words, the number of iterations until the results converge) may be set as the value of N mentioned above.

[0165] Furthermore, the above-mentioned value of N may be set to a value that prevents overlearning of the local model for the input data in the client device 1A. Generally, when overlearning occurs in a client device, a phenomenon may occur in which the inference accuracy of other client devices different from the client device in question decreases. Therefore, by setting the value of N in the client device with reference to the inference accuracy of other client devices, the value of N can be set so as to prevent overlearning in the client device.

[0166] As mentioned above, the value of N may be set to a different value for each client device 1A. Alternatively, the value of N may be set according to the type of input data targeted by the client device 1A. For example, if the input data is captured by an automobile traveling on the ground, N may be set to 5, and if the input data is captured by a drone flying in the sky, N may be set to 10.

[0167] (Processing Example by Pseudo Label Generator 14) Next, a specific example of the pseudo label generation process by the pseudo label generator 14 of the client device (information processing device) 1A will be described with reference to Fig. 11. Fig. 11 is a diagram for explaining the pseudo label generation process by the pseudo label generator 14.

[0168] The upper part of Fig. 11 schematically illustrates the pseudo label generation process performed by the pseudo label generation unit 14. For example, the pseudo label generation unit 14 inputs input data acquired by the acquisition unit 13 into a source model acquired from the server device 3A or a model obtained by directly or indirectly referencing the source model (for example, a model obtained by applying at least one of the integration process performed by the integration unit 11 and the learning process performed by the learning unit 12 to the source model), and clusters the output data of the model in a predetermined space (hyperspace). The upper part of Fig. 11 illustrates cluster A and cluster B, where the data points of the output data indicated by diamonds belong to cluster A, and the data points of the output data indicated by triangles belong to cluster B.

[0169] The pseudo label generation unit 14 then calculates the value of the center of gravity of each cluster and determines a pseudo label using the calculated center of gravity value. As an example, the pseudo label generation unit 14 uses the value of the center of gravity of each cluster as a pseudo label. The upper part of FIG. 11 shows the center of gravity CA of cluster A and the center of gravity CB of cluster B. The pseudo label generation unit 14 assigns a label A similar to the center of gravity CA of cluster A to each data point belonging to cluster A and the input data from which the data points originate. Similarly, the pseudo label generation unit 14 assigns a label B similar to the center of gravity CB of cluster B to each data point belonging to cluster B and the input data from which the data points originate.

[0170] On the other hand, the lower part of Fig. 11 shows a pseudo label generation process according to a comparative example. In this pseudo label generation process, data points of the output data are divided by region division (boundary identification and boundary generation) in a predetermined space. As a result, classification errors occur in the data points marked "Error" in the lower part of Fig. 11.

[0171] As described above, the pseudo label generation unit 14 according to this exemplary embodiment can generate more appropriate pseudo labels because it generates pseudo labels according to the results of the clustering process on the input data acquired by the acquisition unit 13. In other words, the pseudo label generation unit 14 according to this exemplary embodiment generates highly accurate pseudo labels, and the learning unit 12 executes the learning process using the pseudo labels, thereby generating a highly accurate local model.

[0172] (Additional Notes Regarding the Learning Unit 12) As described above, the learning unit 12 inputs the input data acquired by the acquisition unit 13 to the integrated model, and updates (trains) the integrated model so that the difference between the output of the integrated model and the pseudo label becomes smaller. Here, the learning unit 12 may use cross-entropy loss as a loss function referenced in the update process (learning process), or may be configured to add an additional term to the loss function so as to increase the confidence level of the output. Using a loss function including such an additional term can achieve a more effective learning process.

[0173] As the additional term, for example, in a classification problem, an additional term may be added that makes the output result of the model approach a one-hot vector (an output result in which the confidence in a certain label is high and the confidence in others is low). For example, in a 3-classification problem, an additional term may be added to the loss function that indicates that the learning accuracy is higher when the model output result is [0.1, 0.8, 0.1] than when it is [0.2, 0.6, 0.2] (in other words, an additional term that makes the model output result approach the one-hot vector [0.0, 1.0, 0.0]).

[0174] (Additional Note 1 Regarding Client Device (Information Processing Device) 1A) In the above explanation, much has been said about the update processing (integration processing and learning processing) of the local model in the client device (information processing device) 1A. However, this exemplary embodiment also includes a device specialized for the inference phase as a client device (information processing device). Taking the configuration of the above-described client device 1A as an example, the control unit 10A may be configured to include only an acquisition unit 13 that acquires input data (data for inference) and an inference unit 22 that performs inference processing on the input data using a trained model. Here, the trained model may be configured to use a model of any generation generated by the above-described integration processing and learning processing, or may be configured to use an aggregated model provided by the server device 3A.

[0175] As described above, according to this exemplary embodiment, highly accurate local models and aggregated models can be generated through integration processing and learning processing, and inference processing can be performed using such highly accurate models.

[0176] (Additional Note 2 Regarding Client Device (Information Processing Device) 1A) The present exemplary embodiment also includes a configuration including a pseudo label generation unit 14 that performs a pseudo label generation process without performing the above-described integration process and learning process, or in addition to the above-described integration process and learning process. As an example, the client device (information processing device) 1A according to the present exemplary embodiment may include: an acquisition unit 13 that acquires input data; a pseudo label generation unit 14 that generates pseudo labels by referring to the input data; and a learning unit 12 that performs a learning process by using a model provided by the server device 3A and by referring to the input data and the pseudo labels, and the pseudo label generation unit 14 generates the pseudo labels according to a result of a clustering process on the input data.

[0177] According to the above configuration, the pseudo label generation unit 14 generates pseudo labels according to the results of the clustering process on the input data, thereby generating more appropriate pseudo labels. In other words, the pseudo label generation unit 14 according to this exemplary embodiment generates highly accurate pseudo labels, and the learning unit 12 performs learning processing using the pseudo labels, thereby generating highly accurate local models.

[0178] (Application Examples) Specific application examples of the information processing system 100A according to this exemplary embodiment will be described below. The information processing system 100A can be applied to a variety of industries, and several examples will be described below. However, these examples do not limit this exemplary embodiment, and the information processing system 100A can, of course, be applied to other industries. Furthermore, the information processing system 100A can also be applied across several industries.

[0179] (Example 1: Finance-related) The information processing system 100A according to this exemplary embodiment may be applied to the finance-related field, for example.

[0180] For example, a configuration may be adopted in which multiple client devices 1A, 2A, ... are each located at a respective branch of a bank, and in each client device 1A, the integration unit 11 and the learning unit 12 integrate and learn a model that predicts loan default risk based on the characteristics (loans, business status) of borrowers at that branch. In such a configuration, the data of each branch or each group corresponds to the above-mentioned source domain data. Then, as an example, the server device 3A generates a source model by referring to the source domain data and distributes it to each of the client devices 1A, 2A, ....

[0181] As another example, multiple client devices 1A, 2A, ... may be installed at each of multiple insurance company branches, and each client device 1A may integrate and learn a model that predicts insurance premiums for customers based on data such as the medical history, age, blood pressure, and genes of the customers at that branch using the integration unit 11 and learning unit 12. In such a configuration, the data of each branch or each insurance company corresponds to the source domain data described above. Then, as an example, the server device 3A generates a source model by referring to the source domain data and distributes it to each of the client devices 1A, 2A, ....

[0182] The results of the prediction of the default risk and the prediction of the insurance premium using the above model are examples of information that assists the user in making decisions according to this exemplary embodiment.

[0183] (Example 2: Medical and Healthcare Related) The information processing system 100A according to this exemplary embodiment may be applied to the medical field, for example.

[0184] For example, a configuration may be adopted in which multiple client devices 1A, 2A, ... are respectively located in multiple clinics, and each client device 1A integrates and learns a model that infers the cause of a disease and suggests a treatment method based on symptoms recorded in the medical records of patients at the clinic using the integration unit 11 and the learning unit 12. In such a configuration, the data of each clinic corresponds to the above-mentioned source domain data. Then, as an example, the server device 3A generates a source model by referring to the source domain data and distributes it to each client device 1A, 2A, ....

[0185] As another example, a configuration may be adopted in which multiple client devices 1A, 2A, ... are located at multiple pharmaceutical companies, respectively, and each client device 1A integrates and learns a model that predicts the activity of a compound based on the structure of the compound (ligand), the structure of a protein, etc., using the integration unit 11 and the learning unit 12. In such a configuration, data from each pharmaceutical company (e.g., data regarding activity against a compound library) corresponds to the above-mentioned source domain data. Then, as an example, the server device 3A generates a source model by referring to the source domain data and distributes it to each client device 1A, 2A, ....

[0186] The proposed treatment methods and predicted compound activity results from the above model are examples of information that assists a user in making decisions according to this exemplary embodiment.

[0187] (Example 3: Machine-related) The information processing system 100A according to this exemplary embodiment may be applied to a machine-related field, for example.

[0188] For example, a configuration may be adopted in which multiple client devices 1A, 2A, ... are respectively located in multiple factories, and in each client device 1A, the integration unit 11 and the learning unit 12 integrate and learn a model that controls the operation of a robot in the factory by referring to the situation (manufacturing situation, transportation situation) in the factory. In such a configuration, the data of each factory and each warehouse corresponds to the above-mentioned source domain data. Then, as an example, the server device 3A generates a source model by referring to the source domain data and distributes it to each of the client devices 1A, 2A, ....

[0189] As another example, a configuration may be adopted in which multiple client devices 1A, 2A, ... are respectively installed on multiple transportation vehicles (cars, airplanes, ships, etc.), and each client device 1A integrates and learns a model for controlling the transportation vehicle or traffic signals, etc., based on data such as the scenery from the transportation vehicle, the measurement status of the transportation vehicle, or the congestion level, using the integration unit 11 and the learning unit 12. In such a configuration, the data from each transportation vehicle, the data on traffic signals, etc., corresponds to the above-mentioned source domain data. Then, as an example, the server device 3A generates a source model by referring to the source domain data and distributes it to each client device 1A, 2A, ....

[0190] As another example, multiple client devices 1A, 2A, ... may be located at multiple logistics companies, and each client device 1A may integrate and learn a model that derives (changes to) a transportation route based on the status of the transported goods and the status of the transport equipment at the logistics company using the integration unit 11 and the learning unit 12. In this configuration, data related to the status of the transport equipment or the status of the transported goods at each logistics company (each site) corresponds to the above-mentioned source domain data. Then, as an example, the server device 3A generates a source model by referring to the source domain data and distributes it to each client device 1A, 2A, ....

[0191] (Example 4: Court-Related) The information processing system 100A according to this exemplary embodiment may be applied to the court-related field, for example.

[0192] For example, a configuration may be adopted in which multiple client devices 1A, 2A, ... are respectively located in multiple courts, and in each client device 1A, the integration unit 11 and the learning unit 12 integrate and learn a model that derives sentences and other information from data such as the circumstances of crimes handled by that court, the circumstances of evidence, laws, and court precedents. In such a configuration, the data that forms the basis of each trial in each court, or data such as recidivism rates, corresponds to the above-mentioned source domain data. Then, as an example, the server device 3A generates a source model by referencing the source domain data and distributes it to each of the client devices 1A, 2A, ...

[0193] The predicted results of sentencing and the like using the above model are an example of information that supports the user's decision-making according to this exemplary embodiment.

[0194] (Notes on each exemplary embodiment) The configurations described in each exemplary embodiment are not limited to the examples described above. The following configurations may be included to resolve some secondary issues that may arise when implementing federated learning in practice.

[0195] For example, when transmitting model parameters from each client device to a server device, it is preferable to have a configuration that ensures the confidentiality of the model parameters. For example, the information processing system described in each exemplary embodiment may have a configuration related to homomorphic encryption or the like that allows calculations to be performed while keeping the model parameters confidential, so that the confidentiality of the model parameters themselves can be ensured.

[0196] As an example, the providing unit 17 of each client device 1A may include an encryption unit that encrypts the model parameters using homomorphic encryption or the like, and the aggregating unit 32 of the server device 3A may aggregate the encrypted model parameters while keeping them confidential. Alternatively, the providing unit 33 of the server device 3A may include an encryption unit that encrypts the model parameters of the source model using homomorphic encryption or the like, and the acquiring unit 13 of each client device 1A may decrypt the model parameters.

[0197] Furthermore, it is preferable to have a configuration that can minimize the data size of model parameters when they are transmitted from each client device to the server device. For example, each client device 1A and the server device 3A may be configured to compress the model parameters. Furthermore, at least one of each client device 1A and the server device 3A may be configured to include a model reconfiguration unit that reduces the size of the model by reconfiguring the model (generating a distilled model, a derived model, a pseudo model, or a higher-level model). For example, such a configuration is suitable in situations where the processing performance of the client device is limited.

[0198] Furthermore, as an example, when each client device 1A is realized as a wearable device, the client device 1A may be configured to control the transmission of model parameters to the server device 3A depending on the remaining battery power. For example, when the remaining battery power is below a predetermined level, the client device 1A may be configured to transmit to the server device 3A only model parameters whose change from the previous value is equal to or greater than a predetermined value (percentage). Alternatively, the client device 1A may be configured to apply a sampling process to the acquired data (for example, randomly sampling to extract only 10%) and perform a learning process by the learning unit 12 using only the sampled data. Alternatively, the client device 1A may be configured to store the acquired data in another device (for example, the server device 3A).

[0199] Furthermore, since there is a possibility that noise may be present in the learning data and the parameters, each client device 1A and the server device 3A may have a configuration for improving resistance to the noise. For example, each client device 1A and the server device 3A may have a configuration for performing error correction on the learning data or the model parameters.

[0200] [Example of implementation by software] Some or all of the functions of the information processing devices (client devices) 1, 1-1, 1-2, 1A, 2A, 1A-1, 1A-2, ..., 2, 2A and the server devices 3, 3A (hereinafter also referred to as "the above-mentioned devices") may be implemented by hardware such as an integrated circuit (IC chip), or by software.

[0201] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 12. Figure 12 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.

[0202] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.

[0203] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0204] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0205] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0206] [Additional Notes] This disclosure includes the technologies described in the following supplementary notes. However, the present invention is not limited to the technologies described in the following supplementary notes, and various modifications are possible within the scope of the claims.

[0207] (Appendix A1) An information processing device comprising: an integration means for integrating a plurality of models included in a model group including one or more models obtained by directly or indirectly using a model provided from a server device; and a learning means for executing a learning process using the model integrated by the integration means.

[0208] (Appendix A2) The information processing device described in Appendix A1, wherein each of the one or more models included in the model group is obtained by one or more learning steps, the integration means integrates a model obtained by the nth learning step with one or more models obtained by a learning step prior to the nth step, and the learning means generates the (n+1)th model by executing a learning process using the models integrated by the integration means.

[0209] (Appendix A3) The information processing device according to Appendix A2, wherein the integration means performs integration processing by taking a weighted sum of a plurality of parameters defining a model obtained by the n-th learning step and a plurality of parameters defining a model obtained by the (n-1)-th learning step.

[0210] (Supplementary Note A4) The information processing device according to Supplementary Note A3, wherein the integration means sets weights in the weighted sum according to categories to which each of the plurality of parameters belongs.

[0211] (Appendix A5) An information processing device according to any one of Appendices A1 to A4, comprising: an acquisition means for acquiring input data; and a pseudo label generation means for generating pseudo labels by referring to the input data, wherein the learning means executes a learning process by referring to the input data and the pseudo labels.

[0212] (Supplementary Note A6) The information processing device according to Supplementary Note A5, wherein the pseudo label generating means generates the pseudo labels according to a result of a clustering process on the input data.

[0213] (Appendix A7) The information processing device according to any one of Appendices A1 to A6, further comprising a presentation means for presenting information relating to at least one of the following: at least one model included in the model group; a model integrated by the integration means; and a trained model obtained by applying a learning process to the integrated model.

[0214] (Supplementary Note A8) The information processing device according to any one of Supplementary Notes A2 to A4, further comprising a repetition control means for controlling repetition of the learning step, wherein the repetition control means repeats the learning step until a predetermined condition is satisfied.

[0215] (Appendix A9) An information processing device comprising: an acquisition means for acquiring input data; and an inference means for executing an inference process on the input data using a trained model obtained by a learning process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device.

[0216] (Appendix A10) An information processing system including a server device and a plurality of client devices, wherein the server device comprises: an acquisition means for acquiring a model from each of the plurality of client devices; an aggregation means for aggregating the models from each of the client devices; and a provision means for providing the model aggregated by the aggregation means to each of the plurality of client devices, and each of the plurality of client devices comprises: an integration means for integrating a plurality of models included in a model group including one or more models obtained directly or indirectly using a model provided from the server device; and a learning means for executing a learning process using the model integrated by the integration means.

[0217] (Supplementary Note A11) The information processing system according to Supplementary Note A10, wherein the aggregating means clusters the models from the client devices, and aggregates the models for each cluster obtained by the clustering.

[0218] (Supplementary Note A12) The information processing device according to Supplementary Note A10 or A11, wherein the server device includes a generation unit that generates a pre-trained model to be provided to each of the client devices.

[0219] (Appendix A13) An information processing method including: integrating a plurality of models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device; and executing a learning process using the integrated model.

[0220] (Appendix A14) An information processing method including: acquiring input data; and performing inference processing on the input data using a trained model obtained by a training process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device.

[0221] (Appendix A15) A program for causing a computer to function as an information processing device, the program causing the computer to: integrate multiple models included in a model group including one or more models obtained by directly or indirectly using a model provided from a server device; and execute a learning process using the integrated model.

[0222] (Appendix A16) A program for causing a computer to function as an information processing device, the program causing the computer to acquire input data and perform inference processing on the input data using a trained model obtained by a training process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device.

[0223] DESCRIPTION OF SYMBOLS 1, 2, 1A, 2A ... Information processing device (client device) 11 ... Integration unit 12 ... Learning unit 13, 21 ... Acquisition unit 14 ... Pseudo label generation unit 15 ... Presentation unit 16 ... Repetition control unit 17 ... Provision unit 22 ... Inference unit 3, 3A ... Server device 31 ... Acquisition unit 32 ... Aggregation unit 33 ... Provision unit 34 ... Generation unit 100A ... Information processing system

Claims

1. An information processing device comprising: an integration means for integrating a plurality of models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device; and a learning means for executing a learning process using the model integrated by the integration means.

2. The information processing device of claim 1, wherein each of the one or more models included in the model group is obtained by one or more learning steps, the integration means integrates a model obtained by the nth learning step with one or more models obtained by a learning step prior to the nth step, and the learning means generates the (n+1)th model by executing a learning process using the model integrated by the integration means.

3. The information processing device according to claim 2, wherein the integration means performs integration processing by taking a weighted sum of a plurality of parameters defining the model obtained in the nth learning step and a plurality of parameters defining the model obtained in the (n-1)th learning step.

4. The information processing apparatus according to claim 3, wherein said integration means sets weights in said weighted sum according to a category to which each of said plurality of parameters belongs.

5. An information processing device as claimed in any one of claims 1 to 4, comprising: an acquisition means for acquiring input data; and a pseudo label generation means for generating pseudo labels by referring to the input data, wherein the learning means executes a learning process by referring to the input data and the pseudo labels.

6. The information processing device according to claim 5, wherein the pseudo label generating means generates the pseudo labels according to a result of a clustering process performed on the input data.

7. An information processing device as claimed in any one of claims 1 to 6, further comprising a presentation means for presenting information relating to at least any of the following: at least any of the models included in the model group; a model integrated by the integration means; a trained model obtained by applying a learning process to the integrated model; and information obtained using at least any of the models, the integrated model, and the trained model, the information assisting a user in making decisions.

8. The information processing device according to any one of claims 2 to 4, further comprising a repetition control means for controlling the repetition of the learning step, wherein the repetition control means repeats the learning step until a predetermined condition is satisfied.

9. An information processing device comprising: an acquisition means for acquiring input data; and an inference means for executing inference processing on the input data using a trained model obtained by a learning process using a model obtained by integrating a plurality of models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device.

10. An information processing system including a server device and a plurality of client devices, wherein the server device comprises an acquisition means for acquiring a model from each of the plurality of client devices, an aggregation means for aggregating the models from each of the client devices, and a provision means for providing each of the plurality of client devices with the model aggregated by the aggregation means, and each of the plurality of client devices comprises an integration means for integrating a plurality of models included in a model group including one or more models obtained directly or indirectly using the model provided from the server device, and a learning means for executing a learning process using the model integrated by the integration means.

11. The information processing system according to claim 10, wherein the aggregation means clusters the models from each client device, and aggregates the models for each cluster obtained by the clustering.

12. The information processing system according to claim 10 or 11, wherein the server device is provided with a generation means for generating a pre-trained model to be provided to each of the client devices.

13. An information processing method comprising: integrating a plurality of models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device; and executing a learning process using the integrated model.

14. An information processing method including: acquiring input data; and performing inference processing on the input data using a trained model obtained by a learning process using a model obtained by integrating multiple models included in a model group including one or more models obtained directly or indirectly using a model provided from a server device.

15. A program for causing a computer to function as the information processing device according to claim 1, the program causing the computer to function as said integration means and said learning means.

16. A program for causing a computer to function as the information processing device according to claim 9, the program causing a computer to function as said acquisition means and said inference means.

Citation Information

Patent Citations

  • Personalized federal learning method based on parameter layering

    CN115587633A

  • Personalized federal learning method and device based on graph data

    CN116011597A

  • Model learning method, model learning system, and computer program

    JP2022076275A

  • Information processing method, information processing device and server device

    JP2023093838A

  • State prediction device and state prediction control method

    WO2019163141A1