Information processing system, client device, server device, information processing method, and program

The system enables personalized model adaptation in federated learning by training local models on branched global layers and updating global models with integrated weight coefficients, addressing inefficiencies and cost issues in existing techniques.

WO2025158560A1PCT designated stage Publication Date: 2025-07-31NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/001982
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing federated learning techniques fail to ensure that each client device accurately obtains a model tailored to its specific needs, leading to inefficiencies and increased computational costs.

Method used

A system where client devices train a local model based on a global model with branched layers, integrating branches using weight coefficients, and a server updates the global model based on local models, ensuring each client device receives a personalized model without excessive computational overhead.

Benefits of technology

Each client device obtains a model accurately adapted to its own device, reducing computational costs and improving model precision while reflecting knowledge from other devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024001982_31072025_PF_FP_ABST
    Figure JP2024001982_31072025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing system includes a server device and a plurality of client devices. Each client device comprises a local model training unit that trains, by machine learning, a local model in which a prescribed layer among a plurality of layers is not branched, such training being performed on the basis of a global model which is provided from a server device and in which the prescribed layer is branched into a plurality of branches. The server device comprises a global model updating unit that updates the global model on the basis of local models provided from the plurality of client devices.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, client device, server device, information processing method, and program

[0001] The present disclosure relates to an information processing system, an information processing method, a client device, a server device, and a program.

[0002] A technique called federated learning is known as a machine learning technique for distributed data sets. In federated learning, each client device learns a model distributed from a server device and transmits it to the server device. The server device then aggregates the received models and distributes them to each client device. This process is repeated. Each client device can obtain a model that reflects the knowledge of other client devices without disclosing its own knowledge to the other client devices.

[0003] In addition, as opposed to federated learning in which multiple client devices obtain the same model, federated learning in which each client device obtains a model that is compatible with its own device is known. For example, Patent Document 1 describes a technology for performing federated learning using multiple different types of common models for performing image analysis. In this technology, a server device selects one of multiple common models to distribute to each client device in accordance with the image analysis performed on the client device.

[0004] Japanese Patent Application Publication No. 2023-51723

[0005] In the technology described in Patent Literature 1, the model that is suited to a certain client device is not necessarily the only one of the multiple common models selected by the server device. Therefore, simply distributing the single common model selected by the server device to each client device and performing federated learning leaves room for improvement in terms of each of the multiple client devices obtaining a model that is tailored to their own device with high accuracy.

[0006] The present disclosure has been made in consideration of the above-mentioned problems, and one exemplary purpose thereof is to provide a technology that enables each of multiple client devices to obtain a model that accurately matches the device itself in federated learning.

[0007] An information processing system according to an exemplary aspect of the present disclosure includes a plurality of client devices equipped with a local model learning means that trains a local model in which a specified layer among multiple layers is not branched by machine learning based on a global model provided from a server device, the specified layer being branched into a plurality of branches, and the server device that includes a global model update means that updates the global model based on the local model provided from the plurality of client devices.

[0008] A client device according to an exemplary aspect of the present disclosure includes a receiving means for receiving from a server device a global model in which a predetermined layer among a plurality of layers is branched into a plurality of branches, a local model learning means for training a local model in which the predetermined layer is not branched by machine learning based on the global model, and a transmitting means for transmitting the local model to the server device.

[0009] A server device according to an exemplary aspect of the present disclosure includes a receiving means for receiving, from a plurality of client devices, a local model trained by machine learning based on a global model in which a predetermined layer among a plurality of layers branches into a plurality of branches, the predetermined layer not branching, a global model update means for updating the global model based on the local model, and a transmitting means for transmitting the global model to the plurality of client devices.

[0010] Another client device according to an exemplary aspect of the present disclosure includes an input data acquisition means for acquiring input data, and an inference means for making inferences about the input data using a local model trained by machine learning based on a global model provided by a server device, in which a predetermined layer among multiple layers is branched into multiple branches, and in which the predetermined layer is not branched.

[0011] An information processing method according to an exemplary aspect of the present disclosure includes a local model learning process in which at least one processor included in a plurality of client devices uses machine learning to train a local model in which a predetermined layer among multiple layers is not branched, based on a global model provided from a server device, the global model having the predetermined layer branched into a plurality of branches; and a global model update process in which at least one processor included in the server device updates the global model based on the local models provided from the plurality of client devices.

[0012] An information processing method according to an exemplary aspect of the present disclosure includes a receiving process in which at least one processor included in a client device receives from a server device a global model in which a specified layer among multiple layers is branched into multiple branches; a local model learning process in which the at least one processor trains a local model in which the specified layer is not branched using machine learning based on the global model; and a transmitting process in which the at least one processor transmits the local model to the server device.

[0013] An information processing method according to an exemplary aspect of the present disclosure includes a receiving process in which at least one processor included in a server device receives, from a plurality of client devices, a local model trained by machine learning based on a global model in which a predetermined layer among a plurality of layers branches into a plurality of branches, and the predetermined layer does not branch; a global model update process in which the at least one processor updates the global model based on the local model; and a transmitting process in which the at least one processor transmits the global model to the plurality of client devices.

[0014] An information processing method according to an exemplary aspect of the present disclosure includes an input data acquisition process in which at least one processor included in a client device acquires input data, and an inference process in which the at least one processor performs inference on the input data using a local model in which a predetermined layer among multiple layers is not branched, the local model being trained by machine learning based on a global model provided by a server device, the global model having the predetermined layer branching into multiple branches.

[0015] A program according to an exemplary aspect of the present disclosure is a program for causing a computer to function as the above-described client device, and causes the computer to function as the receiving means, the local model learning means, and the transmitting means.

[0016] A program according to an exemplary aspect of the present disclosure is a program for causing a computer to function as the server device described above, and causes the computer to function as the receiving means, the global model updating means, and the transmitting means.

[0017] A program according to an exemplary aspect of the present disclosure is a program for causing a computer to function as the other client device described above, and causes the computer to function as the input data acquisition means and the inference means.

[0018] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that a technology can be provided in which, in federated learning, each of multiple client devices can obtain a model that is accurately adapted to the device itself.

[0019] FIG. 1 is a diagram schematically illustrating an example configuration of a global model and a local model according to the present disclosure. FIG. 2 is a block diagram illustrating a configuration of an information processing system according to the present disclosure. FIG. 3 is a flow diagram illustrating the flow of an information processing method according to the present disclosure. FIG. 4 is a block diagram illustrating a configuration of a client device according to the present disclosure. FIG. 5 is a flow diagram illustrating the flow of an information processing method according to the present disclosure. FIG. 6 is a block diagram illustrating a configuration of a client device according to the present disclosure. FIG. 7 is a flow diagram illustrating the flow of an information processing method according to the present disclosure. FIG. 8 is a block diagram illustrating a configuration of a server device according to the present disclosure. FIG. 9 is a flow diagram illustrating the flow of an information processing method according to the present disclosure. FIG. 10 is a block diagram illustrating a configuration of an information processing system according to the present disclosure. FIG. 11 is a diagram schematically illustrating an example configuration of a global model according to the present disclosure. FIG. 12 is a diagram schematically illustrating an example configuration of an integrated model and a local model according to the present disclosure. FIG. 13 is a diagram schematically illustrating a specific example of a global model update process according to the present disclosure. FIG. 14 is a flow diagram illustrating the flow of an information processing method according to the present disclosure. FIG. 15 is a block diagram illustrating an example hardware configuration of a computer functioning as each device according to the present disclosure.

[0020] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.

[0021] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0022] (Overview of Information Processing System 100) The information processing system 100 is a system that generates a model (hereinafter also referred to as a personalized model) suited to each individual client device 1 while mutually providing a global model and a local model between each of a plurality of client devices 1 and a server device 3. The information processing system 100 can also be considered as a system that performs so-called federated learning, but this term does not limit this exemplary embodiment. Note that providing a model includes providing some or all of the parameters that make up the model.

[0023] Any model configured with multiple layers can be applied as the global model and the local model. As an example, a neural network model can be applied as the global model and the local model. However, the global model and the local model are not limited to a neural network model, and may be other types of models configured with multiple layers. Furthermore, the global model and the local model may be a model that combines multiple models so that each of the multiple models is included as a layer. In this case, the multiple models may or may not include a neural network model. Furthermore, the multiple models may be models of the same type, or at least one model may be a model of a different type from the other models.

[0024] An example of the configuration of a global model and a local model will be described with reference to FIG. 1. FIG. 1 is a diagram schematically illustrating an example of the configuration of a global model and a local model. Hereinafter, each of the global model and the local model may be simply referred to as a model. As shown in FIG. 1, each model has multiple layers corresponding to each other, and in FIG. 1, each model includes three layers, a first layer, a second layer, and a third layer, which correspond to each other. An input to each model is input to the first layer, and an output from the first layer is input to the second layer. An output from the second layer is input to the third layer, and an output from the third layer is output from the model. Note that while FIG. 1 illustrates an example in which each model includes three layers, each model may include four or more layers.

[0025] Also, as shown in FIG. 1 , in the global model, a predetermined layer branches into multiple branches. In this example, the second layer branches into three branches. "A predetermined layer branches into multiple branches" means that the predetermined layer includes multiple branches that may have different outputs in response to inputs from a previous layer, and the outputs from the multiple branches are superimposed and output to a subsequent layer. "Outputs from the multiple branches" may refer to "outputs from at least two of the multiple branches" or "outputs from each of the multiple branches." The type of processing by which one branch converts an input into an output may be the same as or different from the type of processing by another branch.

[0026] While FIG. 1 illustrates an example in which the second layer branches into three branches, the number of branches (hereinafter also referred to as the number of branches) may be two, four, or more. Although FIG. 1 illustrates an example in which the second layer branches in the global model, the first or third layer may also branch. The global model may also include multiple branched layers. For example, two or more of the multiple layers included in the global model may each be branched, or all of the multiple layers may each be branched. When the global model includes multiple branched layers, the number of branches in a certain layer may be the same as or different from the number of branches in another certain layer.

[0027] (Configuration of Information Processing System 100) The configuration of the information processing system 100 will be described with reference to FIG. 2. FIG. 2 is a block diagram showing the configuration of the information processing system 100. As shown in FIG. 2, the information processing system 100 includes a plurality of client devices 1-1, 1-2, 1-3, ... and a server device 3. When it is not necessary to distinguish between the client devices 1-1, 1-2, 1-3, ..., they will also be referred to as client devices 1. Although FIG. 2 shows three client devices 1, the number of client devices 1 included in the information processing system 100 may be two, four or more.

[0028] 2, each client device 1 and the server device 3 are communicatively connected via a network N. The specific configuration of the network N does not limit the present exemplary embodiment, but examples thereof include a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination of these networks.

[0029] The client device 1 includes a local model learning unit 11. The local model learning unit 11 trains a local model by machine learning, based on a global model provided by the server device 3, in which a predetermined layer among multiple layers is branched into multiple branches, a local model in which the predetermined layer is not branched. The local model learning unit 11 is an example of a configuration that realizes a local model learning means. For example, if the client device 1 includes at least one processor, the local model learning unit 11 is realized by the at least one processor executing a program.

[0030] For example, the local model learning unit 11 may generate a local model in which the predetermined layer is not branched by integrating each branch of the predetermined layer in the global model, and train the local model by machine learning.

[0031] Furthermore, for example, the local model learning unit 11 may train the local model by machine learning so that the difference between a specified layer of any local model in which the specified layer is not branched and each branch of the specified layer of the global model is small.

[0032] The server device 3 includes a global model update unit 31. The global model update unit 31 updates the global model based on the local models provided by the multiple client devices 1. The "local models provided by the multiple client devices 1" may refer to "the local models provided by at least two of the multiple client devices 1" or "the local models provided by each of the multiple client devices 1." The global model update unit 31 is an example of a configuration that realizes a global model update means. For example, if the server device 3 includes at least one processor, the global model update unit 31 is realized by the at least one processor executing a program.

[0033] For example, the global model updating unit 31 may update a first branch in a predetermined layer of the global model by overlapping a predetermined layer in each of the multiple local models according to a first weighting, and may update a second branch in a predetermined layer of the global model by overlapping a predetermined layer in each of the multiple local models according to a second weighting.

[0034] (Effects of Information Processing System 100) As described above, the information processing system 100 employs a configuration in which each client device 1 includes the above-described local model learning unit 11, and the server device 3 includes the above-described global model updating unit 31. For this reason, it can be expected that a predetermined layer of a local model obtained by a certain client device 1 reflects multiple branches of a predetermined layer in a global model updated based on local models trained by the multiple client devices 1, so as to be adapted to the certain client device 1. Therefore, the information processing system 100 achieves the effect that, in federated learning, each of the multiple client devices 1 can obtain a model that accurately adapts to the device itself.

[0035] As a comparative example to the information processing system 100, when a predetermined layer of a global model branches into multiple branches, a configuration in which each branch is trained individually on a client device, i.e., a configuration in which federated learning is performed for each branch, can be considered. In this comparative example, the computational cost of the client device increases depending on the number of branches. In contrast, the information processing system 100 trains a local model in which the predetermined layer is not branched, thereby achieving the effect of suppressing the increase in computational cost of the client device compared to the comparative example.

[0036] (Flow of Information Processing Method S100) Next, the flow of the information processing method S100 executed by the information processing system 100 will be described with reference to Fig. 3. Fig. 3 is a flow diagram illustrating the flow of the information processing method S100. As shown in Fig. 3, the information processing method S100 includes a local model learning process (step S11) and a global model updating process (step S31).

[0037] In step S11, at least one processor (e.g., local model learning unit 11) provided in each of the multiple client devices 1 trains a local model in which a specified layer is not branched by machine learning based on a global model provided by the server device 3, in which a specified layer among multiple layers is branched into multiple branches.

[0038] In step S31, at least one processor (for example, the global model update unit 31) included in the server device 3 updates the global model described above based on the local models provided by the multiple client devices 1.

[0039] The information processing system 100 repeatedly executes steps S11 and S31. Note that the information processing method S100 may end when an end condition is satisfied.

[0040] (Effects of Information Processing Method S100) As described above, the information processing method S100 includes the local model training process (step S11) and the global model updating process (step S31). Therefore, the information processing method S100 can achieve the same effects as the information processing system 100.

[0041] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0042] (Configuration of Client Device 2) The client device 2 is, as an example, a device that executes inference processing using a local model. The configuration of the client device 2 will be described with reference to FIG. 4. FIG. 4 is a block diagram showing the configuration of the client device 2. As shown in FIG. 4, the client device 2 includes an input data acquisition unit 21 and an inference unit 22. The input data acquisition unit 21 and the inference unit 22 are examples of configurations that realize an acquisition means and an inference means. For example, if the client device 2 includes at least one processor, the input data acquisition unit 21 and the inference unit 22 are realized by the at least one processor executing a program.

[0043] The input data acquisition unit 21 acquires input data. Here, the input data may be, for example, image data, text data, or other data, or a combination thereof. The input data acquired by the input data acquisition unit 21 is, for example, treated as a target for inference processing by the client device 2 in the inference phase.

[0044] The inference unit 22 performs inference on the input data using a local model. The local model is a global model provided by the server device, in which a predetermined layer among multiple layers is branched into multiple branches, and is trained by machine learning based on the global model, in which the predetermined layer is not branched. The local model may be, for example, the local model obtained by the client device 1 in exemplary embodiment 1.

[0045] (Effects of Client Device 2) As described above, the client device 2 is configured to include the above-described input data acquisition unit 21 and the above-described inference unit 22. For this reason, it is expected that a predetermined layer of the local model used for inference processing in the client device 2 reflects multiple branches of a predetermined layer in the global model updated based on the local models trained by the multiple client devices, so as to be adapted to the client device 2. Therefore, the client device 2 has the effect of being able to perform highly accurate inference by utilizing a model that is adapted to the client device 2 with high accuracy.

[0046] (Flow of Information Processing Method S2) Next, the flow of the information processing method S2 executed by the client device 2 will be described with reference to Fig. 5. Fig. 5 is a flow diagram showing the flow of the information processing method S2. As shown in Fig. 5, the information processing method S2 includes an input data acquisition process (step S21) and an inference process (step S22).

[0047] In step S21, at least one processor (for example, the input data acquisition unit 21) acquires input data.

[0048] In step S22, at least one processor (e.g., the inference unit 22) performs inference on the input data using a local model. The local model is a global model provided by a server device, in which a predetermined layer among multiple layers is branched into multiple branches, and the predetermined layer is not branched. The local model is trained by machine learning based on the global model.

[0049] (Effects of Information Processing Method S2) As described above, the information processing method S2 includes the input data acquisition process (step S21) and the inference process (step S22). Therefore, the information processing method S2 can achieve the same effects as the client device 2.

[0050] [Third Exemplary Embodiment] A third exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0051] (Configuration of client device 1A) The client device 1A is a device that functions as one of multiple client devices in federated learning. The configuration of the client device 1A will be described with reference to FIG. 6. FIG. 6 is a block diagram showing the configuration of the client device 1A. As shown in FIG. 6, the client device 1A includes a receiving unit 10, a local model learning unit 11, and a transmitting unit 12. The receiving unit 10, the local model learning unit 11, and the transmitting unit 12 are an example of a configuration that realizes a receiving means, a local model learning means, and a transmitting means. For example, if the client device 1A includes at least one processor, the receiving unit 10, the local model learning unit 11, and the transmitting unit 12 are realized by the at least one processor executing a program.

[0052] The receiving unit 10 receives a global model in which a predetermined layer among multiple layers branches into multiple branches from a server device. The local model learning unit 11 has the same configuration as in the first exemplary embodiment, and therefore detailed description will not be repeated. The transmitting unit 12 transmits the local model to the server device. Note that the server device may be the server device 3 described above or one of the server devices 3A and 3B described below, or another server device.

[0053] (Effects of the Client Device 1A) As described above, the present exemplary embodiment employs a configuration including the above-described receiving unit 10, the above-described local model learning unit 11, and the above-described transmitting unit 12. Therefore, the client device 1A can obtain the same effects as the above-described information processing system 100.

[0054] (Flow of Information Processing Method S1A) Next, the flow of the information processing method S1A executed by the client device 1A will be described with reference to Fig. 7. Fig. 7 is a flow diagram showing the flow of the information processing method S1A. As shown in Fig. 7, the information processing method S1A includes a receiving process (step S10), a local model learning process (step S11), and a transmitting process (step S12).

[0055] In step S10, at least one processor (for example, the receiving unit 10) receives from the server device a global model in which a predetermined layer among multiple layers is branched into multiple branches.

[0056] In step S11, at least one processor (for example, the local model learning unit 11) trains a local model in which a predetermined layer is not branched by machine learning based on the global model.

[0057] In step S12, at least one processor (for example, the transmitting unit 12) transmits the local model to the server device.

[0058] The client device 1A repeatedly executes steps S10 to S12. The information processing method S1A may end when an end condition is satisfied.

[0059] (Effects of Information Processing Method S1A) As described above, the information processing method S1A includes the above-described receiving process (step S10), the above-described local model learning process (step S11), and the above-described transmitting process (step S12). Therefore, the information processing method S1A can achieve the same effects as the client device 1A.

[0060] [Fourth Exemplary Embodiment] A fourth exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0061] (Configuration of Server Device 3A) The server device 3A is a device that functions as a server device in federated learning. The configuration of the server device 3A will be described with reference to FIG. 8. FIG. 8 is a block diagram showing the configuration of the server device 3A. As shown in FIG. 8, the server device 3A includes a receiving unit 30, a global model updating unit 31, and a transmitting unit 32. The receiving unit 30, the global model updating unit 31, and the transmitting unit 32 are an example of a configuration that realizes a receiving means, a global model updating means, and a transmitting means. For example, if the server device 3A includes at least one processor, the receiving unit 30, the global model updating unit 31, and the transmitting unit 32 are realized by the at least one processor executing a program.

[0062] The receiving unit 30 receives, from a plurality of client devices, local models trained by machine learning based on a global model in which a predetermined layer among the plurality of layers branches into a plurality of branches, but in which the predetermined layer does not branch. The global model update unit 31 has the same configuration as in the first exemplary embodiment, and therefore detailed description will not be repeated. The transmitting unit 32 transmits the global model to the plurality of client devices. Note that "receiving from the plurality of client devices" may mean "receiving from each of at least two of the plurality of client devices" or "receiving from each of the plurality of client devices." Note that "transmitting to the plurality of client devices" may mean "transmitting to each of at least two of the plurality of client devices" or "transmitting to each of the plurality of client devices." Note that each of the plurality of client devices may be any of the client devices 1 and 1A described above and the client device 1B described below, or another client device.

[0063] (Effects of Server Device 3A) As described above, the server device 3A is configured to include the above-described receiving unit 30, the above-described global model updating unit 31, and the above-described transmitting unit 32. Therefore, the server device 3A can provide functions as a server device for achieving the same effects as those of the information processing system 100 described above.

[0064] (Flow of Information Processing Method S3A) Next, the flow of the information processing method S3A executed by the server device 3A will be described with reference to Fig. 9. Fig. 9 is a flow diagram showing the flow of the information processing method S3A. As shown in Fig. 9, the information processing method S3A includes a receiving process (step S30), a global model updating process (step S31), and a transmitting process (step S32).

[0065] In step S30, at least one processor (e.g., receiving unit 30) receives local models from multiple client devices, the local models being trained by machine learning based on a global model in which a predetermined layer among multiple layers branches into multiple branches, and in which the predetermined layer does not branch.

[0066] In step S31, at least one processor (for example, the global model update unit 31) updates the global model based on the local model.

[0067] In step S32, at least one processor (for example, the transmitting unit 32) transmits the global model to a plurality of client devices.

[0068] The server device 3A repeatedly executes steps S30 to S32. The information processing method S3A may end when an end condition is satisfied.

[0069] (Effects of Information Processing Method S3A) As described above, the information processing method S3A includes the above-described receiving process (step S30), the above-described global model updating process (step S31), and the above-described transmitting process (step S32). Therefore, the information processing method S3A has the same effects as the server device 3A.

[0070] Fifth Exemplary Embodiment A fifth exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiments will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0071] (Configuration of Information Processing System 100B) The configuration of an information processing system 100B, which is a modification of the information processing system 100 according to the first exemplary embodiment, will be described with reference to FIG. 10. FIG. 10 is a block diagram showing the configuration of the information processing system 100B. As shown in FIG. 10, the information processing system 100B includes a plurality of client devices 1B-k (k=1, 2, ...) and a server device 3B. Also, as shown in FIG. 10, the server device 3B and each client device 1B-k are communicably connected via a network N. A specific example of the network N is as described in the first exemplary embodiment, and therefore detailed description will not be repeated. Details of the functional blocks included in each device will be described later.

[0072] In this exemplary embodiment, a neural network model is applied as the global model c and the local model w_k transmitted and received between each client device 1B-k and the server device 3B. The "_k" in "w_k" is an alternative notation for the subscript k. Similarly, "_" is used below as an alternative notation for the subscript. For example, examples of neural network models include, but are not limited to, models based on machine learning algorithms such as CNN (Convolutional Neural Network) and RNN (Recurrent Neural Network), as well as combinations thereof. The global model c and the local model w_k have the same layer structure, each consisting of multiple layers (e.g., at least three layers: an input layer, an intermediate layer, and an output layer). Each layer includes a linear layer and a nonlinear layer. The linear layer is a layer that performs linear transformation. Examples of linear layers include, but are not limited to, a convolutional layer and a fully connected layer. The nonlinear layer is a layer that performs nonlinear transformation. Examples of nonlinear layers include, but are not limited to, a normalization layer that performs a transformation using a normalization function, an activation layer that performs a change using an activation function, and combinations of these.

[0073] (Client Device 1B-k) The configuration of client device 1B-k will be described with reference to FIG. 10. As shown in FIG. 10, client device 1B-k includes a control unit 110, a storage unit 120, and a communication unit 130. The control unit 110 controls each unit of client device 1B-k. The storage unit 120 stores various data used by the control unit 110. For example, the storage unit 120 stores a weighting coefficient α^k_b, training data, and a local model w_k. The communication unit 130 communicates with server device 3B via network N under the control of the control unit 110.

[0074] The control unit 110 includes a receiving unit 10, a local model learning unit 11, a transmitting unit 12, and a weighting factor learning unit 13. The weighting factor learning unit 13 is an example of a configuration that realizes a weighting factor learning means. If the client device 1B-k includes at least one processor, the weighting factor learning unit 13 is realized by the at least one processor executing a program.

[0075] (Receiving Unit 10) The receiving unit 10 receives a global model c in which a predetermined layer is branched into multiple branches from the server device 3B. The global model c is a model that is updated by the server device 3B based on the local models w_k received from each client device 1B-k. The updated global model c is transmitted from the server device 3B to each client device 1B-k.

[0076] An example of the configuration of the global model c will be described with reference to Fig. 11. Fig. 11 is a diagram schematically illustrating an example of the configuration of the global model c. As shown in Fig. 11, in the i-th layer of the global model c, the linear layer is branched into a plurality of branches C_1, C_2, and C_3. Each of the branches C_1, C_2, and C_3 is configured to include model parameters for performing a predetermined linear transformation on an input.

[0077] 11 shows three branches C_1, C_2, and C_3, assuming that the number of branches is three, but the number of branches is not limited to three and may be two or more. The fact that a linear layer in a certain layer branches into multiple branches is also simply referred to as "branching." The i-th layer among the branched layers in the global model c is referred to as the "i-th layer."

[0078] More specifically, as shown in FIG. 11 , inputs from the previous layer to the i-th layer are input to each branch C_1, C_2, and C_3 and linearly transformed. Furthermore, a superposition (e.g., weighted sum) of outputs from each branch C_1, C_2, and C_3 is input to a nonlinear layer and nonlinearly transformed. The output from the nonlinear layer is output from the i-th layer. Note that if the input layer is branched, the input to the global model c is input to each branch C_1, C_2, and C_3 of the input layer. If the output layer is branched, the output from the i-th layer is output from the global model c.

[0079] Although FIG. 11 shows one branched i-th layer, the global model c may include multiple branched i-th layers (i = 1, 2, ...). In this case, the number of branches included in a certain layer among the multiple branched i-th layers may be the same as or different from the number of branches included in other layers. Furthermore, the type of linear layer in a certain layer among the multiple branched i-th layers may be the same as or different from the type of linear layer in other layers. The global model c is expressed, for example, by the following equation (1):

[0080] In formula (1), C_(i, b) (b = 1, 2, ..., B_i) represents a branch in the i-th layer. B_i is the number of branches in the i-th layer. Note that "branches C_1 to C_3 in the i-th layer" in FIG. 11 can be expressed as branches C_(i, 1) to C_(i, 3) by applying formula (1). Note that of the multiple layers constituting the global model c, all layers may be branched, or some layers may not be branched. In formula (1), non-branched layers and non-linear layers in branched layers are omitted.

[0081] Note that when receiving the global model c for the second or subsequent time, the receiving unit 10 only needs to receive, among the model parameters constituting the global model c, the model parameters of each branch C_(i, b) of the linear layer in the i-th layer. In this case, the receiving unit 10 does not necessarily need to receive other model parameters (e.g., model parameters of layers that are not branched, model parameters of nonlinear layers in branched layers, etc.) from the second or subsequent time.

[0082] (Weighting Coefficient Learning Unit 13) The weighting coefficient learning unit 13 updates the weighting coefficients α^k_(i,b) through machine learning using an integrated model c_k in which multiple branches C_(i,b) in the global model c are integrated using the weighting coefficients α^k_(i,b). Here, "^k" in "α^k_(i,b)" is an alternative notation for the subscript k written as a superscript. Similarly, "^" is used below as an alternative notation for a superscript.

[0083] The weighting coefficient α^k_(i,b) is a weighting coefficient corresponding to each branch C_(i,b) in the global model c used in the client device 1B-k. The sum of the weighting coefficients α^k_(i,b) for the B_i weighting coefficients in the same layer with the same i is 1. When the global model c includes multiple branched i-th layers (e.g., an i1th layer and an i2th layer), even if the number of branches B_i1 and B_i2 are the same, the weighting coefficient α^k_(i1,b) for the i1th layer and the weighting coefficient α^k_(i2,b) for the i2th layer may be different. Furthermore, the weighting coefficient α^k_(i,b) may be different for each client device 1B-k. This is because the weighting coefficient α^k_(i,b) is updated in each client device 1B-k using the integrated model c_k to suit the device itself.

[0084] The integrated model c_k is a model used in machine learning to update the weighting coefficient α^k_b. The integrated model c_k is expressed by the following equation (2).

[0085] The integrated model c_k will be described with reference to FIG. 12 . FIG. 12 is a diagram schematically illustrating an example configuration of the integrated model c_k and a local model w_k, which will be described later. Note that in FIG. 12 , the i in the subscript "_(i, b)" is omitted and "_b" is used, assuming that the description is related to the i-th layer. As shown in FIG. 12 , the integrated model c_k includes a linear layer C_k that is not branched in the i-th layer. The linear layer C_k is a layer in which each branch C_b in the i-th layer of the global model c is integrated using a weighting coefficient α^k_b corresponding to the i-th layer. "Integrating" may mean, for example, using a weighted sum of the model parameters included in each branch C_(i, b) using a weighting coefficient α^k_(i, b) as the model parameter of the linear layer C_k, as shown in Equation (2).

[0086] For example, the weighting coefficient learning unit 13 updates the weighting coefficient α^k_(i,b) by performing machine learning in the integrated model c_k while fixing model parameters other than the weighting coefficient α^k_(i,b). For example, the weighting coefficient learning unit 13 performs several steps using training data to reduce cross-entropy loss. In addition, some or all of the multiple training data stored in the storage unit 120 are used in this machine learning.

[0087] Here, each of the multiple training data stored in the storage unit 120 is configured as a pair of input data that can be input to a local model w_k (described later) and correct output data corresponding to the input data. Because the input and output of the integrated model c_k are the same as those of the local model c_k, some or all of the multiple training data can also be used in machine learning using the integrated model c_k. The multiple training data may be different for each client device 1B-k and do not need to be made public to each other. Therefore, the weighting coefficient α^k_(i, b) updated by machine learning using some or all of such multiple training data becomes a weighting coefficient adapted to each client device 1B-k.

[0088] (Local Model Learning Unit 11) The local model learning unit 11 trains the local model w_k by machine learning based on the global model c and the updated weighting coefficient α^k_(i, b). The trained local model w_k is transmitted from the client device 1B-k to the server device 3B. With reference to FIG. 12, an example configuration of the local model w_k will be described. As shown in FIG. 12, the local model w_k has an unbranched linear layer W_k in the i-th layer. The linear layer W_k is configured to include model parameters for performing a predetermined linear transformation on the input.

[0089] For example, the local model learning unit 11 may integrate the differences between the multiple branches C_(i,b) and the ith layer of the local model w_k using the updated weighting coefficient α^k_(i,b), and train the local model w_k using the integrated difference as a regularization term. The "difference between the multiple branches C_(i,b) and the ith layer of the local model w_k" may be "the difference between each of at least two of the multiple branches C_(i,b) and the ith layer of the local model w_k," or may be "the difference between each of the multiple branches C_(i,b) and the ith layer of the local model w_k." As an example, the local model learning unit 11 may train the local model w_k using the loss function shown in the following equation (3).

[0090] In equation (3), L_ce represents the cross-entropy. λ is a hyperparameter. W^k_i represents the i-th layer in the local model w_k. The second term in equation (3) represents a regularization term. The regularization term represents a value obtained by adding up, for each i-th layer, a weighted sum calculated by assigning a weighting coefficient α^k_(i,b) to the difference between the i-th layer W^k_i and each branch C_(i,b) in the local model w_k. Note that any value may be used as the initial value of the model parameters constituting W^k_i. The local model learning unit 11 performs several steps of machine learning to update the local model w_k so as to reduce the loss calculated by equation (3).

[0091] Furthermore, the local model learning unit 11 trains the local model w_k using a plurality of training data. As described above, the plurality of training data is stored in the storage unit 120. Details of each training data will be explained in the same manner as for each training data used in the machine learning of the weighting coefficients α^k_(i,b), and therefore detailed explanations will not be repeated. Note that the plurality of training data used in the machine learning of the local model w_k may be partially or entirely the same as or partially or entirely different from the plurality of training data used in the machine learning of the weighting coefficients α^k_(i,b).

[0092] (Variation of the local model learning unit 11) When training the local model w_k based on the global model c and the updated weighting coefficients α^k_(i, b), the local model learning unit 11 can be modified as follows instead of using the regularization term described above.

[0093] The local model learning unit 11 may train the local model w_k using the above-described integrated model c_k to which the updated weight coefficient α^k_(i, b) is applied as the initial model of the local model w_k. In this case, machine learning of the local model w_k is performed so as to reduce the cross-entropy. In other words, only the first term on the right-hand side of Equation (3) is referenced, and the second term is unnecessary. This eliminates the need for the hyperparameter λ, thereby reducing the computational cost of adjusting the hyperparameter.

[0094] (Transmitter 12) The transmitter 12 transmits the trained local model w_k, the updated weighting coefficients α^k_(i,b), and the number n_k of training data to the server device 3B. The number n_k of training data may be the number of training data used to train the local model w_k, the number of training data used to update the weighting coefficients α^k_(i,b), the total number of these, or a total without duplication.

[0095] The transmitter 12 is only required to transmit at least the model parameters of the linear layer W_k of the i-th layer that have been updated through training, among the model parameters constituting the local model w_k. In this case, the transmitter 12 does not necessarily need to transmit other model parameters (e.g., model parameters of layers corresponding to layers that are not branched in the global model c, model parameters of nonlinear layers in the i-th layer, etc.).

[0096] (Server Device 3B) The configuration of the server device 3B will be described with reference to FIG. 10 again. As shown in FIG. 10, the server device 3B includes a control unit 310, a storage unit 320, and a communication unit 330. The control unit 310 controls the various units of the server device 3B. The storage unit 320 stores various data used by the control unit 310. For example, the storage unit 320 stores a global model c. The communication unit 330 communicates with the client device 1B-k via the network N under the control of the control unit 310.

[0097] The control unit 310 includes a receiving unit 30, a global model updating unit 31, a transmitting unit 32, and a global model generating unit 33. When the server device 3B includes at least one processor, the global model generating unit 33 is realized by the at least one processor executing a program.

[0098] (Receiving Unit 30) The receiving unit 30 receives the trained local model w_k, the weighting coefficient α^k_(i, b), and the number of training data n_k from each client device 1B-k. The number of training data n_k has been described above, and therefore will not be described in detail again. The receiving unit 30 is only required to receive at least the model parameters of the linear layer W_k of the i-th layer among the model parameters constituting the local model w_k. In this case, the receiving unit 30 does not necessarily need to receive other model parameters (e.g., model parameters of layers corresponding to layers not branched in the global model c, model parameters of nonlinear layers in the i-th layer, etc.).

[0099] (Global Model Generation Unit 33) The global model generation unit 33 generates an initial model by branching the i-th layer as an initial model of the global model c. For example, the global model generation unit 33 generates initial parameters for each branch C_(i, b) with the branch number B_i that functions as a linear layer in the i-th layer. Any value can be applied as the initial parameter.

[0100] (Global Model Update Unit 31) The global model update unit 31 updates the global model c by integrating the i-th layer in the multiple local models w_k for each branch C_(i,b) based on the weighting coefficients α^k_(i,b) provided from the multiple client devices 1B-k. Note that the "weighting coefficients α^k_(i,b) provided from the multiple client devices 1B-k" may be "weighting coefficients α^k_(i,b) provided from at least two of the multiple client devices 1B-k" or "weighting coefficients α^k_(i,b) provided from each of the multiple client devices 1B-k." Furthermore, the "i-th layer in the multiple local models w_k" may be "the i-th layer in each of at least two of the multiple local models w_k" or "the i-th layer in each of the multiple local models w_k." This makes it possible to update the branch C_(i, b) of the branched global model c based on multiple non-branched local models w_k, taking into consideration the degree to which each client device 1B-k places importance on each branch C_(i, b).

[0101] Furthermore, the global model update unit 31 updates the global model c by integrating the i-th layer of the multiple local models w_k for each branch C_(i, b) based on the number n_k of training data provided from the multiple client devices 1B-k. Note that the "number n_k of training data provided from the multiple client devices 1B-k" may be the "number n_k of training data provided from each of at least two of the multiple client devices 1B-k" or the "number n_k of training data provided from each of the multiple client devices 1B-k." Furthermore, the "i-th layer in the multiple local models w_k" may be the "i-th layer in each of at least two of the multiple local models w_k" or the "i-th layer in each of the multiple local models w_k." This makes it possible to update the branch C_(i, b) of the branched global model c by prioritizing the local model w_k trained using more training data among the multiple non-branched local models w_k.

[0102] For example, the global model update unit 31 updates the global model c using the following equation (4).

[0103] In the numerator on the right side of equation (4), n_k indicates the number of training data received from client device 1B-k, α^k_(i,b) indicates a weighting coefficient received from client device 1B-k, and W^k_i indicates the i-th linear layer received from client device 1B-k. In the denominator on the right side, n_m indicates the number of training data received from client device 1B-m, and α^m_(i,b) indicates a weighting coefficient received from client device 1B-m. The denominator on the right side is a term for normalizing the sum calculated in the numerator, taking into account the possibility that the sum of the weighting coefficients α^m_(i,b) received from each client device 1B-m for the same branch with the same b may not be 1. In equation (4), C_(i,b) on the left side is applied as the b-th branch in the i-th layer of the updated global model c.

[0104] The update of the global model c using equation (4) will be described with reference to Fig. 13. Fig. 13 is a diagram schematically illustrating a specific example of the global model update process. Note that in Fig. 13, the i in the subscript "_(i, b)" is omitted and the subscript is shown as "_b", assuming that the explanation is about the i-th layer. 13, the updated branch C_1 in the global model c is calculated by normalizing the sum of (i) model parameters for the linear layer W_1 trained on the client device 1B-1, weighted by the number n_1 of training data in the client device 1B-1 and a weighting coefficient α^1_1 for the branch C_1, (ii) model parameters for the linear layer W_2 trained on the client device 1B-2, weighted by the number n_2 of training data in the client device 1B-2 and a weighting coefficient α^2_1 for the branch C_1, and (iii) model parameters for the linear layer W_3 trained on the client device 1B-3, weighted by the number n_3 of training data in the client device 1B-3 and a weighting coefficient α^3_1 for the branch C_1. The same applies to the updated branches C_2 and C_3.

[0105] (Transmitter 32) The transmitter 32 transmits the global model c, in which the i-th layer is branched, to each client device 1B-k. Note that when transmitting the global model c for the second or subsequent times, the transmitter 32 only needs to transmit, among the model parameters constituting the global model c, the model parameters of each branch C_(i, b) of the linear layer in the i-th layer. In this case, the transmitter 32 does not necessarily need to transmit other model parameters from the second or subsequent times.

[0106] (Flow of Information Processing Method S100B) Next, the flow of information processing method S100B executed by information processing system 100B will be described with reference to Fig. 14. Fig. 14 is a flow diagram showing the flow of information processing method S100B. As shown in Fig. 14, information processing method S100B includes steps S101 to S105 and steps S301 to S304.

[0107] In step S301, the global model generating unit 33 of the server device 3B generates an initial model of a branched global model c.

[0108] Step S302 is an example of a transmission process and a reception process. In step S302, the transmission unit 32 of the server device 3B transmits the global model c to each of the client devices 1B-k. The reception unit 10 of each of the client devices 1B-k receives the global model c.

[0109] An example of the global model c generated / transmitted / received in steps S301 and S302 has been described with reference to equation (1) and FIG.

[0110] Step S101 is an example of a weighting coefficient learning process. In step S101, the weighting coefficient learning unit 13 of each client device 1B-k updates the weighting coefficient α^k_(i, b) through machine learning using the integrated model c_k. The integrated model c_k used in step S101 is as described with reference to equation (2) and FIG. 12.

[0111] Step S102 is an example of a local model learning process. In step S102, the local model learning unit 11 of each client device 1B-k trains a non-branched local model w_k by machine learning based on the branched global model c and the updated weighting coefficient α^k_(i, b).

[0112] More specifically, in step S102, the local model learning unit 11 may train the local model w_k using the regularization term described above. An example of training the local model w_k using the regularization term is as described with reference to Equation (3) and FIG. 12 .

[0113] Alternatively, instead of using a regularization term, the local model learning unit 11 may use the above-mentioned integrated model c_k to which the updated weighting coefficient α^k_(i, b) has been applied as the initial model of the local model w_k, and train the local model w_k to reduce the cross-entropy.

[0114] Step S103 is an example of a transmission process and a reception process. In step S103, the transmission unit 12 of each client device 1B-k transmits the trained local model w_k, the updated weighting coefficient α^k_(i, b), and the number n_k of training data to the server device 3B. The reception unit 30 of the server device 3B receives this information from each client device 1B-k.

[0115] Step S303 is an example of a global model update process. In step S303, the global model update unit 31 of the server device 3B integrates the i-th layer in each of the multiple local models w_k for each branch C_(i,b) based on the weighting coefficient α^k_(i,b) of each of the multiple client devices 1B-k and the number n_k of training data. As a result, the global model update unit 31 sets the integrated i-th layer for each branch C_(i,b) as the updated branch C_(i,b). The branch C_(i,b) updated in step S303 is as described with reference to equation (4) and FIG. 13.

[0116] In step S304, the control unit 310 of the server device 3B determines whether or not to terminate the federated learning. Whether or not to terminate may be determined, for example, based on whether or not the number of repetitions of steps S302 to S303 exceeds a threshold, or whether or not the accuracy of the local model w_k in each client device 1B-k exceeds a threshold, but is not limited to these. Furthermore, when the server device 3B determines to terminate the federated learning, it may notify each client device 1B-k of the termination, and each client device 1B-k that receives the notification may determine to terminate the federated learning.

[0117] If the answer to step S304 is No, the server device 3B repeats the process from step S302. If the answer to step S304 is Yes, the server device 3B ends the process.

[0118] In step S104, the control unit 110 of each client device 1B-k determines whether or not to terminate the federated learning. Whether or not to terminate may be determined, for example, based on whether or not the number of repetitions of steps S101 to S103 exceeds a threshold, or whether or not the accuracy of the local model w_k exceeds a threshold, but is not limited to this. Furthermore, when each client device 1B-k determines that it has terminated, it may notify the server device 3B of the termination, and the server device 3B, upon receiving the notification, may determine to terminate the federated learning.

[0119] If the answer to step S104 is No, the client device 1B-k repeats the process from step S101, whereas if the answer to step S104 is Yes, the next step S105 is executed.

[0120] In step S105, the control unit 110 of the client device 1B-k determines a model to be used as the personalized model. For example, the control unit 110 may apply the local model w_k trained in the most recent step S102 as the personalized model. In this case, the transmission process in step S103 after the most recent step S102 is executed may be omitted. Furthermore, for example, the control unit 110 may apply the integrated model c_k generated using the weighting coefficient α^k_(i, b) updated in the most recent step S101 as the personalized model. In this case, the local model learning process in step S102 after the most recent step S101 is executed and the transmission process in step S103 may be omitted.

[0121] (Effects of Information Processing System 100B) As described above, in the information processing system 100B, each of the multiple client devices 1B-k further includes a weighting coefficient learning unit 13 that updates the weighting coefficients α^k_(i,b) through machine learning using an integrated model c_k in which multiple branches C_(i,b) in a global model c are integrated using weighting coefficients α^k_(i,b). The local model learning unit 11 trains the local model w_k through machine learning based on the global model c and the updated weighting coefficients α^k_(i,b). The global model update unit 31 updates the global model c by integrating the i-th layer in the multiple local models w_k (e.g., each of them) for each branch C_(i,b) based on the weighting coefficients α^k_(i,b) provided from the multiple client devices 1B-k (e.g., each of them). Therefore, the information processing system 100B provides the following effects in addition to the effects provided by the information processing system 100. That is, it is possible to perform federated learning that takes into account the degree of importance that each client device 1B-k places on each of the multiple branches C_(i, b), without increasing the computational cost of each client device 1B-k. As a result, each client device 1B-k can obtain a personalized model that accurately matches that device.

[0122] As a comparative example to the information processing system 100B, a server device may provide multiple global models, and a client device may train a local model using weighting coefficients for each of the multiple global models. However, since the degree to which each client device emphasizes each global model may differ for each layer, it is not necessarily possible to assign appropriate weighting coefficients to each global model. In contrast, the information processing system 100B can consider the degree to which each client device emphasizes each branch for each layer, making it possible to obtain a personalized model that more accurately matches the device than the comparative example.

[0123] Furthermore, in the information processing system 100B, the local model learning unit 11 integrates the differences between multiple branches C_(i,b) (for example, each of them) and the i-th layer of the local model w_k using the updated weighting coefficient α^k_(i,b), and trains the local model w_k using the integrated difference as a regularization term. Therefore, the information processing system 100B achieves the following effect in addition to the effects achieved by the information processing system 100. That is, the local model w_k in each client device 1B-k is trained so that the difference with the branch C_(i,b) that the device places more importance on is smaller. As a result, each client device 1B-k can obtain a personalized model that accurately matches the device.

[0124] Furthermore, the information processing system 100B employs a configuration in which the local model learning unit 11 trains the local model w_k by using the integrated model c_k to which the updated weighting coefficient α^k_(i, b) has been applied as an initial model of the local model w_k. Therefore, the information processing system 100B achieves the following effect in addition to the effects achieved by the information processing system 100. That is, compared to when a regularization term is used, the calculation cost when each client device 1B-k performs machine learning of the local model w_k can be reduced.

[0125] Furthermore, in the information processing system 100B, the local model learning unit 11 trains a local model w_k using multiple training data, and the global model update unit 31 updates the global model c by integrating the i-th layer in multiple local models w_k (e.g., each of them) for each branch C_(i, b) based on the number n_k of training data provided from multiple client devices 1B-k (e.g., each of them). Therefore, the information processing system 100B achieves the following effect in addition to the effect achieved by the information processing system 100. That is, the more training data a local model w_k trained using, the greater the degree to which it is reflected in the global model c. As a result, the accuracy of the global model updated during the associative learning process can be improved.

[0126] [Application Examples] Application examples of the information processing systems 100 and 100B according to each exemplary embodiment will be described below. The information processing systems 100 and 100B can be applied to a variety of industries, and several examples will be described below. However, these examples do not limit the exemplary embodiments, and the information processing systems 100 and 100B can, of course, be applied to other industries. Furthermore, the information processing systems 100 and 100B can also be applied across several industries.

[0127] (Example 1: Autonomous Driving Assistance) As an example, the information processing system 100 or 100B may be applied to the field of autonomous driving assistance.

[0128] For example, the information processing system 100 or 100B can be applied when multiple companies in different locations jointly create a model that supports autonomous driving based on images from an in-vehicle camera. For example, multiple client devices 1 or 1B-k are located in each company. Furthermore, the server device 3 or 3B is located in an organization that provides a model generation service using federated learning. In the following, for ease of explanation, "client device 1 or 1B-k owned by each company" will simply be referred to as "each company."

[0129] The server device 3 or 3B creates a global model in which a predetermined layer branches into multiple branches as a global model that takes images from an onboard camera as input and outputs control signals for autonomous driving, and distributes the created global model to each company.

[0130] Here, for example, the multiple companies in different locations include a company located in a snowy region, a company located in a coastal area, and a company located in a mountainous area. The company located in the snowy region has "in-vehicle camera images in the snowy region" as training data. The company located in the coastal area has "in-vehicle camera images in the coastal area" as training data. The company located in the mountainous area has "in-vehicle camera images in the mountainous area" as training data.

[0131] Each company uses its own training data to train a non-branched local model by machine learning based on the branched global model distributed from the server device 3 or 3B. The server device 3 or 3B also updates the branched global model based on the trained local model provided by each company. Each company and the server device 3 or 3B repeat the training of the local model and the updating of the global model multiple times.

[0132] This allows each company (for example, a company located in a snowy region) to obtain a personalized model that appropriately reflects knowledge from locations different from the company's own (coastal areas and mountainous areas) and that accurately matches the company's own location, without increasing calculation costs. Furthermore, by obtaining control signals for autonomous driving from the personalized model, user decision-making can be supported.

[0133] (Example 2: Finance-related) The information processing system 100 or 100B according to each exemplary embodiment may be applied to the finance-related field, for example.

[0134] For example, the information processing system 100 or 100B can be applied when multiple branches of a bank jointly create a model that predicts loan default risk based on borrower characteristics (loans, business status). For example, multiple client devices 1 or 1B-k are deployed at each branch. Furthermore, the server device 3 or 3B is deployed in an organization that provides a model generation service using federated learning. In the following, for ease of explanation, the "client device 1 or 1B-k owned by each branch" will be simply referred to as "each branch." It is also possible to apply each bank affiliate instead of each bank branch. In this case, the following explanation will be the same if "branch" is read as "affiliate."

[0135] The server device 3 or 3B creates a global model in which a specified layer is branched into multiple branches as a global model that predicts loan default risk based on the borrower's characteristics (loan amount, business situation), and distributes the created global model to each company.

[0136] Here, for example, each store has a combination of its own borrower's characteristic data and bad debt history as training data. Each store uses its own training data to train a non-branched local model through machine learning based on the branched global model distributed from the server device 3 or 3B. In addition, the server device 3 or 3B updates the branched global model based on the trained local model provided by each store. Each store and the server device 3 or 3B repeat the training of the local model and the updating of the global model multiple times.

[0137] This allows each store to obtain a personalized model that appropriately reflects knowledge about borrowers at other stores and accurately matches the borrowers at that store, without increasing computational costs. In addition, obtaining default risk from the personalized model can support user decision-making.

[0138] As another example of application to the financial field, information processing system 100 or 100B can also be applied when multiple insurance company branches jointly create a model to predict insurance premiums based on customer health data such as medical history, age, blood pressure, genes, etc. In this case, in the explanation of the example of the "model to predict default risk" above, the explanation can be similarly given by replacing "borrower," "characteristic data," and "default (risk)" with "customer," "health data," and "insurance premium," respectively.

[0139] This allows each store to obtain a personalized model that appropriately reflects knowledge about customers of other stores and accurately matches the customers of that store without increasing calculation costs. Furthermore, obtaining insurance premiums from the personalized model can support user decision-making.

[0140] It is also possible to apply each of a plurality of insurance companies instead of each store of an insurance company. In that case, the same explanation can be given by further replacing "store" with "insurance company."

[0141] (Example 3: Medical and Healthcare Related) The information processing system 100 or 100B according to this exemplary embodiment may be applied to the medical field, for example.

[0142] For example, the information processing system 100 or 100B can also be applied when each clinic jointly creates a model that presents the cause or treatment of a disease based on symptom data of a patient recorded in a medical record, etc. In this case, in the explanation of the example of the "model for predicting default risk" above, the explanation can be similarly given by replacing "store," "borrower," "characteristic data," and "default (risk)" with "clinic," "patient," "symptom data," and "cause or treatment of disease," respectively.

[0143] This allows each clinic to obtain a personalized model that appropriately reflects knowledge about patients at other clinics and that accurately matches the patients at that clinic without increasing computational costs.Furthermore, by obtaining the cause or treatment of a disease from the personalized model, it is possible to support user decision-making.

[0144] As another example of application in the medical and healthcare fields, the information processing system 100 or 100B can also be applied when multiple pharmaceutical companies jointly create a model that predicts the activity of a compound from data such as the structure of the compound (ligand), the structure of a protein, etc. In this case, in the explanation of the example of the "model that predicts the risk of default" above, the same explanation can be obtained by replacing "store," "borrower characteristic data," and "default (risk)" with "pharmaceutical company," "data such as the structure of the compound (ligand) and the structure of the protein," and "activity of the compound," respectively.

[0145] This allows each pharmaceutical company to obtain a personalized model that appropriately reflects knowledge about compounds targeted by other pharmaceutical companies and that accurately matches the compounds targeted by that pharmaceutical company, without increasing computational costs. Furthermore, obtaining compound activity from the personalized model can assist users in making decisions.

[0146] (Example 4: Machine-related) The information processing system 100 or 100B according to this exemplary embodiment may be applied to a machine-related field, for example.

[0147] For example, the information processing system 100 or 100B can also be applied when each factory jointly creates a model for controlling the operation of a robot based on factory status data such as the manufacturing status, transportation status, etc. In this case, in the explanation of the example of the "model for predicting default risk" above, the same explanation can be obtained by replacing "store," "borrower characteristic data," and "default risk" with "factory," "factory status data," and "control signal for controlling the operation of the robot," respectively.

[0148] This allows each factory to obtain a personalized model that appropriately reflects knowledge about the situations of other factories and that accurately matches the situation of that factory, without increasing calculation costs. Furthermore, by obtaining a "control signal for controlling the operation of a robot" from the personalized model, user decision-making can be supported.

[0149] As another example of application to machinery-related fields, the information processing system 100 or 100B can also be applied when multiple companies that manufacture transportation equipment jointly create a model to control the transportation equipment (cars, airplanes, ships, etc.) based on data on scenery, equipment measurement status, and congestion. In this case, in the explanation of the example of the "model to predict default risk" above, the explanation can be similarly given by replacing "store," "borrower characteristic data," and "default risk" with "company," "data on scenery, equipment measurement status, and congestion status," and "control signal to control the transportation equipment," respectively.

[0150] This allows each company to obtain a personalized model that appropriately reflects the knowledge of other companies about the transportation equipment they target and that accurately matches the transportation equipment they target, without increasing computational costs. Furthermore, by obtaining a "control signal for controlling the transportation equipment" from the personalized model, it is possible to support user decision-making.

[0151] As another example of application to the machinery-related field, the information processing system 100 or 100B can also be applied when multiple logistics companies jointly create a model for calculating a transportation route from data on the status of transported goods and the status of transportation equipment. In this case, in the explanation of the example of the "model for predicting default risk" above, the same explanation can be obtained by replacing "store," "borrower characteristic data," and "default risk" with "logistics company," "data on the status of transported goods and the status of transportation equipment," and "transportation route," respectively.

[0152] This allows each logistics company to obtain a personalized model that accurately matches the goods or transportation equipment targeted by the logistics company, and that appropriately reflects the knowledge of other logistics companies regarding the goods or transportation equipment targeted by the logistics company, without increasing computational costs. Furthermore, obtaining a "transportation route" from the personalized model can assist users in making decisions.

[0153] (Example 5: Court-Related) The information processing system 100 or 100B according to this exemplary embodiment may be applied to the court-related field, for example.

[0154] For example, the information processing system 100 or 100B can also be applied when multiple courts jointly create a model that predicts sentencing, etc., based on data that forms the basis of the trial, such as the circumstances of the crime, etc., in the trial, the circumstances of the evidence, laws, court precedents, etc. In this case, in the explanation of the example of the "model that predicts default risk" above, the explanation can be similarly given by replacing "store," "borrower characteristic data," and "default risk" with "court," "data that forms the basis of the trial," and "sentencing, etc."

[0155] This allows each court to obtain a personalized model that appropriately reflects knowledge about trials in other courts and accurately matches the trials in question without increasing computational costs. Furthermore, obtaining "sentencing, etc." from the personalized model can assist users in making decisions.

[0156] (Notes on each exemplary embodiment) The configurations described in each exemplary embodiment are not limited to the examples described above. The following configurations may be included to resolve some secondary issues that may arise when implementing federated learning in practice.

[0157] For example, when transmitting model parameters from each client device to a server device, it is preferable to have a configuration that ensures the confidentiality of the model parameters. For example, the information processing system described in each exemplary embodiment may have a configuration related to homomorphic encryption or the like that allows calculations to be performed while keeping the model parameters confidential, so that the confidentiality of the model parameters themselves can be ensured.

[0158] As an example, each client device (1, 1A, 1B-k) may be configured to include an encryption unit that encrypts model parameters of a local model after training using homomorphic encryption or the like, and a global model update unit 31 of the server device (3, 3A, 3B) may update the global model while keeping the encrypted model parameters confidential. Alternatively, the server device (3, 3A, 3B) may be configured to include an encryption unit that encrypts model parameters of the global model using homomorphic encryption or the like, and the model parameters are decrypted in each client device (1, 1A, 1B-k).

[0159] Furthermore, it is preferable to have a configuration that can minimize the data size of model parameters when they are transmitted from each client device to the server device. For example, each client device (1, 1A, 1B-k) and the server device (3, 3A, 3B) may be configured to compress the model parameters. Furthermore, at least one of each client device (1, 1A, 1B-k) and the server device (3, 3A, 3B) may be configured to include a model reconfiguration unit that reduces the size of the model by reconfiguring the model (generating a distilled model, a derived model, a pseudo model, or a higher-level model). For example, such a configuration is suitable in situations where the processing performance of the client device is limited.

[0160] Furthermore, as an example, when each client device (1, 1A, 1B-k) is implemented as a wearable device, the client device (1, 1A, 1B-k) may be configured to control the transmission of model parameters to the server device (3, 3A, 3B) according to the remaining battery charge. As an example, when the remaining battery charge is below a predetermined level, only model parameters that have changed from their previous values ​​by a predetermined value (percentage) or more may be transmitted to the server device (3, 3A, 3B). Alternatively, the client device (1, 1A, 1B-k) may be configured to apply a sampling process to the acquired data (for example, randomly sampling to extract only 10%) and perform a learning process using the local model learning unit 11 using only the sampled data. Alternatively, the client device (1, 1A, 1B-k) may be configured to store the acquired data in another device (for example, the server device 3, 3A, 3B).

[0161] Furthermore, since there is a possibility that noise may be present in each learning data and each parameter, each client device (1, 1A, 1B-k) and server device (3, 3A, 3B) may have a configuration for improving resistance to the noise. As an example, each client device (1, 1A, 1B-k) and server device (3, 3A, 3B) may have a configuration for performing error correction on the learning data or model parameters.

[0162] [Example of implementation by software] Some or all of the functions of the client devices 1, 1A, 1B-k, 2 and the server devices 3, 3A, 3B (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as an integrated circuit (IC chip), or by software.

[0163] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 15. Figure 15 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.

[0164] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.

[0165] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0166] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0167] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0168] [Appendix 1] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0169] (Supplementary Note 1) An information processing system comprising: a plurality of client devices equipped with a local model learning means that trains a local model by machine learning based on a global model provided from a server device, the global model having a predetermined layer among multiple layers branched into a plurality of branches; and the server device equipped with a global model update means that updates the global model based on the local models provided from the plurality of client devices.

[0170] (Supplementary Note 2) The information processing system according to Supplementary Note 1, wherein the predetermined layer of the global model includes the plurality of branches whose outputs may differ from each other in response to inputs from a previous layer, and is branched so that outputs from the plurality of branches are superimposed and output to a subsequent layer.

[0171] (Supplementary Note 3) The information processing system described in Supplementary Note 1 or 2, wherein the plurality of client devices further comprise a weight coefficient learning means for updating the weight coefficients through machine learning using an integrated model in which the plurality of branches in the global model are integrated using weight coefficients; the local model learning means trains the local model through machine learning based on the global model and the updated weight coefficients; and the global model updating means updates the global model by integrating the predetermined layers in the plurality of local models for each branch based on the weight coefficients provided from the plurality of client devices.

[0172] (Supplementary Note 4) The information processing system according to Supplementary Note 3, wherein the local model learning means trains the local model using the integrated model to which the updated weighting coefficients have been applied as an initial model of the local model.

[0173] (Supplementary Note 5) The information processing system according to Supplementary Note 3, wherein the local model learning means integrates differences between the plurality of branches and the predetermined layer of the local model using the updated weighting coefficients, and trains the local model using the integrated difference as a regularization term.

[0174] (Supplementary Note 6) The information processing system described in any one of Supplementary Notes 3 to 5, wherein the local model learning means trains the local model using a plurality of training data, and the global model update means updates the global model by integrating the predetermined layers in the plurality of local models for each branch based on the number of the training data provided from the plurality of client devices.

[0175] (Supplementary Note 7) A client device comprising: a receiving means for receiving from a server device a global model in which a predetermined layer among a plurality of layers is branched into a plurality of branches; a local model learning means for training a local model in which the predetermined layer is not branched by machine learning based on the global model; and a transmitting means for transmitting the local model to the server device.

[0176] (Supplementary Note 8) A server device comprising: a receiving means for receiving, from a plurality of client devices, a local model trained by machine learning based on a global model in which a predetermined layer among a plurality of layers branches into a plurality of branches, the predetermined layer not branching; a global model updating means for updating the global model based on the local model; and a transmitting means for transmitting the global model to the plurality of client devices.

[0177] (Supplementary Note 9) A client device comprising: an input data acquisition means for acquiring input data; and an inference means for performing inference on the input data using a local model in which a predetermined layer among multiple layers is not branched, the local model being trained by machine learning based on a global model provided by a server device, the global model having the predetermined layer branched into multiple branches.

[0178] (Supplementary Note 10) An information processing method including: a local model learning process in which at least one processor included in a plurality of client devices trains a local model by machine learning, based on a global model provided from a server device, in which a predetermined layer among multiple layers is branched into a plurality of branches; and a global model update process in which at least one processor included in the server device updates the global model based on the local models provided from the plurality of client devices.

[0179] (Supplementary Note 11) An information processing method including: a receiving process in which at least one processor included in a client device receives from a server device a global model in which a predetermined layer among a plurality of layers is branched into a plurality of branches; a local model learning process in which the at least one processor trains a local model in which the predetermined layer is not branched by machine learning based on the global model; and a transmitting process in which the at least one processor transmits the local model to the server device.

[0180] (Supplementary Note 12) An information processing method including: a receiving process in which at least one processor included in a server device receives, from a plurality of client devices, a local model trained by machine learning based on a global model in which a predetermined layer among a plurality of layers is branched into a plurality of branches, and the predetermined layer is not branched; a global model updating process in which the at least one processor updates the global model based on the local model; and a transmitting process in which the at least one processor transmits the global model to the plurality of client devices.

[0181] (Supplementary Note 13) An information processing method including: an input data acquisition process in which at least one processor included in a client device acquires input data; and an inference process in which the at least one processor performs inference on the input data using a local model in which a predetermined layer among multiple layers is not branched, the local model being trained by machine learning based on a global model in which the predetermined layer is branched into multiple branches, the local model being provided by a server device.

[0182] (Supplementary Note 14) A program for causing a computer to function as the client device according to Supplementary Note 7, the program causing a computer to function as the receiving means, the local model learning means, and the transmitting means.

[0183] (Supplementary Note 15) A program for causing a computer to function as the server device according to Supplementary Note 8, the program causing a computer to function as the receiving means, the global model updating means, and the transmitting means.

[0184] (Supplementary Note 16) A program for causing a computer to function as the client device according to Supplementary Note 9, the program causing the computer to function as the input data acquisition means and the inference means.

[0185] (Supplementary Note 17) An information processing system comprising: a plurality of client devices each having at least one processor; and a server device having at least one processor; wherein the at least one processor included in each of the plurality of client devices executes a local model learning process to train a local model by machine learning, based on a global model provided by the server device, in which a predetermined layer among multiple layers is branched into a plurality of branches; and the at least one processor included in the server device executes a global model update process to update the global model, based on the local model provided by each of the plurality of client devices.

[0186] Each of the plurality of client devices may further include a memory. The memory may store a program for causing the at least one processor to execute the local model learning process. The server device may further include a memory. The memory may store a program for causing the at least one processor to execute the global model update process.

[0187] 100, 100B Information processing system 1, 1A, 1B, 2 Client device 3, 3A, 3B Server device 10, 30 Receiving unit 11 Local model learning unit 12, 32 Transmitting unit 13 Weighting coefficient learning unit 21 Input data acquisition unit 22 Inference unit 31 Global model update unit 33 Global model generation unit 110, 310 Control unit 120, 320 Storage unit 130, 330 Communication unit C1 Processor C2 Memory

Claims

1. A plurality of client devices including local model learning means for training a local model in which a predetermined layer among a plurality of layers is not branched, based on a global model provided from a server device and in which a predetermined layer among the plurality of layers branches into a plurality of branches, by machine learning; and the server device including global model update means for updating the global model based on the local models provided from the plurality of client devices. An information processing system comprising:

2. The information processing system according to claim 1, wherein the predetermined layer of the global model includes the plurality of branches whose outputs can be different from each other for an input from a previous layer, and branches so as to output by overlapping the outputs from the plurality of branches to a subsequent layer.

3. The plurality of client devices further include weight coefficient learning means for updating the weight coefficient by machine learning using an integrated model in which the plurality of branches in the global model are integrated by a weight coefficient; the local model learning means trains the local model by machine learning based on the global model and the updated weight coefficient; and the global model update means updates the global model by integrating the predetermined layers in the plurality of local models for each of the branches based on the weight coefficients provided from the plurality of client devices. The information processing system according to claim 1 or 2.

4. The information processing system according to claim 3, wherein the local model learning means trains the local model using the integrated model to which the updated weight coefficient is applied as an initial model of the local model.

5. The information processing system according to claim 3, wherein the local model learning means integrates a difference between the plurality of branches and the predetermined layer of the local model by the updated weight coefficient, and trains the local model using the integrated difference as a regularization term.

6. The local model learning means trains the local model using a plurality of training data, and the global model updating means updates the global model by integrating the predetermined layers in the plurality of local models for each branch based on the number of the training data provided from the plurality of client devices. The information processing system according to any one of claims 3 to 5.

7. A receiving means for receiving from a server device a global model in which a predetermined layer among a plurality of layers branches into a plurality of branches, a local model learning means for training, by machine learning, a local model in which the predetermined layer does not branch based on the global model, and a transmitting means for transmitting the local model to the server device. A client device comprising the above.

8. A local model trained by machine learning based on a global model in which a predetermined layer among a plurality of layers branches into a plurality of branches, the local model in which the predetermined layer does not branch, a receiving means for receiving from a plurality of client devices, a global model updating means for updating the global model based on the local model, and a transmitting means for transmitting the global model to the plurality of client devices. A server device comprising the above.

9. An input data acquisition means for acquiring input data, and an inference means for performing an inference on the input data using a local model that is a global model provided from a server device and is trained by machine learning based on a global model in which a predetermined layer among a plurality of layers branches into a plurality of branches and the predetermined layer does not branch. A client device comprising the above.

10. An information processing method including a local model learning process in which at least one processor included in a plurality of client devices trains, by machine learning, a local model in which a predetermined layer does not branch based on a global model provided from a server device and in which a predetermined layer among a plurality of layers branches into a plurality of branches, and a global model updating process in which at least one processor included in the server device updates the global model based on the local model provided from the plurality of client devices.

11. An information processing method including: a receiving process in which at least one processor included in a client device receives from a server device a global model in which a predetermined layer among a plurality of layers branches into a plurality of branches; a local model learning process in which the at least one processor trains, by machine learning, a local model in which the predetermined layer does not branch, based on the global model; and a transmitting process in which the at least one processor transmits the local model to the server device.

12. An information processing method including: a receiving process in which at least one processor included in a server device receives from a plurality of client devices a local model that is trained by machine learning based on a global model in which a predetermined layer among a plurality of layers branches into a plurality of branches, and in which the predetermined layer does not branch; a global model updating process in which the at least one processor updates the global model based on the local model; and a transmitting process in which the at least one processor transmits the global model to the plurality of client devices.

13. An information processing method including: an input data acquisition process in which at least one processor included in a client device acquires input data; and an inference process in which the at least one processor performs inference on the input data using a local model that is trained by machine learning based on a global model provided from a server device, in which a predetermined layer among a plurality of layers branches into a plurality of branches, and in which the predetermined layer does not branch.

14. A program for causing a computer to function as the client device according to claim 7, the program for causing a computer to function as the receiving means, the local model learning means, and the transmitting means.

15. A program for causing a computer to function as the server device according to claim 8, the program for causing a computer to function as the receiving means, the global model updating means, and the transmitting means.

16. A program for causing a computer to function as the client device according to claim 9, the program for causing a computer to function as the input data acquisition means and the inference means.

Citation Information

Patent Citations

  • Dynamic weight updating for neural network

    JP2022171603A

  • Server device

    JP2023179169A

  • Cooperative neural networks with spatial confinement constraints

    JP2023521695A

  • Federated learning for automated selection of high band mm wave sectors

    WO2023164208A1