Computing System and Computing Method

The computing system addresses the risk of inferring compound data in federated learning by anonymizing model parameters and integrating them through secret calculations, thereby enhancing security and maintaining process accuracy.

JP7697582B2Active Publication Date: 2025-06-24NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024505753
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-10
Publication Date
2025-06-24
Estimated Expiration
2042-03-10

AI Technical Summary

Technical Problem

There is a risk that malicious users can infer compound data used in federated learning by obtaining the parameters of machine learning models.

Method used

A computing system and method that includes a confidentiality means for generating a model from compound data at client terminals and anonymizing the model parameters, followed by secret calculation to integrate the models using these anonymized parameters.

Benefits of technology

The solution effectively reduces the risk of inferring compound data used in federated learning, while maintaining the accuracy and efficiency of the machine learning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697582000001
    Figure 0007697582000001
  • Figure 0007697582000002
    Figure 0007697582000002
  • Figure 0007697582000003
    Figure 0007697582000003
Patent Text Reader

Abstract

Provided are a computation system and a computation method which reduce the risk that compound data used in federated learning will be inferred. A computation system (1) comprises: a concealment unit (11) that carries out a first process in which after respective models are generated from compound data at a plurality of client terminals, parameters of the models are concealed; and a secure computation unit (12) that uses the concealed parameters to carry out a secure computation for integrating the models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a computing system and a computing method.

Background Art

[0002] In recent years, in the fields of drug discovery and chemistry, in order to reduce development costs, it has been expected to link the structural data of compounds held by multiple organizations. Therefore, the use of federated learning, which performs machine learning locally and integrates machine learning models on the server side, has been expected.

[0003] Note that Patent Document 1 discloses a secure computing system that can perform calculations while keeping data encrypted.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] By the way, it has been pointed out that there is a risk that a malicious user may obtain the parameters of a machine learning model and infer the compound data used in the machine learning.

[0006] Therefore, one of the objects to be achieved by the embodiments disclosed in this specification is to provide a computing system and a computing method capable of reducing the risk that the compound data used in federated learning is inferred.

Means for Solving the Problems

[0007] The computing system according to the first aspect of the present disclosure a confidentiality means for performing a first process of generating a model from a set of compound data at each of a plurality of client terminals and then anonymizing the parameters of the model; Secret calculation means for performing secret calculation to integrate the model using the anonymized parameter, is provided.

[0008] In the calculation method according to the second aspect of the present disclosure, after generating a model from a set of compound data on each of a plurality of client terminals, a first process of anonymizing the parameters of the model is performed, secret calculation is performed to integrate the model using the anonymized parameter.

Advantages of the Invention

[0009] According to the present disclosure, a calculation system and a calculation method capable of reducing the risk of speculation of compound data used in federated learning can be provided.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Mode for Carrying Out the Invention

[0011] <Background Leading to the Embodiment> First, an overview of federated learning will be described. FIG. 1 is a block diagram showing the functional configuration of a related computing system 1. The computing system 1 includes client terminals 2a, 2b, and 2c and a computing server 3.

[0012] The client terminal 2a generates a machine learning model (referred to as a local model a) from the data owned by organization A. The client terminal 2a transmits the parameters of the local model a to the computing server 3.

[0013] The client terminal 2b generates a machine learning model (referred to as a local model b) from the data owned by organization B. The client terminal 2b transmits the parameters of the local model b to the computing server 3.

[0014] The client terminal 2c generates a machine learning model (referred to as a local model c) from the data owned by organization C. The client terminal 2c transmits the parameters of the local model c to the computing server 3.

[0015] The computing server 3 generates a global model by integrating the local models a, b, and c. The computing server 3 may generate a global model, for example, by taking the arithmetic mean of the parameters. Note that the method of integrating the parameters is not limited to the arithmetic mean. The computing server 3 transmits the global model to the client terminals 2a, 2b, and 2c.

[0016] According to the computing system 1, there is a problem that the parameters of the local model a, the parameters of the local model b, and the parameters of the local model c are aggregated in one computing server 3, and the risk of information leakage is high. Based on the above considerations, the inventor of the present application came up with the invention according to Embodiment 1.

[0017] <Embodiment 1> FIG. 2 is a schematic diagram showing an example of the configuration of the computing system 10 according to Embodiment 1. The computing system 10 includes client terminals 20a, 20b, and 20c and a computing server group 30. Each client terminal is a terminal of an organization (for example, a pharmaceutical company or a chemical company) that uses the computing system 1. The computing server group 30 includes computing servers 31_1, 31_2, and 31_3.

[0018] The client terminals 20a, 20b, and 20c and the computing server group 30 are communicably connected via a network (not shown). The network may be wired or wireless. The network may be, for example, a VPN (Virtual Private Network).

[0019] Hereinafter, when the client terminals 20a, 20b, and 20c are not distinguished from each other, they may simply be referred to as the client terminal 20. Note that the number of client terminals 20 is not limited to three, and may be two or four or more. Similarly, when the computing servers 31_1, 31_2, and 31_3 are not distinguished from each other, they may simply be referred to as the computing server 31. The number of computing servers 31 is not limited to three, and may be two or four or more. In FIG. 2, the number of client terminals 20 and the number of computing servers 31 match, but they do not have to match.

[0020] Next, the client terminal 20 will be described in detail with reference to FIG. 3. The client terminal 20 includes a model generation unit 21, an anonymization unit 22, an acquisition unit 23, and a prediction unit 24.

[0021] The model generation unit 21 generates a local model from a set of compound data within the self-organization. The local model is also referred to as a local artificial intelligence (AI) model. The model generation unit 21 may use the set of compound data as training data. The set of compound data includes a plurality of items, for example, an item related to the structure of the compound and an item related to the properties of the compound. The structure of the compound is represented by, for example, a fixed-length bit string. Each bit of the bit string represents the presence or absence of a predetermined structure (for example, a benzene ring). The properties are represented by property values (for example, the value of tensile strength). The property value may be a value obtained through experiments or a value obtained through simulations or theoretical calculations. Since machine learning is performed on the client terminal 20, the compound data within the self-organization does not go outside.

[0022] A set of compound data typically includes items related to the purpose for which the compound is used (such as headache medicine, abdominal pain medicine, etc.), items related to the structure and composition of the compound, and items related to theoretical calculations and simulation results (such as simulation results of properties). The set of compound data further includes items related to the manufacturing process of the compound, items of data for materials informatics (also referred to as data for machine learning), items related to the functions and properties of the compound, and so on.

[0023] The anonymization unit 22 divides each parameter of the local model into a plurality of shares and transmits the plurality of shares to the computing server group 30. Since the original parameter cannot be restored from a single share, it can be said that the client terminal 2 anonymizes the parameter.

[0024] The acquisition unit 23 acquires the global model from the calculation results of the computing server group 30. The acquisition unit 23 acquires the global model by combining the calculation results of the computing server 31_1, the computing server 31_2, and the computing server 31_3.

[0025] The prediction unit 24 predicts the properties, structure, etc. of the compound using the global model. The prediction unit 24 may, for example, predict the properties from the structure of the compound using the global model. Also, the prediction unit 24 may predict the structure from the properties of the compound using the global model. The prediction unit 24 may output the prediction result to a display, a monitor (not shown), etc. The prediction unit 24 can predict the properties of the compound with high accuracy by using the global model.

[0026] Note that the client terminal 20 includes a processor, a memory, and a storage device as a configuration not shown. The processor causes the memory to read a computer program from the storage device and executes the computer program. Thereby, the processor realizes the functions of the model generation unit 21, the anonymization unit 22, the acquisition unit 23, and the prediction unit 24.

[0027] Next, with reference to FIG. 4, the functions of the calculation server 31 will be described in detail. The calculation server 31 includes a shared storage unit 311 and a secure calculation unit 312.

[0028] The shared storage unit 311 is a storage that stores the shares generated by the anonymization unit 22 of the client terminal 20. The three shares generated for one parameter are distributed and stored in the shared storage unit 311 of the calculation server 31_1, the shared storage unit 311 of the calculation server 31_2, and the shared storage unit 311 of the calculation server 31_3.

[0029] The secure calculation unit 312 performs secure calculations for integrating the models using the shares stored in the shared storage unit 311. The secure calculation unit 312 may perform model integration at a predetermined time. The local model parameters cannot be known from the shares, and the calculation using the shares can be said to be a secure calculation. The secure calculation units 312 of the calculation server 31_1, the secure calculation unit 312 of the calculation server 31_2, and the secure calculation unit 312 of the calculation server 31_3 may cooperate to perform multi-party calculation (MPC). The secure calculation unit 312 transmits the calculation result to the client terminal 20.

[0030] Note that the calculation server 31 also includes a processor, a memory, and a storage device as a configuration (not shown), similar to the client terminal 20. The processor causes the memory to read a computer program from the storage device and executes the computer program. Thereby, the processor realizes the functions of the secure calculation unit 312.

[0031] Next, with reference to FIGS. 5 to 8, the operation of the calculation system 10 will be specifically described. FIG. 5 is a diagram for explaining the processing performed by the anonymization unit 22 of the client terminal 20a. The anonymization unit 22 of the client terminal 20a divides the parameters of the local model into shares Sa1, Sa2, and Sa3. The anonymization unit 22 of the client terminal 20a transmits the share Sa1 to the calculation server 31_1, transmits the share Sa2 to the calculation server 31_2, and transmits the share Sa3 to the calculation server 31_3.

[0032] Note that the client terminal 20b also similarly transmits the share Sb1 to the calculation server 31_1, the share Sb2 to the calculation server 31_2, and the share Sb3 to the calculation server 31_3. The client terminal 20c also similarly transmits the share Sc1 to the calculation server 31_1, the share Sc2 to the calculation server 31_2, and the share Sc3 to the calculation server 31_3.

[0033] FIG. 6 is a diagram for explaining the shares stored in the share storage unit 311 of the calculation server 31_1. The share storage unit 311 of the calculation server 31_1 stores the share Sa1 received from the client terminal 20a, the share Sb1 received from the client terminal 20b, and the share Sc1 received from the client terminal 20c.

[0034] Note that the calculation server 31_2 also similarly stores the share Sa2, the share Sb2, and the share Sc2. The calculation server 31_3 also similarly stores the share Sa3, the share Sb3, and the share Sc3.

[0035] FIG. 7 is a diagram for explaining the processing performed by the secret calculation unit 312 of the calculation server 31_1. The secret calculation unit 312 of the calculation server 31_1 performs calculations for integrating the model using the shares Sa1, Sb1, and Sc1. The secret calculation unit 312 of the calculation server 31_1 transmits the calculation result g1 to the client terminals 20a, 20b, and 20c.

[0036] Note that the calculation server 31_2 also performs similar calculations using the shares Sa2, Sb2, and Sc2, and transmits the calculation result g2 to the client terminals 20a, 20b, and 20c. The calculation server 31_3 also performs similar calculations using the shares Sa3, Sb3, and Sc3, and transmits the calculation result g3 to the client terminals 20a, 20b, and 20c.

[0037] FIG. 8 is a diagram for explaining the processing performed by the acquisition unit 23 of the client terminal 20a. The acquisition unit 23 of the client terminal 20a calculates the parameters of the global model from the calculation results g1 of the calculation server 31_1, the calculation result g2 of the calculation server 31_2, and the calculation result g3 of the calculation server 31_3. The acquisition unit 23 may, for example, calculate the sum of g1, g2, and g3. Similarly, the client terminals 20b and 20c can also calculate the parameters of the global model. Note that any one of the calculation servers 31_1, 31_2, and 31_3 may calculate the parameters of the global model from g1, g2, and g3 and distribute them to the client terminals 20a, 20b, and 20c.

[0038] The calculation system 1 can periodically update the global model by repeating the processing of FIGS. 5 to 8. The client terminal 20 first updates the global model by performing machine learning using new compound data and generates a new local model. Next, the client terminal 20 secretly distributes the parameters of the new local model. Note that the client terminal 20 may secretly distribute the difference between the parameters of the local model and the parameters of the global model. Next, the calculation server group 30 executes secret calculation.

[0039] FIG. 9 is a block diagram showing the minimum functional configuration of the calculation system 1. The calculation system 1 includes an anonymization unit 11 and a secret calculation unit 12.

[0040] The anonymization unit 11 performs a first process of anonymizing the parameters of the model after generating the model from the set of compound data on each of the plurality of client terminals. The anonymization unit 22 of the client terminal 20 described above is a specific example of the anonymization unit 11. Note that when other servers are provided in addition to the calculation server group 30, the anonymization unit 11 may be provided in other servers. The anonymization unit 11 may anonymize the parameters of the local model by a method other than secret sharing (for example, a homomorphic encryption method).

[0041] The secret calculation unit 12 performs secret calculations for integrating the model using the encrypted parameters. The secret calculation units 312 of the calculation servers 31_1, 31_2, and 31_3 described above cooperate to function as the secret calculation unit 12. Further, the secret calculation unit 12 may perform secret calculations on data encrypted by the homomorphic encryption method. In such a case, the calculation system 1 may not include the calculation server group 30.

[0042] Next, the effects of the calculation system 1 will be described. In the calculation system 1, federated learning is performed while keeping the parameters of the local model encrypted. As a result, the risk of inferring the compound data used for learning within each organization from the parameters of the local model can be reduced.

[0043] In secret calculation, calculations can be executed while keeping the data encrypted, but there is a problem that the execution time of the calculations is long. However, since the amount of calculation required for integrating the local model is sufficiently small, it is considered that the calculation system 1 can execute secret calculations in a realistic time.

[0044] The inventors and applicants of the present application verified the accuracy and calculation time of the calculation system 1. The number of clients was set to 2, the secret calculation method was set to the secret sharing method, and the number of shares was set to 3. It was verified that the calculation system 1 can achieve the same inference accuracy in the same calculation time as the related art.

[0045] <Embodiment 2> In Embodiment 1, the global model generated by secret calculation is distributed to each organization, and each organization uses the global model to predict the properties of compounds and the like. Therefore, there remains a risk of inferring the compound data used for learning from the global model by the participating organizations in federated learning. Therefore, it is preferable not to perform federated learning using highly confidential data.

[0046] In addition, Embodiment 1 executes a process (first process) of anonymizing the parameters of the local model. However, secure computation has a problem of long execution time, and there may be cases where it is preferable to generate the global model without anonymization. For example, when the number of parameters of the local model is large, the execution time of secure computation may become long. Also, when integrating parameters by arithmetic mean, the execution time is considered to be short, but when integrating parameters by more complex calculations, the execution time is considered to become long. For example, when considering outliers of the parameters of the local model, complex calculations may be required.

[0047] As described above, the set of compound data used in machine learning may include a plurality of items. The plurality of items are, for example, purpose, structure, results of theoretical calculations, manufacturing process, materials informatics, characteristics, and the like. Among these, there are items with low confidentiality such as the results of theoretical calculations, and items with high confidentiality such as purpose, structure, and manufacturing process. Also, among these, there are items such as the results of theoretical calculations and data for materials informatics, which are considered to have a large amount of data and a large number of model parameters. In the calculation system according to Embodiment 2, the process to be applied to each item is selected from among a plurality of processes including the first process.

[0048] FIG. 10 is a block diagram showing the configuration of the calculation system 100 according to Embodiment 2. The calculation system 100 includes client terminals 200a, 200b, and 200c, a group of calculation servers 30, and a server 400. Comparing the calculation system 10 shown in FIG. 2 with the calculation system 100, the server 400 is added to the calculation system 100. Also, the client terminals 20a, 20b, and 20c are replaced with the client terminals 200a, 200b, and 200c. Also, the calculation servers 31_1, 31_2, and 31_3 are replaced with the calculation servers 32_1, 32_2, and 32_3.

[0049] Note that, similar to Embodiment 1, when the client terminals 200a, 200b, and 200c are not distinguished from each other, they may simply be referred to as the client terminal 200. When the computing servers 32_1, 32_2, and 32_3 are not distinguished from each other, they may simply be referred to as the computing server 32.

[0050] Next, the server 400 will be described in detail with reference to FIG. 11. The server 400 and the client terminal 200 are communicably connected via a network (not shown).

[0051] The server 400 includes a storage unit 410 and a computing unit 420. The storage unit 410 stores the data of each item received from the client terminal 200 (hereinafter also referred to as item data) and the parameters of the local model.

[0052] The computing unit 420 has a function of performing calculations using the item data and a function of integrating the parameters of the local model. The computing unit 420 executes calculations in a state where the item data and the parameters of the local model are not encrypted.

[0053] First, the function of performing calculations using the item data itself will be described. In Embodiment 1, a machine learning model was used to predict the properties of a compound, etc., but the computing unit 420 predicts the properties of a compound, etc. using the item data itself. For example, when predicting the properties of a compound having a certain structure, the properties of the compound can be predicted by calculating the average value of the properties of compounds having a similar structure. The calculations using the item data are not limited to the calculation of the average value, and complex calculations may be performed.

[0054] Next, the function of integrating the local models will be described. The computing unit 420 performs a process of integrating the parameters stored in the storage unit 410 at a predetermined timing (for example, once a day). Then, the computing unit 420 transmits the parameters of the global model to the client terminals 200a, 200b, and 200c.

[0055] Next, the calculation server 32 will be described with reference to FIG. 12. The calculation server 32 includes a shared memory unit 321 and a secure calculation unit 322. Comparing the calculation server 31 shown in FIG. 4 with the calculation server 32, the shared memory unit 311 is replaced by the shared memory unit 321, and the secure calculation unit 312 is replaced by the secure calculation unit 322.

[0056] The shared memory unit 321 stores the shares of the item data in addition to the shares of the parameters of the local model. The shared memory unit 321 may store the item data of a plurality of items. In such a case, it is not necessary for all items to be anonymized, and at least one item may be anonymized.

[0057] In addition to the function of performing secure calculation for integrating the models, the secure calculation unit 322 has a function of performing calculations using the shares of the item data stored in the shared memory unit 321. The secure calculation unit 322 executes secure calculation in response to a calculation request from the client terminal 200 and outputs the calculation result. The secure calculation units 322 of the calculation server 32_1, the secure calculation unit 322 of the calculation server 32_2, and the secure calculation unit 322 of the calculation server 32_3 may cooperate to perform multi-party calculation.

[0058] Next, the client terminal 200 will be described with reference to FIG. 13. Comparing the client terminal 20 shown in FIG. 3 with the client terminal 200, the anonymization unit 22 is replaced by the anonymization unit 220, the acquisition unit 23 is replaced by the acquisition unit 230, and the prediction unit 24 is replaced by the prediction unit 240. Also, a transmission unit 250 and a selection unit 260 are added.

[0059] In addition to the function of anonymizing the parameters of the local model, the anonymization unit 220 has a function of anonymizing item data. The acquisition unit 230 has a function of acquiring the global model from the calculation server group 30 and, in addition, has a function of acquiring the global model from the server 400. The prediction unit 240 has a function of predicting the properties of a compound and the like using the global model, and in addition, has a function of predicting the properties of a compound and the like using the item data stored in the server 400 or the calculation server group 30. The prediction unit 240 has a function of transmitting a calculation request to the server 400 or the calculation server group 30 and acquiring the calculation result.

[0060] The transmission unit 250 has a function of transmitting the item data and the parameters of the local model to the server 400 without anonymizing them.

[0061] The selection unit 260 selects a process to be applied to each item of the compound data set from among a first process, a second process, a third process, and a fourth process. In the first process, after generating a local model based on the data of each item, the parameters of the local model are secretly distributed. In the second process, after generating a local model based on the data of each item, the parameters of the local model are transmitted to the server 400 without anonymization. In the third process, the data of each item itself is anonymized. In the fourth process, the data of each item is transmitted to the server 400 without anonymization.

[0062] Note that the selection unit 260 may select a process to be applied to each item from among a plurality of processes including the first process. The plurality of processes do not necessarily include all of the second process, the third process, and the fourth process, and may include at least any one of them.

[0063] When performing the first process, the model generation unit 21 generates a local model based on the item data, and the anonymization unit 220 generates a plurality of shares from the model parameters and transmits them to the calculation server group 30. When performing the second process, the model generation unit 21 generates a local model based on the item data, and the transmission unit 250 transmits the model parameters to the server 400. When performing the third process, the anonymization unit 220 generates a plurality of shares from the item data and transmits them to the calculation server group 30. When performing the fourth process, the transmission unit 250 transmits the item data to the server 400 without anonymization.

[0064] The selection unit 260 may select the process to be applied according to the confidentiality of the data of each item. For example, the selection unit 260 may select the third process or the fourth process that does not perform federated learning instead of the first process that performs federated learning for items with high confidentiality. Also, the selection unit 260 may select the second process that does not anonymize the model parameters instead of the first process that anonymizes the model parameters for items with low confidentiality.

[0065] The level of confidentiality may be set for each item by the user who operates the client terminal 200 when inputting compound data. Also, the level of confidentiality may be set in advance for each item of the compound data set.

[0066] Also, the selection unit 260 may select whether to apply the first process of anonymizing the parameters or the second process of not anonymizing the parameters according to the amount of calculation required when integrating the local models. When the amount of calculation required when integrating the local models is large (for example, when including processes other than arithmetic operations or when the number of parameters is large), the selection unit 260 may select the second process instead of the first process.

[0067] The amount of calculation required when integrating the local models may be determined according to the size of each item data. Also, for each item, the amount of calculation required when integrating the model may be estimated in advance.

[0068] Further, the selection unit 260 may select whether to apply a third process of anonymizing item data or a fourth process of not anonymizing item data according to the magnitude of the assumed computational complexity of the data for each item. The selection unit 260 may select the fourth process out of the third process and the fourth process for an item for which a large computational complexity is assumed. The selection unit 260 may have a function of estimating the computational complexity applied to the data of each item. The selection unit 260 determines the process to be applied to each item based on the estimation result.

[0069] The computational complexity may be determined according to the assumed calculation content for each item. It is known that secret calculation can be processed in a realistic time if it is about arithmetic operations, but the logarithmic coefficient cannot be processed in a realistic time. When the prediction unit 240 makes a calculation request including a process other than arithmetic operations, the selection unit 260 may select the fourth process.

[0070] The selection unit 260 may cause the calculation server group 30 to actually perform the calculation and select whether to apply the third process or the fourth process based on the time taken. In such a case, the selection unit 260 transmits a part of the data of each item to the calculation server group 30, actually executes a predetermined calculation (for example, calculation of an average value, etc.), and measures the computational complexity based on the execution result.

[0071] Further, the selection unit 260 may select the process to be applied to the data of each item in consideration of the desired processing time set for each item. For example, when the desired processing time is short, the selection unit 260 may select the fourth process instead of the third process. Also, when the desired processing time is short, the first process or the second process may be selected. Further, when the priority of confidentiality and computational complexity is set for each item, the selection unit 260 may determine the process to be applied to the data of each item in consideration of the priority.

[0072] Note that, for each item of the set of compound data, which process to apply may be determined in advance. The selection unit 260 selects the process to be applied to each item based on the determination result.

[0073] The selection unit 260 may determine to apply the first process to the items related to the characteristics of the compound. This is because the data related to the characteristics of the compound is not highly confidential and the computational load when integrating the local model is not large.

[0074] FIG. 14 is a flowchart showing an example of the selection method by the selection unit 260. Note that FIG. 14 is merely an example. In FIG. 14, the computational load is determined after determining the confidentiality, but the confidentiality may be determined after determining the computational load.

[0075] First, the selection unit 260 acquires a set of compound data (step S11). Next, the selection unit 260 determines whether the confidentiality of each item data is high (step S12).

[0076] If the confidentiality is high (YES in step S12), the selection unit 260 determines whether the computational load when the prediction unit 240 makes a prediction is large (step S13). If the computational load is large (YES in step S13), the selection unit 260 selects the fourth process of transmitting the item data to the server 400 without anonymizing it. If the computational load is not large (NO in step S13), the selection unit 260 selects the third process of anonymizing the item data and transmitting it to the computing server group 30.

[0077] If the confidentiality is not high (NO in step S12), the selection unit 260 determines whether the computational load required when integrating the local model is large (step S14). If the computational load is large (YES in step S14), the selection unit 260 selects the second process of transmitting the model parameters generated based on the item data to the server 400. If the computational load is not large (NO in step S14), the selection unit 260 selects the first process of anonymizing the model parameters generated based on the item data and outputting them to the computing server group 30.

[0078] According to the computing system 100 according to Embodiment 2, an optimal process can be selected for each compound data. According to the computing system 100, data with high confidentiality can be secretly distributed and stored, so security can be improved.

[0079] Note that when the above-described program is loaded into a computer, it includes a set of instructions (or software code) for causing the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, a computer-readable medium or a tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD), or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray (registered trademark) disc, or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices. The program may be transmitted on a transient computer-readable medium or a communication medium. By way of example and not limitation, a transient computer-readable medium or a communication medium includes electrical, optical, acoustic, or other forms of propagated signals.

[0080] The present invention has been described above with reference to the embodiments, but the present invention is not limited thereto. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the invention.

[0081] Some or all of the above embodiments may be described as follows, but are not limited thereto. (Appendix 1) A confidentiality means for performing a first process of anonymizing the parameters of the model after generating the model from a set of compound data at each of a plurality of client terminals; A secure computing means for performing secure computing for integrating the model using the anonymized parameters; A computing system comprising the same. (Appendix 2) The computing system is Further comprising selection means for selecting a process to be applied to each item of the set of compound data from among the first process and one or more processes. The one or more processes include at least any one of three processes: a second process of transmitting the parameters to the server without anonymizing them after generating the model based on the data of each item, a third process of anonymizing the data of each item itself, and a fourth process of transmitting the data of each item to the server without anonymizing it. The computing system according to Supplementary Note 1. (Supplementary Note 3) The selection means selects a process to be applied to each item according to the confidentiality of the data of each item and the assumed computational amount for the data of each item. The computing system according to Supplementary Note 2. (Supplementary Note 4) The selection means estimates the computational amount and selects a process to be applied to each item based on the estimation result. The computing system according to Supplementary Note 3. (Supplementary Note 5) The selection means estimates the computational amount by actually executing a calculation using a part of the data of each item. The computing system according to Supplementary Note 4. (Supplementary Note 6) The selection means selects a process to be applied to each item in consideration of a specified desired processing time. The computing system according to Supplementary Note 3. (Supplementary Note 7) The set of compound data includes items related to the structure of the compound, items related to simulation results, items related to the production process of the compound, and items related to the properties of the compound. Which process to apply for each item is determined in advance. The computing system according to Supplementary Note 2. (Supplementary Note 8) The selection means selects to apply the first process to the items related to the properties of the compound. The computing system according to Supplementary Note 7. (Appendix 9) A means for predicting properties from the structure of a compound using the model, The calculation system according to Appendix 8. (Appendix 10) A means for predicting the structure from the properties of a compound using the model, The calculation system according to Appendix 8. (Appendix 11) After generating a model from a set of compound data on each of a plurality of client terminals, a first process of anonymizing the parameters of the model is performed, Performing secret calculation for integrating the model using the anonymized parameters, Calculation method.

Explanation of symbols

[0082] 1, 10, 100 Calculation system 2, 2a, 2b, 2c, 20, 20a, 20b, 20c, 200, 200a, 200b, 200c Client terminal a, b, c Local model 30 Calculation server group 3, 31, 31_1, 31_2, 31_3, 32, 32_1, 32_2, 32_3 Calculation server 11, 22, 220 Anonymization unit 311, 321 Shared memory unit 12, 312, 322 Secret calculation unit 21 Model generation unit 23, 230 Acquisition unit 24, 240 Prediction unit 250 Transmission unit 260 Selection unit 400 Server 410 Memory unit 420 Calculation unit

Claims

1. After generating a model from a set of compound data on each of a plurality of client terminals, a anonymization means for performing a first process of anonymizing the parameters of the model; A secure computing means for performing secure computing for integrating the model using the anonymized parameters; A selection means for selecting a process to be applied to each item of the set of compound data from among the first process and one or more processes; Comprising; The one or more processes include at least any one of a second process of transmitting the parameters to a server without anonymizing them after generating the model based on the data of each item, a third process of anonymizing the data of each item, and a fourth process of transmitting the data of each item to the server without anonymizing them; A computing system.

2. The selection means; Selects a process to be applied to each item according to the anonymization of the data of each item and the assumed amount of computation for the data of each item. The computing system according to Claim 1.

3. The selection means; Estimates the amount of computation and selects a process to be applied to each item based on the estimation result. The computing system according to Claim 2.

4. The selection means; Estimates the amount of computation by actually executing a computation using a part of the data of each item. The computing system according to Claim 3.

5. The selection means; Selects a process to be applied to each item in consideration of a specified desired processing time. The computing system according to Claim 2.

6. The set of compound data includes items related to the structure of the compound, items related to simulation results, items related to the production process of the compound, and items related to the properties of the compound. Which process to apply is determined in advance for each item. The computing system according to Claim 1.

7. The selection means selects to apply the first process to the item related to the properties of the compound. The computing system according to Claim 6.

8. Comprising means for predicting properties from the structure of a compound using the model. The computing system according to Claim 7.

9. Comprising means for predicting the structure from the properties of a compound using the model. The computing system according to Claim 7.

10. After generating a model from a set of compound data on each of a plurality of client terminals, performing a first process of anonymizing the parameters of the model; Performing secret calculations for integrating the model using the anonymized parameters; Selecting a process to be applied to each item of the set of compound data from among the first process and one or more processes; including; the one or more processes include at least any one of three processes: a second process of transmitting the parameters to the server without anonymizing them after generating the model based on the data of each item, a third process of anonymizing the data of each item, and a fourth process of transmitting the data of each item to the server without anonymizing it; Calculation method.

Citation Information

Patent Citations

  • Secure computation conversion device, secure computation system, secure computation conversion method, and secure computation conversion program

    JP6795863B1

  • Systems and Methods for Providing a Modified Loss Function in Federated-Split Learning

    US20220029971A1

  • Material descriptor generation method, material descriptor generation device, material descriptor generation program, prediction model building method, prediction model building device, and prediction model building program

    WO2020031671A1

  • Integrated analysis method, integrated analysis device, and integrated analysis program

    WO2021090789A1