Privacy-enhancing machine learning

By deploying a multi-participant machine learning system in a trusted execution environment, selecting the first part of the training data for training models, and keeping the second part private, the problem of limited availability of diverse labeled data is solved, and privacy protection and performance improvement of machine learning models is achieved.

CN114430835BActive Publication Date: 2025-05-06MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080065917.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-18
Filing Date
2020-06-18
Publication Date
2025-05-06
Estimated Expiration
2040-06-18

AI Technical Summary

Technical Problem

In machine learning, especially in deep learning, the main obstacle to training well-performing models is the limited availability of diversely labeled training data and is difficult to utilize because the data is distributed and owned by multiple parties.

Method used

By deploying a multi-participant machine learning system in a trusted execution environment, the selector component is used to select the first part of the training data from the training data uploaded by multiple participants, and only that part of the data is used to train the machine learning model while keeping the second part of the training data private. Use the participation controller to calculate the contribution metric of each participant to the machine learning model to control access to the model.

Benefits of technology

Enhanced data privacy, preventing malicious parties from reverse-engineering unused training data by accessing trained models, while ensuring that the performance of the machine learning model is not compromised, and may even improve in some cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114430835B_ABST
    Figure CN114430835B_ABST
Patent Text Reader

Abstract

A method for selecting data for privacy preserving machine learning includes storing training data from a first party, storing a machine learning model, and storing criteria from the first party or another party. The method includes filtering the training data to select a first portion of the training data to be used to train the machine learning model and selecting a second portion of the training data. The selection is done by calculating a measure of the contribution of the data to the performance of the machine learning model using the criteria.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In machine learning, and in particular deep learning, the main obstacle to training well-performing machine learning models is often the limited availability of sufficiently diverse labeled training data. However, the data needed to train a good model often exists, but is not easily exploitable because it is distributed and owned by multiple parties. For example, in the medical field, important data about patients that can be used to learn a cancer diagnosis support system may be owned by different hospitals, where each of these hospitals maintains different data from a specific geographic area with different demographic characteristics.

[0002] The embodiments described below are not limited to implementations that solve any or all shortcomings of known machine learning systems. Summary of the invention

[0003] A simplified summary of the present disclosure is presented below to provide the reader with a basic understanding. This summary is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Its sole purpose is to present a selection of concepts disclosed herein in a simplified form as a prelude to a more detailed description presented later.

[0004] In various examples, there is a method of selecting data for privacy-preserving machine learning, comprising: storing training data from a first party, storing a machine learning model, and storing criteria from the first party or from another party. The method comprises: filtering the training data to select a first portion of the training data to be used to train the machine learning model and selecting a second portion of the training data. The selection is done by calculating a measure of the contribution of the data to the performance of the machine learning model using the criteria.

[0005] Many of the attendant features will be more readily appreciated as they become better understood by reference to the following detailed description considered in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] This description will be better understood through the following detailed description read in conjunction with the accompanying drawings, wherein

[0007] Figure 1 is a schematic diagram of the multi-party machine learning system used by two hospitals;

[0008] Figure 2 yes Figure 1 A more detailed diagram of a multi-party machine learning system;

[0009] Figure 3 is a schematic diagram of another multi-party machine learning system;

[0010] Figure 4is a flowchart of a method for privacy-preserving training performed by a multi-party machine learning system, wherein a single machine learning model is being trained;

[0011] Figure 5 is a flow chart of a method for privacy-preserving training performed by a multi-party machine learning system, wherein multiple machine learning models are being trained;

[0012] Figure 6 is a flow chart of a method of controlling access to a single trained machine learning model;

[0013] Figure 7 is a flow chart of a method of controlling access to a plurality of trained machine learning models;

[0014] Figure 8 An exemplary computing-based device in which an embodiment of a multi-party machine learning system is implemented is illustrated.

[0015] In the drawings, the same reference numerals are used to denote the same components. DETAILED DESCRIPTION

[0016] The specific embodiments provided below in conjunction with the accompanying drawings are intended to be a description of this example, and are not intended to represent the only form of constructing or utilizing this example. This description sets forth the functions of the example and the order of operations for constructing and operating the example. However, the same or equivalent functions and sequences may be implemented by different examples.

[0017] As mentioned above, in the medical field, important data about patients that can be used to learn a cancer diagnosis support system may be owned by different hospitals, where each hospital maintains different data from a specific geographic area with different demographic characteristics. By pooling the available data, hospitals can train better machine learning models for their applications than if they only used their own data. Since all hospitals will benefit from better machine learning models obtained through data sharing, there is a need for collaborative machine learning.

[0018] Of course, this type of collaboration raises technical challenges related to one or more of the following: security of the individual participants' data, privacy of the individual participants' data, quality of the machine learning results. It is difficult to deploy a collaborative machine learning system that respects the integrity of the individual participants' data. In this context, integrity involves ensuring that a participant's training data is not modified during collaborative training and that the data submitted by the participant is indeed used for collaborative training.

[0019] although Figure 1Described and shown herein as being implemented for use by a hospital, the described system is provided as an example rather than a limitation. Those skilled in the art will appreciate that this example is applicable to applications in a variety of different types of multi-party machine learning scenarios.

[0020] Figure 1 is a schematic diagram of a multi-party machine learning system 100 being used by two hospitals 108, 112. The multi-party machine learning system 100 is deployed in a trusted execution environment 104 in the cloud or at any location that communicates with the hospital via a communication network 106 such as the Internet, an intranet, or any other communication network 106. The trusted execution environment 104 is implemented using secure hardware and / or software and includes at least one secure memory area. In this example, for clarity, one trusted execution environment 104 is shown, although in practice many trusted execution environments 104 may be deployed and located at computing units in a data center such as servers with disk storage devices or virtual machines connected via a network within the data center. Where there are many trusted execution environments 104, the multi-party machine learning system 100 is distributed among them.

[0021] In one example, the trusted execution environment 104 includes a secure memory area, which is a processor-protected memory area within the address space of a regular process. The processor monitors memory access to the trusted execution environment so that only code running in the trusted execution environment 104 can access data in the trusted execution environment 104. When inside the physical processor package (in the processor's cache), the memory of the trusted execution environment 104 is provided in plain text, but when written to the system memory (random access memory RAM), it is encrypted and integrity protected. External code can only call code inside the trusted execution environment at statically defined entry points (using a call-gate mechanism).

[0022] In some examples, the trusted execution environment 104 is implemented using hardware so that the secure memory area is isolated from any other code including the operating system and the hypervisor. In some examples, the trusted execution environment 104 is implemented using a trusted virtual machine.

[0023] Within the trusted execution environment 104 are one or more trained machine learning models 102 that have been computed by the multi-party machine learning system 100 using training data from multiple parties, such as Figure 1In the example of FIG. 1 , there is a first hospital 108 and a second hospital 112. The first hospital 108 has training data 110, such as medical images of patients, where the medical images are labeled as depicting or not depicting cancer. The training data 110 is confidential and securely stored at the first hospital. When the hospital wants to participate in multi-party machine learning, the training data 110 is encrypted and uploaded to the trusted execution environment 104.

[0024] The second hospital 112 has training data 114, which includes medical images of different patients, where the medical images are labeled as depicting or not depicting cancer. The training data 114 is confidential data and is securely stored at the second hospital. When the second hospital wants to participate in multi-party machine learning, the training data 114 is encrypted and uploaded to the trusted execution environment 104.

[0025] One or more malicious parties, such as malicious party 116, may exist and have fake training data 118. Fake training data is any training data of poor quality, such as having inaccurate labels, or a duplication of training data that has been uploaded to the trusted execution environment by the party.

[0026] One or more parties upload the training data to the trusted execution environment 104. The multi-party machine learning system 100 trains one or more machine learning models 102 using at least some of the training data. One or more of the parties are then able to access the trained machine learning model and use it to compute predictions to label medical images for tumor detection or other tasks, depending on the application area. In this way, a first party, such as hospital one, is able to benefit from a high-performance machine learning model that has been trained using data from multiple parties. The performance of the resulting machine learning model is lower if the first party uses only its own training data, where the amount and / or variety of data is typically lower than the amount and / or variety of data available to the multiple parties.

[0027] Figure 1 The multi-party machine learning system allows multiple parties to jointly train machine learning models based on the training data provided by all parties and achieve improved performance on their respective tasks. The multi-party machine learning system supports single validation task scenarios, for example, hospitals pool their data together to train a single model for detecting cancer. In addition, the multi-party machine learning system also supports scenarios where one party's data contributes to multiple tasks.

[0028] Assume that training data from the first hospital 110 and the second hospital 112 are uploaded to the trusted execution environment 104 and used by the multi-party machine learning system to train one or more of the machine learning models 102. Then, assume that the first hospital and the second hospital can access the resulting trained machine learning model 102 via the communication network 106. Then, the first hospital may discover information about the training data used to train the machine learning model 102. Therefore, the first hospital is able to discover the confidential training data of the second hospital. Among them, attacks that obtain confidential training data from prediction application programming interfaces are known, such as described in "Stealing machine learning models via prediction APIs" by Tramer et al. in USENIX Security 2016.

[0029] Various examples described herein use a selector component within a multi-party machine learning system to enhance privacy. The selector component selects a first portion of training data from training data uploaded by multiple parties, and uses only the first portion of the training data to train one or more machine learning models. The second portion of the training data remains private in the trusted execution environment. The selection is completed based on one or more criteria submitted by individual parties. In this way, at least some of the training data that has been uploaded to the trusted execution environment 104 is not used to train a specific machine learning model instance. Therefore, privacy is enhanced because malicious parties who access the trained machine learning model cannot discover unused training data. Although some but not all training data is used, the performance of the machine learning model is not affected by carefully designing the selection process. In some cases, the criteria include verification data, and the benefits of using a selector are as follows: only information related to the verification task of the verification data is published by the model, thereby limiting the possibility of copying the training data and reusing it for other tasks.

[0030] Various examples described herein use a participant controller within a multi-party machine learning system in order to improve the quality of the resulting trained machine learning model 102 and prevent errors such as Figure 1 A spoofing attack in which a malicious participant 116 uses fake training data (such as training data that has already been used) to gain access to the trained machine learning model 102. The participant controller calculates a measure of the contribution of an individual participant to a particular trained machine learning model and uses the measure to control access to that model or other machine learning models.

[0031] Figure 2 Such as Figure 1A more detailed diagram of a multi-party machine learning system of a multi-party machine learning system. The trusted execution environment 104 includes the multi-party machine learning system 100.

[0032] The multi-party machine learning system 100 includes a memory that stores training data 200 and stores a model library 202 including at least one machine learning model. The multi-party machine learning system 100 optionally includes a selector 204, and it includes a standard storage device 206, a training engine 208, a participant controller 210, and a storage device for holding one or more trained machine learning models 102 calculated by the training engine 208.

[0033] Because it is stored in the trusted execution environment, the stored training data 200 is stored in plain text. The training data 200 includes multiple examples, such as images, videos, documents, sensor data values, or other training examples. The training data 200 includes labeled training data in the case where the training engine 208 uses supervised training and / or unlabeled training data in the case where unsupervised training is used. The stored training data 200 has been received from two or more participants at the trusted execution environment 104. Figure 2 In the example, two parties are shown as 108, 112, but there may be more parties in practice. When a party uploads training data, the training data is encrypted. The training data 200 stored at the trusted execution environment is tagged or marked to indicate which party it originated from.

[0034] The model library 202 is a storage device for one or more machine learning models, such as a neural network, a random decision forest, a support vector machine, a classifier, a regressor, or other machine learning models.

[0035] The selector 204 is optional and is included in situations where the multi-party machine learning system enhances privacy by selecting some of the training data 200 to be used to train a particular instance of a machine learning model, but not all of the training data 200. The selector uses one or more criteria provided by individual ones of the parties. Figure 2 It shows that Party 1 uploads standards 220 and training data 222 to the trusted execution environment. It also shows that Party 2 uploads standards 224 and training data 226 to the trusted execution environment.

[0036] The criteria storage device 206 maintains the criteria uploaded by each individual participant in each participant. A criterion is a quality, threshold, value, metric, statistic or other standard used to select training data and / or indicate the performance level of a machine learning model.

[0037] The training engine 208 is one or more training processes for training machine learning models from the model library 202. In some examples, the training process is a well-known conventional training process.

[0038] Participation controller 210 includes functionality for calculating a measure of the contribution of individual participants' training data to the performance of a particular trained machine learning model. In some examples, the participation controller uses a metric. More details about the participation controller are given later in this document.

[0039] Trained machine learning model 102 is a stored architecture, parameter values, and other data that specifies a single trained machine learning model.

[0040] The access controller 212 is a firewall, network card, or other functionality that enables control of access to the trained machine learning model 102 by individual parties 108, 112.

[0041] The selector of a multi-party machine learning system operates in an unconventional manner to enhance the privacy of a trained machine learning model without compromising the performance of the machine learning model.

[0042] A selector of a multi-party machine learning system improves the functionality of an underlying computing device by selecting a first portion of training data to be used to train the machine learning model and selecting a second portion of the training data to remain private, thereby maintaining the performance of a trained machine learning model.

[0043] Participating controllers of a multi-party machine learning system operate in an unconventional manner to protect access to trained machine learning models.

[0044] A participant controller of a multi-party machine learning system improves the functionality of underlying computing devices by increasing the security of access to trained machine learning models and preventing spoofing attacks in which malicious participants spoof training data in an attempt to gain access to trained machine learning models.

[0045] Alternatively or additionally, reference Figure 2 The functions described herein are at least partially performed by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and graphics processing units (GPUs).

[0046] Figure 3 is with Figure 2Schematic diagram of a multi-party machine learning system that is very similar and in which like components have like reference numerals. In this example, the criteria uploaded by each party are tasks, where the tasks are verification data for verifying the machine learning tasks. Figure 3 As shown, participant 1 108 uploads the encrypted training data to the training data server 200 in the trusted execution environment, and uploads the task (in this example, the task is verification data) to the standard storage device 206 (in Figure 3 There are multiple participants, each of which can participate in multi-participant machine learning only when it uploads a standard such as a task (i.e., verification data in this example) to the multi-participant machine learning system. Participant M 112 uploads the encrypted training data to the training data server 200 in the trusted execution environment, and uploads the task (in this example, the task is verification data) to the standard storage device 206 (in Figure 3 is called a task server). Figure 3 The padlocks 300, 302 in the figure represent access control mechanisms, such as firewalls, network cards, software or other access controls, which prevent each participant from accessing the multi-party machine learning system unless it has uploaded standards, such as verification data in this example.

[0047] Using the uploaded training data, the multi-party machine learning system performs machine learning. In some examples, it trains a single machine learning model, which each of the parties may then access. In some examples, it trains multiple machine learning models, one for each of the parties.

[0048] In some examples, the use is Figure 2 The selector 204 is a task and data matching component of the selector 204. The selector uses the uploaded criteria to select which training data items are used to train specific models in the model and which training data items remain private and are not used to train specific models in the model. As mentioned above, the criteria are qualities, thresholds, values, metrics, statistics, or other criteria used to select training data and / or indicate the performance level of the machine learning model. In one example, the criteria include validation data, and in this case, the validation data is used to evaluate the performance of the trained machine learning model in a conventional manner. In one example, the criteria include qualities such as roundness, where it is known that for a specific result, images depicting circular objects will be better for training the machine learning system. In one example, the criteria include a number of corrupted bits, where it is known that for a specific result, an audio signal with a high level of corrupted bits will be worse for training the machine learning system.

[0049] Customized machine learning model training components and Figure 2The training engine 208 is the same as the training engine 208. It uses the selected training data to train the machine learning model from the machine learning model library 202. The output of the training engine 208 is stored in the trained machine learning model 102.

[0050] The stored trained machine learning model 102 is Figure 2 Access controller 212 Figure 3 The participation controller 210 calculates scores that indicate whether (and optionally for how long) each individual participant has access to each individual trained machine learning model 102 in the trained machine learning model 102.

[0051] Since the number of participants changes over time as different participants leave or join the multi-party machine learning system, Figure 3 The device is shown for a specific time instance. Thus, the training data at storage device 200, the criteria at storage device 206, and the trained machine learning models at storage device 102 change over time as the device operates. It is possible that, depending on the calculation results of the participating controller 210, a participant may initially access individual ones of the machine learning models, but lose that access over time (and potentially regain access again). A participant is able to participate in the multi-participant machine learning system as long as it submits criteria.

[0052] Figure 4 4 is a flow chart of a method performed by a multi-party machine learning system for privacy-preserving training, wherein a single machine learning model is being trained. The machine learning model is stored 400, such as by selecting the machine learning model from a model library 202 according to one or more rules or according to a selection parameter value given by one of the parties. Figure 2 The first participant of participant 1 receives 402 training data. Figure 2 One or more participants of participant 2 receive 404 training data.

[0053] The multi-party machine learning system 100 checks 406 whether criteria have been received from the first party. If not, the multi-party machine learning system waits and continues to check for arrival of criteria. If criteria have been received from party one, the process continues by using a selector 204. The selector selects a first portion of training data to be used to train the stored machine learning model. The selector selects a second portion of training data to remain private and not be used to train the stored machine learning model. The selection is done based on the criteria from party one.

[0054] The multi-party machine learning system trains 410 the stored machine learning model using the first portion of the training data. The selection of the criteria to use is done in a manner that does not harm the performance of the resulting trained model compared to the performance of the model trained using all available training data.

[0055] In some examples, the resulting trained machine learning model is deployed 412 by retaining the resulting trained machine learning model in a trusted execution environment and allowing access to the trained machine learning model via the access controller 212. Parties that send queries to and obtain results from the trained machine learning model in the trusted execution environment cannot reverse engineer training data that is kept private and is not used to train the machine learning model. In some cases, the access controller 212 and the scores from the participation controller 208 are used to control 414 the ability of parties to send queries to and obtain results from the trained machine learning model, as described in more detail later in this document.

[0056] In some examples, the resulting trained machine learning model is deployed 412 by installing it on an end-user device or on a server outside the trusted execution environment. In this case, security is enhanced compared to deploying the trained machine learning model after training it on all available training data. A malicious party that attacks the deployed machine learning model to obtain the used training data cannot obtain the training data that remains private in the trusted execution environment.

[0057] Figure 5 is a flow chart of a method for privacy-preserving training performed by a multi-party machine learning system, wherein multiple machine learning models are being trained. Figure 3 An example is given in which each party has a machine learning model.

[0058] The multi-party machine learning system receives 500 training data from a first party and receives 502 training data from one or more other parties. The multi-party machine learning system checks 504 whether it has received standards from each party. Each party that has submitted standards can participate. Figure 3 Access controls 300, 302 prevent entities that have not submitted standards from accessing the multi-party machine learning system, and therefore the entity is not a participant.

[0059] For each participant, the multi-participant machine learning system selects 506 some training data but not all training data based on the criteria of the corresponding participant. For each participant, the multi-participant machine learning system trains 508 a machine learning model using the appropriate selected training data.

[0060] After the individual machine learning models have been trained, they are deployed 510. Deploying the individual models is accomplished by enabling access controller 212 to enable parties associated with the individual models to send queries and receive responses from the individual models. In some cases, access is controlled 512 based on a score calculated by engagement controller 210, as described below. However, use of engagement controller 210 is not required.

[0061] Figure 6 6 is a flow chart of a method for controlling access to a single trained machine learning model that has been trained by a multi-party machine learning system. The multi-party machine learning system stores 600 machine learning models. The machine learning model is selected from a model library 202 according to one or more rules or according to a selection parameter value given by one of the participants. Figure 2 A first participant of participant one receives 602 training data and receives training data from a user such as Figure 2 One or more participants of participant two receive 604 training data.

[0062] The multi-party machine learning system 100 checks 606 whether the criteria have been received from the first party. If not, the multi-party machine learning system waits and continues to check for the arrival of the criteria.

[0063] The multi-party machine learning system uses some or all of the training data to train 608 the stored machine learning model.

[0064] For each participant, the participant controller calculates 610 a measure of the contribution of the training data submitted by the participant to the performance of the trained machine learning model. The measure of the contribution is calculated using the criteria submitted by participant 1.

[0065] The resulting trained machine learning model is deployed 412 by retaining the resulting trained machine learning model in a trusted execution environment and allowing access to the trained machine learning model via the access controller 212. The access granted to a party is related to a measure of contribution calculated for that party. For each party, a check is made 612 to see if the measure of contribution is above a threshold. If so, access to the trained model is granted 616. If not, at 614, access is blocked.

[0066] In some cases, Figure 4 and Figure 6That is, the machine learning model uses the method based on the above reference Figure 4 The selected training data is selected for training according to the description. Then, according to the reference Figure 6 Controls access to trained models.

[0067] Figure 7 is a flow chart of a method performed by a multi-party machine learning system in which multiple machine learning models are being trained and in which access to the individually trained models is controlled using a participation controller 210 and an access controller 212.

[0068] The multi-party machine learning system receives 700 training data from a first party and receives 702 training data from one or more other parties. The multi-party machine learning system checks 704 whether it has received standards from each party. Each party that submitted a standard can participate. Figure 3 Access controls 300, 302 prevent entities that have not submitted standards from accessing the multi-party machine learning system and are therefore not participants.

[0069] For each participant, the multi-participant machine learning system trains 706 a machine learning model using all or some of the training data (thus, it may be trained using all of the training data submitted by all participants).

[0070] For each participant, the multi-participant machine learning system calculates 708 a measure of the contribution of the participant's training data to the performance of each of the machine learning models.

[0071] For each participant and each model, the multi-participant machine learning system checks 712 to see if the metric of contribution is above a threshold. If so, the particular participant is given 716 access to the trained model. If not, at 714, access is blocked.

[0072] In various examples, the selector 204 and the participant controller 210 calculate Shapley values. The Shapley value is the output of a function that takes a characteristic function and a participant i as parameters. The characteristic function used by the selector 204 is different from the characteristic function used by the participant controller 210.

[0073] The characteristic function υ and the Shapley value of the participant i∈M are as follows:

[0074]

[0075] Expressed in words as follows: the characteristic function υ and the Shapley value of a participant i that is a member of a set M of participants of a multi-participant machine learning system is given by the sum of the factorial of the cardinality of the set S over every possible set S of M participants that does not include i, multiplied by the factorial of the number of participants M minus the cardinality of the set S minus 1, divided by the factorial of the number of participants M, multiplied by the difference in the output of the characteristic function of S with i and S without i.

[0076] The Shapley values ​​quantify the average marginal contribution of participant i to a subset of all possible participants. The inventors have recognized that Shapley values ​​are not robust to replication, that is, they do not take into account participants that submit the same training data multiple times.

[0077] The selector 204 uses the following feature function when calculating the Shapley value for the single machine learning model case and the case of one machine learning model per participant:

[0078]

[0079] In words: The characteristic function used by the selector 204 when calculating the Shapley value is considered as a parameter of the number of possible sets S of participants and is equal to the gain function The output of a specific machine learning model The training data available from all the participants in the combination of participants in one of the sets j in the possible set S has been used After training, when using the standard (such as verification data provided by participant i) for this specific machine learning model When evaluating, this gain function expresses the The notation used in this paper is To represent the characteristic function, so as to express the gain function Use as a characteristic function.

[0080] With a single machine learning model, the participating controller 210 uses the following feature function when calculating the Shapley value:

[0081]

[0082] Expressed in words as follows: The characteristic function used by the participating controller 210 for S possible sets of participants in a multi-participant machine learning system is equal to the sum of the performance of the model plus the performance of the model of each individual participant. The symbol v is used to refer to the characteristic function used by the participating controller for a single machine learning model.

[0083] The feature function is the value of the model trained on all datasets in S plus the marginal benefit of each participant. Note that for a single participant, the value of the data is expressed as the value of the model trained on its own training dataset.

[0084] In the case of one machine learning model per participant, the participant controller 210 uses the following feature function when calculating the Shapley value:

[0085]

[0086] In words: the characteristic function used by the participating controller 210 when calculating the Shapley value in the case of one machine learning model per participant is equal to the sum of the performance of all models plus the sum of the performance gain of each model of each individual participant based only on the data of the individual participant. The symbol ω is used to refer to the characteristic function used by the participating controller in the case of multiple machine learning models.

[0087] Figure 8 Various components of an exemplary computing-based device 800 are illustrated, which is implemented as any form of computing device and / or electronic device and in which, in some examples, embodiments of a multi-party machine learning system are implemented.

[0088] The computing-based device 800 includes one or more processors 802, which are microprocessors, controllers, or any other suitable type of processors for processing computer-executable instructions to control the operation of the device to train one or more machine learning models using training data from one or more participants. In some examples, for example, where a system-on-chip architecture is used, the processor 802 includes one or more fixed function blocks (also referred to as accelerators) that are implemented in hardware (rather than software or firmware) Figures 4 to 7 A portion of the method of any of the figures in the drawings. Platform software including an operating system 804 or any other suitable platform software is provided at a computing-based device to enable application software 806 to execute on the device. A trusted execution environment 104 is provided to maintain training data and machine learning models, as described previously herein.

[0089] Computer executable instructions are provided using any computer readable medium accessible to the computing-based device 800. Computer readable media include, for example, computer storage media such as memory 808 and communication media. Computer storage media such as memory 808 include volatile and non-volatile media, removable media, and non-removable media implemented in any method or technology for storing information such as computer readable instructions, data structures, program modules, etc. Computer storage media include, but are not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electronic erasable programmable read-only memory (EEPROM), flash memory or other storage technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, cassettes, tapes, disk storage or other magnetic storage devices, or any other non-transmission media for storing information for access by a computing device. In contrast, communication media embody computer readable instructions, data structures, program modules, etc. in modulated data signals such as carrier waves or other transmission mechanisms. As defined herein, computer storage media do not include communication media. Therefore, computer storage media themselves should not be interpreted as propagating signals. Although computer storage media (memory 808 ) is shown within computing-based device 800 , it should be appreciated that in some examples the storage is distributed or remotely located and accessed via a network or other communications link (eg, using communications interface 810 ).

[0090] The computing-based device 800 also includes an input / output controller 812, which is arranged to output display information to a display device 814, which can be separate from or integrated into the computing-based device 800. The display information can provide a graphical user interface. The input / output controller 812 is also arranged to receive and process input from one or more devices such as a user input device 816 (e.g., a mouse, keyboard, camera, microphone or other sensor).

[0091] Alternatively or in addition to other examples described herein, examples include any combination of:

[0092] A method for selecting data for privacy-preserving machine learning, comprising:

[0093] storing training data from the first party;

[0094] Storing machine learning models;

[0095] storing criteria from the first party or another party;

[0096] selecting training data to select a first portion of the training data to be used to train the machine learning model and selecting a second portion of the training data;

[0097] The selection is done by using a standard measure of the contribution of the data to the performance of the machine learning model.

[0098] In this way, privacy is enhanced because the first portion of the training data can be used to train the machine learning model without using the second portion. Therefore, the second portion cannot be revealed through access to the trained model. By using the criteria for selection, the performance of the model is improved even if the model is not trained on all available training data.

[0099] The method described above is performed in a trusted execution environment and includes training a machine learning model using a first portion of training data such that a second portion of the training data remains private in the trusted execution environment. Security is enhanced by using a trusted execution environment.

[0100] The method described above, wherein the criteria include one or more of the following: quality, threshold, value, metric, statistic. Since the trusted computing environment is a resource-constrained entity, it is efficient to calculate and store these in a multi-party machine learning system.

[0101] The method described above, wherein the criterion is adapted to select the training data based on the likelihood of the performance of the machine learning model when trained using the selected training data. This helps the performance of the machine learning model even if not all available training data is used. Efficiency is gained.

[0102] The method described above, wherein a criterion is applied to indicate the performance level of a machine learning model.

[0103] The method described above, wherein the standard is validation data used to validate the machine learning task for which the machine learning model is to be trained.

[0104] The method described above, where the metric is a Shapley value computed using a characteristic function, where the characteristic function is equal to the performance of the machine learning model when evaluated using a criterion given by participant i after the model has been trained using training data available from all participants in a combination of participants in a set of multiple possible participants S. This provides an efficient and practical way to select training data that has been found to perform well in empirical testing.

[0105] The method described above, wherein there are multiple machine learning models. Using multiple machine learning models provides flexibility and enables different parties to train different models.

[0106] The method described above includes: calculating a measure of the contribution of the first party's training data to the performance of the machine learning model, and controlling access to the machine learning model based on the calculated measure. In this way, malicious parties that submit duplicate training data and / or poor quality training data are prevented from accessing the results.

[0107] The method described above, wherein the contribution metric is a Shapley value calculated using a characteristic function, wherein the characteristic function is equal to the sum of the performance of the machine learning model plus the performance of the machine learning model for each individual participant. The characteristic function used in this article was found to perform well in empirical testing.

[0108] The method described above includes storing multiple machine learning models, one machine learning model for each participant, and wherein the metric is a Shapley value calculated using a characteristic function, wherein the characteristic function is equal to the sum of the performance of all machine learning models plus the sum of the performance of each machine learning model for each individual participant. The characteristic function used in this article is found to perform well in practice.

[0109] An apparatus for selecting data for privacy-preserving machine learning, comprising:

[0110] A memory for storing training data from the first party;

[0111] Memory, which stores machine learning models;

[0112] The memory stores criteria from the first party or another party;

[0113] A selector configured to select training data to select a first portion of the training data to be used to train the machine learning model and to select a second portion of the training data; wherein the selection is performed by using a standard to calculate a measure of the contribution of the data to the performance of the machine learning model.

[0114] An apparatus for controlling access to a machine learning model, the apparatus comprising:

[0115] Trusted computing environment, storing machine learning models and training data;

[0116] Access controllers, which are configured to allow or deny access to machine learning models;

[0117] a storage for storing criteria submitted by a party requesting access to a machine learning model;

[0118] Participating controllers, using standard calculation scores;

[0119] And wherein the access controller uses the calculated score to allow or deny access to the machine learning model.

[0120] The apparatus described above, wherein the criterion is suitable for indicating the performance of a machine learning model.

[0121] The apparatus described above, wherein the access controller is configured to prevent parties who submit training data to the trusted computing environment but do not submit standards to the trusted computing environment from accessing the machine learning model.

[0122] The apparatus described above, wherein the access controller uses the calculated score to grant timed access to the machine learning model, the timing being related to the score.

[0123] The apparatus described above, wherein the training data has been submitted by one or more parties, and wherein the access controller prevents access to the machine learning model by a malicious party that submitted the submitted training data.

[0124] The apparatus described above, wherein the participating controller calculates a score as a Shapley value using a characteristic function, wherein the characteristic function is equal to the sum of the performance of the machine learning model plus the performance of the machine learning model for each individual participant.

[0125] The apparatus described above, wherein the trusted computing environment stores multiple machine learning models, one machine learning model for each participant, and the participating controller calculates a score using a characteristic function as a Shapley value, wherein the characteristic function is equal to the sum of the performance of all machine learning models plus the sum of the performance of each machine learning model for each individual participant.

[0126] A method for controlling access to a machine learning model, the method comprising:

[0127] Storing machine learning models and training data in a trusted computing environment;

[0128] Use access controllers to allow or deny access to machine learning models;

[0129] storing, at a memory, criteria submitted by a party requesting access to a machine learning model;

[0130] Scores were calculated using criteria;

[0131] And use the calculated score to allow or deny access to the machine learning model.

[0132] The term 'computer' or 'computing-based device' is used herein to refer to any device having processing capabilities, such that the device executes instructions. Those skilled in the art will appreciate that such processing capabilities are incorporated into many different devices, and thus the terms 'computer' or 'computing-based device' each include personal computers (PCs), servers, mobile phones (including smartphones), tablet computers, set-top boxes, media players, game consoles, personal digital assistants, wearable computers, and many other devices.

[0133] In some examples, the methods described herein are performed by software in a machine-readable form on a tangible storage medium (e.g., in the form of a computer program, including computer program code means, which, when the program is run on a computer and in the case where the computer program may be embodied on a computer-readable medium, is adapted to perform all operations of one or more of the methods described herein). The software is suitable for execution on a parallel processor or a serial processor, so that the method operations may be performed in any suitable order or simultaneously.

[0134] This recognizes that software is a valuable, separately tradable commodity. It is intended to cover software that runs on or controls "virtual" or standard hardware to perform a desired function. It is also intended to cover software that "describes" or defines the configuration of hardware, such as HDL (Hardware Description Language) software used to design silicon chips or used to configure general-purpose programmable chips to perform a desired function.

[0135] Those skilled in the art will recognize that the storage device for storing program instructions is optionally distributed on the network. For example, the remote computer can store examples of processes described as software. Local or terminal computers can access the remote computer and download part or all of the software to run the program. Alternatively, the local computer can download software fragments as needed, or execute some software instructions at the local terminal and execute some software instructions at the remote computer (or computer network). Those skilled in the art will also recognize that, by utilizing conventional techniques known to those skilled in the art, all or part of the software instructions can be executed by special circuits such as digital signal processors (DSPs), programmable logic arrays, etc.

[0136] It will be apparent to one skilled in the art that any range or device value given herein may be expanded or altered without losing the effect sought.

[0137] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0138] It should be understood that the benefits and advantages described above may relate to one embodiment or may relate to several embodiments. The embodiments are not limited to embodiments that solve any or all of the problems described or embodiments that have any or all of the benefits and advantages described. It should be further understood that reference to an item refers to one or more of those items.

[0139] The operations of the methods described herein may be performed in any suitable order or simultaneously where appropriate. Additionally, individual boxes may be deleted from any of the methods without departing from the scope of the subject matter described herein. Aspects of any of the examples described above may be combined with aspects of any other example described to form other examples without losing the effects sought.

[0140] The term 'comprising' is used herein to mean including the identified method blocks or elements, but such blocks or elements do not comprise an exclusive list and the method or apparatus may contain additional blocks or elements.

[0141] The term 'subset' is used herein to refer to a proper subset, such that a subset of a set does not include all elements of the set (ie, at least one of the elements of the set is missing from the subset).

[0142] It should be understood that the above description is given as an example only, and various modifications may be made by those skilled in the art. The above description, examples and data provide a complete description of the structure and use of the exemplary embodiments. Although various embodiments have been described above with a certain degree of particularity or with reference to one or more separate embodiments, those skilled in the art may make various changes to the disclosed embodiments without departing from the scope of this specification.

Claims

1. A method for selecting data for privacy-preserving machine learning in a trusted execution environment, the method comprising: storing training data from the first party; Storing machine learning models; storing criteria from the first party or from another party, the criteria including quality; selecting the training data to select a first portion of the training data to be used to train the machine learning model and selecting a second portion of the training data, wherein the selecting is accomplished at least by calculating a measure of contribution of the training data to performance of the machine learning model using the criterion; as well as The machine learning model is trained using the first portion of the training data such that the second portion of the training data remains private in the trusted execution environment.

2. The method of claim 1, wherein the criteria further comprises one or more of the following: a threshold, a value, a metric, a statistic.

3. A method according to claim 1, wherein the criterion is suitable for: selecting training data based on the likelihood of performance of the machine learning model when trained using the selected training data.

4. The method of claim 1, wherein the criterion is suitable for indicating a performance level of a machine learning model.

5. The method of claim 1, wherein the standard is validation data for validating a machine learning task for which the machine learning model is to be trained.

6. A method according to claim 1, wherein the metric is a Shapley value calculated using a characteristic function, wherein the characteristic function is equal to the performance of the machine learning model when evaluated using the criteria given by participant i after the model has been trained using training data available from all participants in a combination of participants in one possible set S of multiple possible sets of participants.

7. The method of claim 1, wherein there are multiple machine learning models.

8. The method according to claim 1, comprising: Calculate a measure of the contribution of participant l's training data to the performance of the machine learning model, and control access to the machine learning model based on the calculated measure.

9. A method according to claim 8, wherein the measure of the contribution is a Shapley value calculated using a characteristic function, wherein the characteristic function is equal to the performance of the machine learning model plus the sum of the performance of the machine learning model for each individual participant.

10. The method according to claim 9, comprising: storing a plurality of machine learning models, one for each participant, and wherein the metric is a Shapley value calculated using a characteristic function, wherein the characteristic function is equal to the sum of the performance of all of the machine learning models plus the sum of the performance of the machine learning model for each individual participant.

11. An apparatus for selecting data for privacy-preserving machine learning in a trusted execution environment, the apparatus comprising: A memory for storing training data from the first party; The memory stores a machine learning model; The memory stores criteria from the first party or from another party, the criteria comprising quality; a selector configured to select the training data to select a first portion of the training data to be used to train the machine learning model and to select a second portion of the training data, wherein the selecting is accomplished at least by calculating a measure of a contribution of the training data to a performance of the machine learning model using the criterion; as well as A training engine configured to train the machine learning model using the first portion of the training data such that the second portion of the training data remains private in the trusted execution environment.

12. A system for selecting data for privacy-preserving machine learning in a trusted execution environment, the system comprising: at least one processor; Memory, which stores the following items: Training data from the first party, Machine learning models, and standards from the first party or from another party, the standards including quality; a selector implemented on the at least one processor, the selector configured to select the training data to select a first portion of the training data to be used to train the machine learning model and to select a second portion of the training data, wherein the selecting is performed at least by calculating a measure of contribution of the training data to performance of the machine learning model using the criterion; as well as A training engine implemented on the at least one processor, the training engine configured to train the machine learning model using the first portion of the training data such that the second portion of the training data remains private in the trusted execution environment.

13. The system of claim 12, wherein the criteria is adapted to select training data based on the likelihood of performance of the machine learning model when trained using the selected training data.

14. The system of claim 12, wherein the criterion is suitable for indicating a performance level of a machine learning model.

15. The system of claim 12, wherein the standard is validation data for validating a machine learning task for which the machine learning model is to be trained.

16. A system according to claim 12, wherein the metric is a Shapley value calculated using a characteristic function, wherein the characteristic function is equal to the performance of the machine learning model when evaluated using the criteria given by participant i after the model has been trained using training data available from all participants in a combination of participants in one possible set S of multiple possible sets of participants.

17. The system of claim 12, further comprising a plurality of machine learning models, the plurality of machine learning models comprising the machine learning model.

18. The system of claim 12, wherein: The selector is further configured to calculate a measure of the contribution of the training data of party l to the performance of the machine learning model, and The system also includes an access controller configured to control access to the machine learning model based on the calculated metric.

19. A system according to claim 18, wherein the measure of contribution is a Shapley value calculated using a characteristic function, wherein the characteristic function is equal to the performance of the machine learning model plus the sum of the performance of the machine learning model for each individual participant.

20. A system according to claim 19, wherein the memory also stores multiple machine learning models, one machine learning model for each participant, and wherein the metric is a Shapley value calculated using a characteristic function, wherein the characteristic function is equal to the sum of the performance of all of the machine learning models plus the sum of the performance of the machine learning model for each individual participant.

Citation Information

Patent Citations

  • Reward augmented model training

    CN109791631A

  • Document Coding Computer System and Method With Integrated Quality Assurance

    US20140279761A1