Methods, devices, apparatuses, and media for protecting data

CN114519209BActive Publication Date: 2026-09-29FACE CUTE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210119229.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-08
Publication Date
2026-09-29
Estimated Expiration
2042-02-08

AI Technical Summary

Technical Problem

目前已经提出了在具有数据的标签信息的第一方与不具有标签信息的第二方之间联合训练模型的技术方案,然而已有技术方案的性能并不理想,因而不能有效地防止敏感信息泄漏

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519209B_ABST
    Figure CN114519209B_ABST
Patent Text Reader

Abstract

According to embodiments of the present disclosure, methods, devices, apparatuses and media for protecting data are provided. The method includes obtaining, by a first device, a feature representation generated by a second device based on sample data according to a second model. The first device has label information for the sample data. The first device and the second device are used for jointly training a first model at the first device and a second model at the second device. The method further includes generating, by the first device, a predicted label for the sample data according to the first model based on the feature representation. The method further includes determining, by the first device, a total loss value for training the first model and the second model based on the feature representation, the label information and the predicted label. In this way, the second device without the label information of the data can be prevented from predicting the label information of the data, so that the data, especially sensitive data, is saved from being leaked.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for protecting data. Background Technology

[0002] With the development of artificial intelligence technology, machine learning technology has been applied to various industries. Processing models for specific functions can be trained based on pre-collected training data. However, some training data may involve user privacy and / or other sensitive data. Therefore, how to collect training data more securely and use it to train processing models has become a research hotspot. Currently, technical solutions have been proposed for jointly training models between a first party with data labeling information and a second party without labeling information; however, the performance of existing solutions is not ideal, and therefore cannot effectively prevent the leakage of sensitive information. Summary of the Invention

[0003] According to an example embodiment of this disclosure, a scheme for protecting data is provided.

[0004] In a first aspect of this disclosure, a method for protecting data is provided. The method includes: acquiring, by a first device, a feature representation generated by a second device based on sample data and according to a second model. The first device has label information for the sample data. The first device and the second device are used to jointly train a first model at the first device and a second model at the second device. The method further includes: generating predicted labels for the sample data by the first device based on the feature representation and according to the first model; and determining a total loss value for training the first model and the second model by the first device based on the feature representation, the label information, and the predicted labels.

[0005] In a second aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the following actions: acquiring a feature representation generated by a second device based on sample data and according to a second model, the electronic device having label information for the sample data, and the electronic device and the second device jointly training a first model at the electronic device and a second model at the second device; generating a predicted label for the sample data based on the feature representation and according to the first model; and determining a total loss value for training the first and second models based on the feature representation, label information, and predicted label.

[0006] In a third aspect of this disclosure, an apparatus for protecting data is provided, the apparatus comprising an acquisition module configured to acquire feature representations generated by a second device based on sample data and according to a second model, the apparatus having label information for the sample data, the apparatus and the second device being used to jointly train a first model at the apparatus and a second model at the second device; a label prediction module configured to generate predicted labels for the sample data based on the feature representations and according to the first model; and a total loss determination module configured to determine a total loss value for training the first model and the second model based on the feature representations, the label information, and the predicted labels.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, which, when executed by a processor, implements the method of the first aspect.

[0008] It should be understood that the content described in this summary section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 A schematic diagram of an example architecture for joint training of models according to some embodiments of the present disclosure is shown;

[0012] Figure 3 A schematic diagram illustrating the results of data protection according to some embodiments of the present disclosure is shown;

[0013] Figure 4 A flowchart of a process for protecting data according to some embodiments of the present disclosure is shown;

[0014] Figure 5 A block diagram of an apparatus for protecting data according to some embodiments of the present disclosure is shown; and

[0015] Figure 6 A block diagram of an apparatus capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0016] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0018] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably. As used in this paper, "determining the parameters of a model" or similar expressions refers to determining the values ​​of the model's parameters (also known as parameter values), including specific values, sets of values, or ranges of values.

[0019] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0020] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating its parameter values ​​until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values ​​of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual inputs based on the trained parameter values ​​to determine the corresponding output.

[0021] In the field of machine learning, training datasets containing large amounts of training data can be used to train predictive models, enabling these models to perform the desired functions. As the volume of data increases, it is typically distributed across different storage locations, such as those belonging to different companies. With increasingly stringent data regulations and growing emphasis on data privacy, traditional centralized machine learning techniques struggle to ensure data privacy protection.

[0022] Distributed machine learning technology has already been proposed. For example, different participants can have their own machine learning models, and they can send model parameters and push data to each other. At the same time, sensitive data, including users' personal information, is stored only at one participant and is not transmitted to other participants.

[0023] For example, in the online advertising industry, media outlets (passive participants) can push advertisements to users. If a user views an advertisement on the media side and subsequently performs actions such as purchasing, registering, or requesting information on the advertiser side (active participant), the advertiser can tag the data associated with that advertisement to indicate the user's actions. The advertiser can use this tagging information to infer information such as sensitive business information, like user base and revenue. Training data can be used to train predictive models, enabling them to make inferences such as user base and revenue scale. However, this tagging information is crucial for advertisers and is sensitive data that needs protection. Advertisers do not want their tagging information to be leaked, for example, they do not want media outlets to steal it. Therefore, it is necessary to prevent the leakage of sensitive data during the acquisition of training data.

[0024] A technical solution has been proposed for jointly training models between a first party with data labeling information and a second party without labeling information. In this paper, the first party with data labeling information is referred to as the "active participant," and the second party without data labeling information is referred to as the "passive participant." In the joint training model scheme (also known as federated learning of models), the data labeling information is stored only at the active participant and is not transmitted to the passive participant. Insensitive data and / or model parameters can be transmitted between the active and passive participants to train the models at the active and passive participants respectively. For example, the active participant can determine the gradient information used to update the prediction model and send this gradient information to the passive participant. The passive participant can then use this gradient information to train the model without acquiring sensitive data, such as labeling information.

[0025] However, attackers can use techniques such as spectral clustering to predict label information from data that lacks label information. To prevent attackers from obtaining label information (e.g., predicting it), and thus to protect sensitive information, several schemes for protecting relational label information have been proposed. For example, the Gradient Depth Leakage (DLG) algorithm and its improved form, iDLG, can be used to protect label information.

[0026] Specifically, the DLG algorithm discovers that passive participants can infer corresponding user label information from the gradients shared by active participants. DLG first generates random dummy training input data and label information, which produces a dummy gradient through forward propagation. By minimizing the difference between the dummy gradient and the true gradient, DLG can infer the training input and its corresponding label information. iDLG is an improvement on DLG. iDLG finds that DLG cannot infer the label information of active participants with high quality. Under the setting of binary classification tasks and using cross-entropy as the loss function, iDLG finds a correlation between label information and the gradient sign of the loss with respect to the last layer's logit. iDLG uses this finding to analyze the gradient of the parameters to infer the label information used for training samples. Both DLG and iDLG require the passive participants to possess the entire model structure of all participants in the federated learning process. This is difficult to achieve in reality.

[0027] Furthermore, solutions have been proposed to add noise to gradient information to prevent the leakage of sensitive information. However, in scenarios where attackers are highly capable, they may be able to predict label information that is very close to the true label information based on the feature representation obtained from the sample data. In such cases, even if adding noise to the gradient information prevents the attacker from predicting the label information through the gradient information, it is still impossible to prevent the attacker from inferring the label information based on the feature representation obtained from the sample data.

[0028] In conclusion, current data protection solutions have many shortcomings and fail to achieve satisfactory results. Therefore, there is a need for more effective methods to prevent the leakage of tag information and protect sensitive data.

[0029] According to embodiments of this disclosure, a scheme for protecting data is provided, aiming to address one or more of the aforementioned problems and other potential problems. In this scheme, a first device (of an active participant) uses a first model to predict labels for sample data based on feature representations generated from sample data by a second model of a second device (of a passive participant). The first device then determines a total loss value based on the feature representation, the true label information of the sample data, and the obtained predicted labels. The first device can then adjust the parameters of the first model at the first device based on the total loss value. Furthermore, the first device can also enable the second device to adjust the parameters of the second model at the second device based on the gradient associated with the total loss value.

[0030] In this scheme, it is not necessary to know the structure of the models of each participant (i.e., the models at each device). The first device and the second device only share the feature representation generated by the second device and the gradient associated with the total loss value, without sharing sensitive information such as label information.

[0031] Furthermore, this scheme adjusts the parameters of the first and / or second models using a total loss value, which is correlated with feature representation, label information, and predicted label. In this way, an appropriate total loss value can be determined by setting a suitable loss function to reduce the correlation between feature representation and label information. This prevents passive participants (or second devices) from predicting the label information of the data based on the feature representation of the sample data, thereby protecting sensitive data.

[0032] In the following text, we will first refer to Figure 1 An example environment is described according to an exemplary implementation of this disclosure.

[0033] Example Environment

[0034] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. Figure 1In environment 100, it is desired to train and apply a first model 140 at a first device 120 and / or a second model 150 at a second device 130. The first device 120 may be an active participant device. In this document, the first device 120 may also be referred to as an active participant device, and the first model 140 may also be referred to as an active participant model. The second device 130 may be a passive participant device. In this document, the second device 130 may also be referred to as a passive participant device, and the second model 150 may also be referred to as a passive participant model. The first device 120 and the second device 130 can interact via a network to receive or send data or messages, etc. In some embodiments, the first model 140 and the second model 150 may be configured to perform the same or different functions or tasks, such as prediction tasks, classification tasks, etc.

[0035] Overall, environment 100 includes a joint training system 110 and optional application systems 160. Figure 1 In the example embodiments described below and in some example embodiments, the joint training system 110 is configured to use training dataset 101 to jointly train the first model 140 and the second model 150 at the first device 120 and the second device 130. Training dataset 101 may include sample data 102-1, 102-2, ..., 102-N and associated label information 104-1, 104-2, ..., 104-N, where N is an integer greater than or equal to 1. For ease of discussion, sample data 102-1, 102-2, ..., 102-N may be collectively referred to as sample data 102, and label information 104-1, 104-2, ..., 104-N may be collectively referred to as label information 104.

[0036] Each sample data 102 may have corresponding tag information 104. The tag information 104 may include sensitive data regarding the user's processing of the sample data 102. In an embodiment where the advertiser is the active participant and the media is the passive participant, after a user clicks on an advertisement on the media side, a processing behavior may occur on the advertiser side. For example, if a user is influenced by an advertisement on the media side and engages in behaviors such as purchasing, registering, or requesting information on the advertiser side, it indicates that a processing behavior (also known as a conversion) has occurred during the advertisement registration process.

[0037] In some embodiments, a positive example label of 1 can be used to indicate that a sample data has undergone the aforementioned transformation behavior, that is, the sample data is a positive example. If no transformation occurs, a negative example label of 0 is used to indicate that the sample data is a negative example. It should be understood that in some embodiments, other methods can also be used to add labels to each sample data. For example, for behaviors such as purchasing, registering, and information demand, labels of 3, 2, and 1 can be added to the sample data respectively, while 0 is used as the label for behaviors that have not undergone transformation. It should be understood that the examples of label information listed above are merely exemplary and are not intended to limit the scope of this disclosure.

[0038] In some embodiments, the tag information 104 described above can only be known by the first device 120. In some embodiments, the first device 120 can use the tag information 104 to infer sensitive business information such as advertisers' user base and revenue scale. The tag information 104 is important information that the active participant (i.e., the first device 120) needs to protect. Therefore, it is not desirable for the second device 130 to obtain the tag information 104. That is, the second device 130 may not have the tag information 104.

[0039] Before joint training, the parameter values ​​of the first model 140 and the second model 150 can be initialized. After joint training, the parameter values ​​of the first model 140 and the second model 150 are updated and determined. After joint training is completed, the trained first model 170 and the trained second model 180 have the trained parameter values. Based on these parameter values, the trained first model 170 and the trained second model 180 can be used for various processing tasks, such as predicting user scale. In some embodiments, the first model 140 and the second model 150 can be any suitable neural network model. For example, in the example of positive and negative label information, the first model 140 and the second model 150 can be models for binary classification. The first model 140 and the second model 150 can each include hidden layers, logit layers, softmax layers, etc.

[0040] Environment 100 may optionally include application system 160. Application system 160 may be configured to perform various processing tasks on source data 162 and / or source label information 164 using a trained first model 170 and a trained second model 180. In some embodiments, the first device 120 or the trained first model 170 has source label information 164. In contrast, the second device 130 or the trained second model 180 does not have source label information 164. Application system 160 may also include other models or operations (not shown) to perform corresponding inference or other tasks.

[0041] exist Figure 1In this context, the joint training system 110 and the application system 160 can be any system with computing capabilities, such as various computing devices / systems, terminal devices, servers, etc. Terminal devices can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. Servers include, but are not limited to, mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0042] It should be understood that Figure 1 The components and arrangements shown in the environment are merely examples, and a computing system suitable for implementing the exemplary embodiments described in this disclosure may include one or more different components, other components, and / or different arrangements. For example, although shown as separate, two or more of the joint training system 110 and the application system 160 may be integrated into the same system or device. Embodiments of this disclosure are not limited in this respect.

[0043] The following will continue to describe example embodiments of joint model training with reference to the accompanying drawings.

[0044] Model training architecture

[0045] Figure 2 A schematic diagram of an example architecture 200 for joint training of models according to some embodiments of the present disclosure is shown. Figure 2 The architecture 200 can be implemented in Figure 1 In the joint training system 110, the various steps in architecture 200 can be implemented by hardware, software, firmware, or any combination thereof. In architecture 200, both the first device 120 and the second device 130 can be used to perform distributed training on the first model 140 and the second model 150.

[0046] Figure 2 The left side shows the processing procedure on the first device 120 side, and Figure 2 The right side illustrates the processing on the second device 130 side. In some embodiments, the second model 150 can encode the sample data 102 to obtain a feature representation 212. This feature representation 212 can be obtained based on various encoders currently known and / or those to be developed in the future. The second device 130 can transmit the feature representation 212 to the first model 140 for subsequent joint training of the first model 140 and the second model 150. It should be understood that, although... Figure 2 Only one sample data 102 is shown, but in some embodiments, a set of sample data or a batch of sample data may be used. In some embodiments, the feature representation for a batch of sample data can be represented as f(X). In this document, the feature representation is also referred to as intermediate layer encoding (e.g., embedding) or intermediate encoding (e.g., embedding).

[0047] In some embodiments, the second device 130 may be aggressive. That is, the second device 130 may attempt to infer or predict the label information 104 using the feature representation 212. For example, the second device 130 may use a spectral attack, such as a 2-cluster outlier detection algorithm, to predict the label information against the feature representation 212.

[0048] In some embodiments, positive sample data can be represented as Negative sample data can be represented as The following can be used to predict labels for feature representation 212:

[0049] | <f(X)-μ F ,v>| (1)

[0050] Where f(X) represents the characteristic representation 212, μ F Let v denote the median of the sample data F (including positive and negative sample data), and let v denote the top singular vector of the covariance matrix of F.

[0051] With a batch of sample data For example, the second device 130 can calculate the empirical median and covariance matrix of the batch of sample data to estimate μ. F and v. Based on the estimated μ F And v, we can use equation (1) to calculate the score for each sample data in the batch | <f(X)-μ F The second device 130 can divide all f(X) into two clusters based on the calculated scores.

[0052] In some embodiments, the second device 130 may also use prior indications or empirical values ​​to determine labels for the two resulting clusters. For example, if the dataset currently in use is imbalanced and predominantly negative, negative examples may be assigned to the larger cluster and positive examples may be assigned to the smaller cluster.

[0053] In some embodiments, the second device 130 can use network structures of arbitrary width and depth to enhance the capabilities of its side model. In some cases, the second device 130 can have strong attack capabilities, enabling it to make f(X) and the label information prediction capabilities of the first device 120 consistent.

[0054] The above describes the process by which a second device 130, acting as an attacker, infers label information 104 based on feature representation 212 in certain situations. In some embodiments, it is undesirable for the second device 130 to infer label information 104 based on feature representation 212. This can be avoided by using a joint training process of a first model 140 and a second model 150, described in detail below. In other words, the leakage of sensitive user information, such as label information, is prevented through the joint training process of the first model 140 and the second model 150 described below.

[0055] like Figure 2 As shown, the first device 120 generates a predicted label 222 for the sample data 102 based on the feature representation 212 and the first model 140. In this paper, the predicted label 222 can be the result of the logit output of the first model 140, and is also denoted as h(f(X)). The first device 120 is configured to determine the total loss value 230 for training the first model 140 and the second model 150 based on the feature representation 212, the label information 104, and the predicted label 222.

[0056] In some embodiments, the first device 120 may determine a first loss based on the tag information 104 and the predicted tag 222. The first loss can be used to represent the degree of difference between the predicted tag 222 and the tag information 104. For example, the cross-entropy between the tag information 104 and the predicted tag 222 can be determined as the first loss. In this document, the first loss can be represented as Lc. In this example, a larger value of the first loss indicates a greater degree of difference between the predicted tag 222 and the tag information 104.

[0057] In some embodiments, the first device 120 may determine a second loss based on the feature representation 212 and the label information 104. The second loss can be used to represent the degree of correlation between the feature representation 212 and the label information 104. For example, in some embodiments, the first device 120 may determine the second loss as the distance correlation coefficient between the feature representation 212 and the label information 104. Determining the second loss using the distance correlation coefficient can be expressed as:

[0058] L d =DCOR(Y,F(X)) (2)

[0059] Where L d Let L represent the second loss, Y represent the label information (104), F(X) represent the feature representation (212), and DCOR() represent the distance correlation coefficient function. It should be understood that any appropriate distance correlation coefficient function can be used to determine the second loss L. d .

[0060] Additional or alternative land may also be used to address the aforementioned second loss L. d The logarithm is taken to obtain the logarithmically processed second loss. In the example above, the larger the value of the second loss, the more relevant feature representation 212 is to label information 104.

[0061] In some embodiments, the first device 120 may determine the total loss value 230 based on the first loss and the second loss. For example, the first device 120 may determine the total loss value 230 as a weighted sum of the first loss and the second loss. For example, the total loss value 230 may be determined using the above equation (3):

[0062] L = L c +α d L d (3)

[0063] Where L (which can also be represented as l) represents the total loss value of 230, L c Indicates the first loss, L d Indicates the second loss, α d It is a value greater than or equal to zero, α d This indicates the weight of the second loss.

[0064] In some embodiments, the weight α of the second loss d It can be a preset value, such as 0.5 or other appropriate values ​​between 0 and 1. Alternatively, the weights α can be dynamically adjusted during joint training based on the results of the training process. d It should be understood that the loss function described above is illustrative, not restrictive.

[0065] In some embodiments, the first device 120 can determine the parameters of the first model 140 by minimizing the total loss value 230. For example, the first device 120 can determine the parameters of the first model 140 by determining the minimum total loss value from the total loss value 230 and at least one other total loss value. The at least one other total loss value is determined by the first device 120 before or after the determination.

[0066] In some embodiments, the first device 120 may perform iterative updates based at least on the loss function of equation (3) to determine the parameter values ​​of the first model 140. For example, training of the first model 140 is complete if the iteration converges or the predetermined number of iterations is reached. In this way, the parameters of the first model 140 can be determined based on the total loss value 230.

[0067] By using the total loss value described above, the correlation between feature representation 212 and label information 104 can be reduced while ensuring the accuracy of the predicted label 222. That is, feature representation 212 and label information 104 are decoupled. In this way, it is possible to prevent a second device 130 or other attackers from inferring or predicting label information 104 based on feature representation 212. Sensitive user information, such as label information, is thus protected.

[0068] After determining the total loss value 230, the first device 120 can also utilize the gradient determination module 250 to determine the gradient 240 based on the total loss value 230. For example, the first device 120 can determine the gradient 240 of the total loss value 230 relative to the feature representation 212. In some embodiments, the gradient 240 can be determined by the following equation (4):

[0069]

[0070] Where g represents the gradient 240, l represents the total loss 230, and f(X) represents the feature representation 212.

[0071] In some embodiments, the first device 120 may send the determined gradient 240 to the second device 130, so that the second device 130 determines the parameters of the second model 150 based on the gradient 240. For example, in response to receiving the gradient 240, the second device 130 may update the parameters of the second model 150 based on the chain rule.

[0072] By utilizing the joint training process described above, the second device 130 can avoid inferring the label information 104 using the feature representation 212. In this way, the technical solution of this disclosure can better protect sensitive data while ensuring the performance of the first and second models. Furthermore, the technical solution of this disclosure does not require knowledge of the structure of each model to perform joint training, thus making it easy to implement.

[0073] Example Results

[0074] Figure 3 A schematic diagram illustrating the results of data protection according to some embodiments of the present disclosure is shown. Result 300 shows that when data is not protected using the techniques of the present disclosure, the second device 130 utilizes attack AUC (Area Under Curve) on different layers, such as spectral attacks. Attack AUC, or simply AUC, can be used to represent the ability of an attacker (e.g., the second device 130) to distinguish between positive and negative labels. An attack AUC close to 0.5 is considered approximately a random guess, meaning that an attack AUC close to 0.5 can be considered non-dangerous, i.e., the data is secure.

[0075] In result 300, curve 302 shows the test AUC (i.e., the AUC of the first device 120), and curves 302, 304, 306, 308, 312, and 314 show the AUC of the attacker (i.e., the second device 130) at layers 5, 4, 3, 2, and 1, respectively. Result 300 shows that the attacker (e.g., the second device 130) has a very strong attack capability. Especially when the attacker has more than 4 layers, the attack capability approaches the model's utility.

[0076] Result 320 illustrates that when data is protected using the techniques disclosed herein, the second device 130 utilizes attack AUCs such as spectral attacks on different layers. In Result 320, the weight α... d It was set to 0.002. In result 320, curve 322 shows the test AUC (i.e., the AUC of the first device 120), and curves 324, 326, 328, 332 and 334 show the layer 5 AUC, layer 4 AUC, layer 3 AUC, layer 2 AUC and layer 1 AUC of the attacker (i.e., the second device 130), respectively.

[0077] Similarly, result 340 illustrates that when data is protected using the techniques of this disclosure, the second device 130 utilizes attack AUCs such as spectral attacks on different layers. In result 340, the weight α... d It was set to 0.005. In result 340, curve 342 shows the test AUC (i.e., the AUC of the first device 120), and curves 344, 346, 348, 352 and 354 show the layer 5 AUC, layer 4 AUC, layer 3 AUC, layer 2 AUC and layer 1 AUC of the attacker (i.e., the second device 130), respectively.

[0078] Results 320 and 340 show that although the attacker (e.g., the second device 130) possesses strong attack capabilities, the data protection scheme disclosed herein can effectively bring the AUC close to the desired value of 0.5. Specifically, result 320 (weight α) d With the value set to 0.002, the AUC of each layer is close to 0.5. This indicates that the data protection method of this scheme is very effective and can effectively prevent the leakage of sensitive data.

[0079] The results 360 show different weights α d The AUC curves are shown below. Curves 362, 364, 366, 368, 372, and 374 in the results show the weights α. dThe AUC curves of the first device 120 with α values ​​of 0.002, 0.005, 0.01, 0.02, 0.05, and 0.1 are the test AUC curves. In some embodiments, a larger α is selected. d This can improve data protection, meaning better data privacy, but the model's utility will decrease. Conversely, choosing a smaller α... d This would compromise data privacy, and the model's utility would be similar to that of a vanilla model, meaning the model would have good utility. Therefore, in some embodiments, this can be mitigated by selecting appropriate weights α. d This allows for a good balance between data privacy and model utility. For example, in some embodiments, the weight α... d It can be selected as a value between 0.002 and 0.005. It should be understood that the weight α... d The possible values ​​are merely illustrative and are not intended to limit the scope of this disclosure.

[0080] The results above demonstrate that using the joint training process described in this disclosure can prevent the second device 130 from inferring the label information 104 using the feature representation 212. In particular, the security of sensitive information can be guaranteed even when the second device 130 is highly aggressive. Furthermore, by appropriately setting the weight of the second loss in the total loss value, sensitive data can be better protected while ensuring the effectiveness of both the first and second models.

[0081] Example process

[0082] Figure 4 A flowchart of a process 400 for protecting data according to some embodiments of the present disclosure is shown. Process 400 may be implemented at the joint training system 110 and / or application system 160.

[0083] At box 410, the first device 120 acquires feature representation 212 generated by the second device 130 based on sample data 102 and according to the second model 150. The first device 120 has label information 104 for the sample data 102. The first device 120 and the second device 130 are used to jointly train the first model 140 at the first device 120 and the second model 150 at the second device 130.

[0084] In some embodiments, the tag information 104 includes sensitive data for the processing of the sample data 102. In some embodiments, the second device 130 does not have the tag information 102.

[0085] At box 420, the first device 120 generates a predicted label 222 for sample data 102 based on feature representation 212 and according to the first model 140. At box 430, the first device 120 determines the total loss value 230 for training the first model 140 and the second model 150 based on feature representation 212, label information 104, and predicted label 222.

[0086] In some embodiments, a first loss can be determined based on label information 104 and predicted label 222. The first loss represents the degree of difference between predicted label 222 and label information 104. In some embodiments, a second loss can be determined based on feature representation 212 and label information 104. The second loss represents the degree of correlation between feature representation 212 and label information 104. For example, in some embodiments, the distance correlation coefficient between feature representation 212 and label information 104 can be determined as the second loss.

[0087] In some embodiments, the total loss value 230 can be determined based on the first loss and the second loss. For example, in some embodiments, the weighted sum of the first loss and the second loss can be determined as the total loss value 230. For example, the total loss value 230 can be determined by equation (3).

[0088] In some embodiments, the first device 120 can determine the parameters of the first model 140 by determining the minimum total loss value from the total loss value and at least one other total loss value. The at least one other total loss value is determined by the first device 120 before or after the determination.

[0089] In some embodiments, the gradient 240 of the total loss value 230 with respect to the feature representation 212 can be determined by the first device 120. For example, the gradient 240 can be determined by using equation (4).

[0090] In some embodiments, the first device 120 may send the determined gradient 240 to the second device 130. The second device 130 may use the gradient 240 to determine the parameters of the second model 150.

[0091] Example devices and equipment

[0092] Figure 5 A block diagram of an apparatus 500 for protecting data according to some embodiments of the present disclosure is shown. The apparatus 500 may be implemented as or included in the joint training system 110 and / or application system 160. The various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0093] As shown in the figure, the device 500 includes a feature representation acquisition module 510, configured to acquire feature representations generated by a second device based on sample data and according to a second model. The device 500 has label information for the sample data. The device and the second device are used to jointly train the first model at the device and the second model at the second device.

[0094] In some embodiments, the tag information includes sensitive data regarding the user's processing results of the sample data. In some embodiments, the second device does not have tag information.

[0095] The apparatus 500 also includes a prediction label generation module 520, configured to generate prediction labels for sample data based on feature representations and according to a first model.

[0096] The device 500 also includes a total loss determination module 530, configured to determine the total loss value used to train the first model and the second model based on feature representation, label information and predicted labels.

[0097] In some embodiments, the total loss determination module 530 includes a first loss determination module configured to determine a first loss based on label information and predicted labels. The first loss represents the degree of difference between the predicted labels and the label information.

[0098] In some embodiments, the total loss determination module 530 further includes a second loss determination module configured to determine a second loss based on feature representation and label information. The second loss represents the degree of correlation between the feature representation and the label information. In some embodiments, the second loss determination module includes a distance correlation coefficient determination module configured to determine the distance correlation coefficient between the feature representation and the label information as the second loss.

[0099] In some embodiments, the total loss value determination module 530, a second total loss value determination module, is configured to determine a total loss value based on a first loss and a second loss. In some embodiments, the second total loss value determination module includes a weighted sum module configured to determine the total loss value by a weighted sum of the first loss and the second loss.

[0100] In some embodiments, the apparatus 500 further includes a minimum total loss value determination module, configured to determine parameters of the first model by determining a minimum total loss value from a total loss value and at least one other total loss value. The at least one other total loss value is determined earlier or later by the first apparatus 120.

[0101] In some embodiments, the apparatus 500 further includes a gradient determination module configured to determine the gradient of the total loss value relative to the feature representation by the first device. In some embodiments, the apparatus 500 further includes a transmission module configured to transmit the gradient to a second device, such that the second device determines the parameters of the second model based on the gradient.

[0102] Figure 6 A block diagram is shown illustrating a computing device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 6 The computing device 600 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 6 The computing device 600 shown can be used to implement Figure 1 The joint training system 110 and / or application system 160.

[0103] like Figure 6 As shown, computing device 600 is in the form of a general-purpose computing device. Components of computing device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage devices 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 600.

[0104] Computing device 600 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 600.

[0105] The computing device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 6As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0106] The communication unit 640 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 600 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0107] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 600 can also communicate as needed with one or more external devices (not shown) via communication unit 640. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 600, or with any device (e.g., network card, modem, etc.) that enables computing device 600 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).

[0108] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0109] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0110] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0111] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0113] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to the technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for protecting data, comprising: The first device acquires feature representations generated by the second device based on sample data and according to the second model. The first device has label information for the sample data. The first device and the second device are used to jointly train the first model at the first device and the second model at the second device. The first model and the second model are configured to perform the same or different functions or tasks. The sample data includes data associated with push information. The label information is determined based on the processing behavior associated with the push information. Based on the feature representation, the first device generates a predicted label for the sample data according to the first model; The first device determines the total loss value for training the first model and the second model based on the feature representation, the label information, and the predicted label. Determining the total loss value includes: determining a first loss based on the label information and the predicted label, where the first loss represents the degree of difference between the predicted label and the label information; determining a second loss based on the feature representation and the label information, where the second loss represents the degree of correlation between the feature representation and the label information; and determining the total loss value based on the first loss and the second loss. The gradient of the total loss value relative to the feature representation is determined by the first device; and The gradient is sent to the second device so that the second device determines the parameters of the second model based on the gradient.

2. The method of claim 1, wherein determining the total loss value comprises: The weighted sum of the first loss and the second loss is determined as the total loss value.

3. The method of claim 1, wherein determining the second loss comprises: The distance correlation coefficient between the feature representation and the label information is determined as the second loss.

4. The method according to claim 1, further comprising: The parameters of the first model are determined by identifying the minimum total loss value from the total loss value and at least one other total loss value, which is determined by the first device either before or after the determination.

5. The method according to claim 1, wherein the tag information includes sensitive data regarding the user's processing of the sample data.

6. The method of claim 1, wherein the second device does not have the tag information.

7. An electronic device, comprising: At least one processing unit; as well as At least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the following actions: The electronic device acquires feature representations generated by the second device based on sample data and according to the second model. The electronic device has label information for the sample data. The electronic device and the second device are used to jointly train a first model at the electronic device and a second model at the second device. The first model and the second model are configured to perform the same or different functions or tasks. The sample data includes data associated with push information. The label information is determined based on processing behavior associated with the push information. Based on the feature representation, a predicted label for the sample data is generated according to the first model; Based on the feature representation, the label information, and the predicted label, a total loss value for training the first model and the second model is determined, wherein determining the total loss value includes: determining a first loss based on the label information and the predicted label, the first loss representing the degree of difference between the predicted label and the label information; determining a second loss based on the feature representation and the label information, the second loss representing the degree of correlation between the feature representation and the label information; and determining the total loss value based on the first loss and the second loss. The electronic device determines the gradient of the total loss value relative to the feature representation; and The gradient is sent to the second device so that the second device determines the parameters of the second model based on the gradient.

8. The electronic device of claim 7, wherein determining the total loss value comprises: The weighted sum of the first loss and the second loss is determined as the total loss value.

9. The electronic device of claim 7, wherein determining the second loss comprises: The distance correlation coefficient between the feature representation and the label information is determined as the second loss.

10. The electronic device of claim 7, wherein the action further comprises: The parameters of the first model are determined by identifying the minimum total loss value from the total loss value and at least one other total loss value, which is determined by the electronic device either previously or subsequently.

11. The electronic device of claim 7, wherein the tag information includes sensitive data regarding the user's processing of the sample data.

12. The electronic device of claim 7, wherein the second device does not have the tag information.

13. An apparatus for protecting data, comprising: The acquisition module is configured to acquire feature representations generated by a second device based on sample data and according to a second model. The device has label information for the sample data. The device and the second device are used to jointly train a first model at the device and a second model at the second device. The first model and the second model are configured to perform the same or different functions or tasks. The sample data includes data associated with push information. The label information is determined based on processing behavior associated with the push information. The label prediction module is configured to generate predicted labels for the sample data based on the feature representation and according to the first model. The total loss determination module is configured to determine the total loss value used to train the first model and the second model based on the feature representation, the label information, and the predicted label, including: determining a first loss based on the label information and the predicted label, wherein the first loss represents the degree of difference between the predicted label and the label information; determining a second loss based on the feature representation and the label information, wherein the second loss represents the degree of correlation between the feature representation and the label information; and determining the total loss value based on the first loss and the second loss. A gradient determination module is configured to determine, by a first device, the gradient of the total loss value relative to the feature representation; and The sending module is configured to send the gradient to the second device so that the second device determines the parameters of the second model based on the gradient.

14. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Neural network training method and device

    CN113705769A