Method and apparatus for protecting privacy information of image sample set
By adjusting the pixel values of the selected images in the image sample set, it approaches the target image in the image recognition model, thereby generating protected image samples, solving the problem of privacy information protection of image sample sets and achieving effective privacy protection and model training effects.
Patent Information
- Application Number
- CN202111415199.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-11-25
AI Technical Summary
The prior art is difficult to effectively protect the privacy information in the image sample set, especially when the image sample set is leaked or open source, an attacker may use the correspondence between the face image and the tag to attack, threatening user privacy.
By determining the image sample to be protected as a target sample, determining an image whose label is different from the target label is used as the selected image, and adjusting the pixel value of the selected image with the pre-trained image recognition model as the target to obtain the adjusted image. Set the target label to adjust the image's label to generate protected image samples for forming the protected image sample set.
It realizes effective protection of the privacy information of the image sample set, preventing attackers from using the image sample set to attack, and ensuring that users of the protected image sample set can train a model with good results.
Smart Images

Figure CN114091104B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of artificial intelligence, and more particularly, to a method and apparatus for protecting privacy information of an image sample set. Background Art
[0002] With the continuous development of computer software and artificial intelligence, the application of machine learning models is becoming more and more extensive. For example, machine learning models are widely used in the field of image processing, and training a model with good performance for processing images requires a large number of image samples. At present, many open source image sample sets on the Internet are open. There are also many corporate or personal image sample sets that are only allowed to be used internally due to privacy or confidentiality. Image sample sets, especially face image sample sets, may contain a large amount of user privacy information. When the data is leaked or open source, the privacy information contained in it may be used maliciously by attackers. For example, attackers can conduct correlation attacks on the correspondence between leaked face images and labels. For example, through AI (Artificial Intelligence, artificial intelligence) face replacement to generate fake videos corresponding to face images, in order to defraud the relatives of the user corresponding to the face image, which greatly threatens user privacy. This type of attack is mainly caused by the attacker's successful acquisition of the correspondence between face images and labels. Based on this, how to protect image sample sets has important practical significance and value. Summary of the invention
[0003] The embodiments of this specification describe a method and device for protecting the privacy information of an image sample set. The method determines the image sample to be protected as the target sample, and determines the image with a label different from the target label as the selected image. When adjusting the pixel value of the selected image, the image recognition model is required to process the selected image close to the processing result of the target image. Therefore, the obtained adjusted image is close to the target image for the image recognition model. Therefore, it can be ensured that the model trained using the protected image sample set is close to the model trained using the image sample set to be protected. From the perspective of human vision, the adjusted image still looks like an image similar to the selected image, and is different from the target image. Based on this, the protected image sample can be used to replace the target sample. In this way, the correspondence between the target image and the target label in the target sample is protected, and the privacy information is protected, and the user of the protected image sample set can be guaranteed to train a model with good results.
[0004] According to a first aspect, a method for protecting the privacy information of an image sample set is provided, comprising: determining an image sample to be protected in the image sample set to be protected as a target sample, wherein the target sample includes a target image and a target label; determining an image whose label is different from the target label as a selected image; adjusting the pixel value of the selected image with the goal of making the processing result of the selected image by a pre-trained image recognition model approach the processing result for the target image to obtain an adjusted image, wherein the image recognition model is trained using the image sample set to be protected; setting the target label as the label of the adjusted image, obtaining a protected image sample including the adjusted image and the target label, for forming a protected image sample set.
[0005] In one embodiment, the determining of the image having a label different from the target label as the selected image comprises: selecting an image having a label different from the target label as the selected image from the to-be-protected image sample set.
[0006] In one embodiment, before adjusting the pixel value of the selected image, the method further includes: in response to determining that the model structure used by the target user corresponding to the protected image sample set is known, using the image sample set to be protected, performing model training based on the model structure to obtain the image recognition model.
[0007] In one embodiment, the processing result is the output vector of the intermediate layer of the image recognition model; the pixel value of the selected image is adjusted to obtain the adjusted image, including: determining the distance between the first output vector for the selected image and the second output vector for the target image; adjusting the pixel value of the selected image with the goal of minimizing the distance.
[0008] In one embodiment, the distance includes one of the following: Euclidean distance, Manhattan distance, and infinite norm of a difference vector.
[0009] In one embodiment, adjusting the pixel value of the selected image with the goal of minimizing the distance includes: determining the gradient of the distance relative to the pixel value; and performing a predetermined number of pixel value adjustments along the direction of gradient descent with a predetermined step length.
[0010] In one embodiment, before adjusting the pixel value of the selected image, the method further includes: in response to determining that the model structure used by the target user corresponding to the protected image sample set is unknown, using the image sample set to be protected, performing model training based on multiple preset predetermined model structures to obtain multiple image recognition models.
[0011] In one embodiment, the processing result is the output vector of the intermediate layer of the multiple image recognition models; the pixel value of the selected image is adjusted to obtain the adjusted image, including: determining the weighted results of the distances between the output vectors of the multiple image recognition models for the selected image and the output vectors for the target image; adjusting the pixel value of the selected image with the goal of minimizing the weighted results.
[0012] In one embodiment, the method further includes: determining a plurality of selected images for the target image to generate a plurality of adjusted images; and setting the labels of the plurality of adjusted images as the target label to obtain a plurality of protected image samples.
[0013] In one embodiment, the image recognition model includes a softmax layer, and the intermediate layer is a previous layer of the softmax layer.
[0014] According to a second aspect, a device for protecting the privacy information of an image sample set is provided, comprising: a first determination unit, configured to determine an image sample to be protected in the image sample set to be protected as a target sample, wherein the target sample includes a target image and a target label; a second determination unit, configured to determine an image whose label is different from the target label as a selected image; an adjustment unit, configured to adjust the pixel value of the selected image with the goal of making the processing result of the selected image by a pre-trained image recognition model approach the processing result for the target image, so as to obtain an adjusted image, wherein the image recognition model is trained using the image sample set to be protected; a generation unit, configured to set the target label as the label of the adjusted image, so as to obtain a protected image sample including the adjusted image and the target label, so as to form a protected image sample set.
[0015] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in any implementation manner in the first aspect.
[0016] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, the method described in any implementation manner in the first aspect is implemented.
[0017] According to the method and device for protecting the privacy information of an image sample set provided by the embodiments of this specification, first, the image sample to be protected is determined as a target sample, and the target sample includes a target image and a target label. After that, an image with a label different from the target label is determined as a selected image. Then, with the goal that the processing result of the pre-trained image recognition model for the selected image is close to the processing result for the target image, the pixel value of the selected image is adjusted to obtain an adjusted image. Finally, the target label is set to the label of the adjusted image, and a protected image sample including the adjusted image and the target label is obtained. Since when adjusting the pixel value of the selected image, the processing result of the image recognition model for the selected image is required to be close to the processing result for the target image, the obtained adjusted image is close to the target image for the image recognition model. Therefore, it can be ensured that the effect of the model trained using the protected image sample set is close to that of the model trained using the image sample set to be protected. From the perspective of human vision, the adjusted image still looks like an image similar to the selected image, and is different from the target image. Based on this, protected image samples can be used to replace target samples. This not only protects the correspondence between the target image and the target label in the target sample, thus protecting privacy information, but also ensures that users of the protected image sample set can train a model with good results. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram showing an application scenario in which the embodiments of this specification can be applied;
[0019] Figure 2 A schematic flow chart of a method for protecting privacy information of an image sample set according to an embodiment is shown;
[0020] Figure 3 A schematic diagram showing the composition of the information contained in the adjustment image is shown;
[0021] Figure 4 A schematic diagram of a process for training an image recognition model is shown;
[0022] Figure 5 A schematic diagram showing the effect of generating multiple protected image samples for the same target image;
[0023] Figure 6 A schematic block diagram of an apparatus for protecting privacy information of an image sample set according to an embodiment is shown. DETAILED DESCRIPTION
[0024] The technical solution provided in this specification is further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant inventions, rather than to limit the inventions. It should also be noted that, for ease of description, only the parts related to the relevant inventions are shown in the accompanying drawings. It should be noted that, in the absence of conflict, the embodiments of this specification and the features in the embodiments can be combined with each other.
[0025] As mentioned above, attackers may use the privacy information contained in open source image sample sets or leaked image sample sets for internal use only to launch attacks. To prevent the above from happening, some owners of image sample sets encrypt the entire or part of the image samples in the image sample set (for example, images or labels) to prevent the samples from being used maliciously. However, the existence of encryption is obvious and easy to be detected by attackers. In addition, since encrypting image samples may affect the availability of samples, encryption is particularly unsuitable for protecting open source image sample sets.
[0026] To this end, the embodiments of this specification provide a method for protecting the privacy information of an image sample set, so as to achieve the protection of the image sample set. For example, the image sample to be protected is a face image, and the label is a name. Figure 1 FIG. 1 is a schematic diagram showing an application scenario in which the embodiments of this specification can be applied. Figure 1As shown, the image sample set 101 to be protected includes multiple image samples to be protected, and each image sample to be protected may include a face image and a label. In this example, each image sample to be protected in the image sample set 101 to be protected may be processed. Taking the image sample to be protected currently to be processed as the face image of Zhang San, and the label is "Zhang San" as an example, first, the image sample to be protected in the image sample set 101 to be protected is determined as a target sample, wherein the target sample includes a target image 102 and a target label "Zhang San". In this example, the target image 102 is the face image of Zhang San. Secondly, a face image with a label different from "Zhang San" is determined as a selected image 103. In this example, the selected image 103 is the face image of Li Si, and its label is "Li Si". Then, with the goal that the processing result of the image recognition model 104 for the selected image 103 is close to the processing result for the target image 102, the pixel value of the selected image 103 is adjusted to obtain an adjusted image 105. Here, the image recognition model 104 can be trained using the image sample set 101 to be protected. Finally, the target label "Zhang San" is set as the label of the adjusted image 105, and the protected image sample 106 including the adjusted image 105 and the target label "Zhang San" is obtained, which is used to form the protected image sample set. Since when adjusting the pixel value of the selected image 103, the processing result of the image recognition model 104 for the selected image 103 is required to be close to the processing result for the target image 102, the obtained adjusted image 105 is close to the target image 102 for the image recognition model 104. Therefore, it can be ensured that the model trained using the protected image sample set is close to the model trained using the image sample set 101 to be protected. From the perspective of human vision, the adjusted image 105 still looks like the face image of Li Si (the selected image), which is different from the face image of Zhang San (the target image). Based on this, protected image samples can be used to replace target samples. In this way, the correspondence between Zhang San’s face image and the target label “Zhang San” is protected, privacy information is protected, and users of the protected image sample set can be guaranteed to train a model with good results.
[0027] Continue to see Figure 2 , Figure 2 FIG. 1 is a flow chart of a method for protecting the privacy information of an image sample set according to an embodiment. It is understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Figure 2 As shown, the method for protecting the privacy information of the image sample set may include the following steps:
[0028] Step 201: determine an image sample to be protected in a set of image samples to be protected as a target sample.
[0029] In this embodiment, the image samples to be protected in the image sample set to be protected may be determined as target samples, wherein the target samples may include target images and target labels. Here, each image sample to be protected in the above-mentioned image sample set to be protected may include an image and a label. For example, the image may be a natural image, an object image, an animal image, a medical image, a face image, a fingerprint image, or other images.
[0030] Step 202: Determine an image with a label different from the target label as a selected image.
[0031] In this embodiment, an image whose label is different from the target label can be determined as the selected image. Here, any image whose label is different from the target label can be determined as the selected image. In practice, in order to make the generated protected image sample more realistic and less likely to be detected by attackers, an image of a type similar to that of the target image can be selected as the selected image. For example, when the target image is a face image of a certain person, the face image of another person can be selected as the selected image. For another example, when the target image is a fingerprint image of a certain person, the fingerprint image of another person can be selected as the selected image.
[0032] In some optional implementations, images with labels different from the target label can be selected from the above-mentioned image sample set to be protected as selected images. Selecting selected images from the image sample set to be protected can make the generated protected image samples more realistic and less likely to be detected by attackers.
[0033] Step 203 , with the goal of making the processing result of the pre-trained image recognition model for the selected image close to the processing result for the target image, the pixel value of the selected image is adjusted to obtain an adjusted image.
[0034] In this embodiment, an image recognition model can be pre-trained, wherein the image recognition model can be trained using the above-mentioned image sample set to be protected. For example, the image recognition model can be a deep neural network (DNN). In this way, the pixel value of the selected image can be adjusted with the goal of making the processing result of the image recognition model for the selected image approach the processing result for the target image, thereby obtaining an adjusted image. Here, the above-mentioned processing result can be various processing results. For example, the image recognition model can include an input layer, multiple intermediate layers, and an output layer. The above-mentioned processing result can be the output result of each layer, for example, it can be the output vector of a certain intermediate layer, or the probability vector of each classification finally output by the output layer. In this example, the process of generating the adjusted image is equivalent to adding relevant knowledge related to the target image to the selected image. Figure 3 As shown, taking face image as an example, Figure 3A schematic diagram showing the composition of the information contained in the adjusted image. Figure 3 In the example shown, the adjusted image 301 includes information in the selected image 302 and a disturbance 303, wherein the disturbance 303 includes information of the target image. Thus, the adjusted image 301 can include relevant knowledge related to the target image.
[0035] Step 204: Set the target tag as the tag of the adjusted image, and obtain a protected image sample including the adjusted image and the target tag, which is used to form a protected image sample set.
[0036] In this embodiment, the target label can be set to the label of the adjusted image obtained in step 203, so that a protected sample image including the adjusted image and the target label can be obtained. The protected image sample set can be formed by obtaining each protected image sample for each image sample to be protected in the image sample set to be protected. In the application scenario of protecting the open source image sample set, the owner of the image sample set to be protected can deploy the protected image sample set online as an open source image sample set for users to use. In the application scenario of internal use only, the protected image sample set can be used as an image sample set for internal use only. In this way, the user of the image sample set can train his own machine learning model based on the protected image sample set.
[0037] In some optional implementations, before adjusting the pixel values of the selected images, the method for protecting the privacy information of the image sample set may further include a process of training an image recognition model. For example, the image recognition model may be trained based on whether the model structure used by the target user corresponding to the protected image sample set is known. Figure 4 As shown, Figure 4 A flow chart of a method for training an image recognition model is shown, which may specifically include the following steps:
[0038] Step 401: determine whether the model structure used by the target user corresponding to the protected image sample set is known.
[0039] In this implementation, the target user corresponding to the protected image sample set may refer to a user who uses the protected image sample set to train a machine learning model. In practice, in some application scenarios, the model structure used by the target user may be known. For example, in an application scenario where the protected image sample set is for internal use only, since the target user may be known, the model structure used by the target user may also be known. In other application scenarios, the model structure used by the target user may be unknown. For example, in an application scenario where the protected image sample set is an open source image sample set, since the target user is unknown, the model structure used by the target user may also be unknown.
[0040] Step 402, in response to determining that the model structure used by the target user corresponding to the protected image sample set is known, the image sample set to be protected is used to perform model training based on the model structure to obtain an image recognition model.
[0041] Step 403, in response to determining that the model structure used by the target user corresponding to the protected image sample set is unknown, the image sample set to be protected is used to perform model training based on multiple preset predetermined model structures to obtain multiple image recognition models.
[0042] In this implementation, since the model structure used by the target user is unknown, it is necessary to use a preset predetermined model structure. Here, the above-mentioned predetermined model structure can be a model structure determined by the technician in various ways. For example, several model structures commonly used at this stage can be determined as the predetermined model structure. In this way, a sample set of images to be protected can be used to perform model training based on multiple predetermined model structures to obtain multiple image recognition models. For example, assuming that there are three predetermined model structures, namely model structure A, model structure B and model structure C, the sample set of images to be protected can be used to perform model training based on model structure A, model structure B and model structure C, so that three image recognition models can be obtained.
[0043] In some optional implementations, in a scenario where the model structure used by the target user is known, the trained model may include an image recognition model, and the processing results of the image recognition model for the selected image and the processing results for the target image may be the output vector of the intermediate layer of the image recognition model. For example, the image recognition model may include a multi-layer structure such as an input layer, multiple intermediate layers, and an output layer, and the above processing results may be the output vector of any intermediate layer. Optionally, the output layer of the above image recognition model may be set to a softmax layer, and the above intermediate layer may refer to the previous layer of the softmax layer. At this time, in the above step 204, adjusting the pixel value of the above selected image to obtain the adjusted image may be specifically performed as follows:
[0044] First, the distance between the first output vector for the selected image and the second output vector for the target image is determined.
[0045] In this implementation, an output vector of a certain intermediate layer of the image recognition model for the selected image can be determined as a first output vector, and an output vector of the intermediate layer of the image recognition model for the target image can be determined as a second output vector. Then, the distance between the first output vector and the second output vector is calculated. Optionally, the distance can include one of the following: Euclidean distance, Manhattan distance, infinite-order norm of a difference vector, etc.
[0046] Then, the pixel values of the selected image are adjusted with the goal of minimizing the above distance.
[0047] In this implementation, various optimization algorithms, such as projected gradient descent (PGD), can be used to adjust the pixel value of the selected image with the goal of minimizing the distance between the first output vector and the second output vector. Optionally, the gradient of the distance between the first output vector and the second output vector relative to the pixel value of the selected image can be first determined. Then, with a predetermined step length, a predetermined number of pixel value adjustments are performed in the direction of gradient descent. Through this implementation, the pixel value of the selected image can be adjusted.
[0048] For example, suppose the image sample set to be protected is Where N represents the size of the sample set; the selected image is x selected , x selected The label is y original ; The target image is x target , x target The label is y target , where y target ≠y original ; then the adjusted image can be x modified , can be obtained by the following formula:
[0049]
[0050] Where W can represent the image dimension, f can represent the mapping from the input to the output of the intermediate layer in the image recognition model, f(x) corresponds to the first output vector, and f(x target ) corresponds to the aforementioned second output vector, d(·) can represent the distance measure between two vectors. In this example, x selected Initialize x. In addition, in this example, the infinite norm of the difference vector can also be used as the distance, x selected After the pixel values are adjusted, the adjusted image x is obtained. modified , that is, d(f(x modified ), f(x target ))=||f(x modified )-f(x target )|| ∞ Through this implementation, a protected image sample corresponding to a target sample can be generated in a scenario where the model structure used by the target user is known.
[0051] In some other optional implementations, in a scenario where the model structure used by the target user is unknown, the trained model may include multiple image recognition models, and the processing results of the image recognition models for the selected image and the processing results for the target image may refer to the output vectors of the intermediate layers of the multiple image recognition models. Optionally, the output layers of the multiple image recognition models may be set to a softmax layer, and the intermediate layer may refer to the previous layer of the softmax layer.
[0052] At this time, in the above step 204, adjusting the pixel value of the selected image to obtain the adjusted image can be specifically performed as follows:
[0053] First, weighted results of the distances between the output vectors of the selected image and the output vectors of the target image of the plurality of image recognition models are determined.
[0054] In this implementation, for multiple image recognition models, the distance between the output vector of the intermediate layer of each image recognition model for the selected image and the output vector for the target image can be determined first to obtain multiple distances. Then, the weighted results of the multiple distances are calculated. Here, the weighted result can be a weighted sum, a weighted average, etc.
[0055] Then, the pixel values of the selected image are adjusted with the goal of minimizing the weighted result.
[0056] For example, assuming there are M image recognition models, adjust the image x modified It can be obtained by the following formula:
[0057]
[0058] In this case, you can use x selected Initialize x. Through this implementation, a protected image sample corresponding to the target sample can be generated in a scenario where the model structure used by the target user is unknown.
[0059] In some optional implementations, the method for protecting the privacy information of the image sample set may further include the following:
[0060] First, a plurality of selected images are determined for a target image, and a plurality of adjusted images are generated.
[0061] In this implementation, for the same target image, multiple selected images may be selected, and step 203 may be repeated multiple times, thereby generating multiple adjusted images.
[0062] Then, the labels of the multiple adjusted images are set as the target label, thereby obtaining multiple protected image samples.
[0063] In this implementation, for the same target image, multiple protected image samples can be generated. Figure 5 As shown, Figure 5 The schematic diagram shows the effect of generating multiple protected image samples for the same target image. Figure 5 In the example of an object image shown, from the perspective of human eyes, the target image 501 shows a boat, and the target label is "boat". The three adjusted images corresponding to the three generated protected image samples show a cat, an alpha and a horse respectively, and the labels of these three adjusted images are all the target label "boat". Figure 5 In the example of another facial image shown, from the perspective of human eyes, the target image 502 shows a man’s face, and the target label is “1133”. The three adjusted images corresponding to the three generated protected image samples show another man’s face, a little boy’s face and a woman’s face respectively. The labels of these three adjusted images are all the target label “1133”.
[0064] Through this implementation, multiple protected image samples can contain information about target samples. In other words, the knowledge about the correspondence between the target image and the target label in the target sample is reflected in multiple protected image samples to enhance the relevant knowledge, so that the image sample set user can better learn the relevant knowledge when training his own machine learning model, making the model trained by the image sample set user more accurate.
[0065] Reviewing the above process, in the embodiments of the present specification, an adjusted image is generated by adjusting the pixel value of the selected image, and the label of the adjusted image is set as the target label, thereby generating a protected image sample. Since when adjusting the pixel value of the selected image, the image recognition model is required to have the processing result for the selected image be close to the processing result for the target image, the obtained adjusted image is close to the target image for the image recognition model. Thus, it can be ensured that the effect of the model trained using the protected image sample set is close to that of the model trained using the image sample set to be protected. From the perspective of human vision, the adjusted image still looks like an image similar to the selected image, and is different from the target image. Based on this, the protected image sample can be used to replace the target sample, so that the correspondence between the target image and the target label in the target sample is protected, the protection of privacy information is achieved, and the user of the protected image sample set can be guaranteed to train a model with good results.
[0066] According to another aspect of the embodiment, a device for protecting the privacy information of an image sample set is provided. The device for protecting the privacy information of an image sample set can be deployed in any device, platform or device cluster with computing and processing capabilities.
[0067] Figure 6 FIG. 1 is a schematic block diagram of an apparatus for protecting privacy information of an image sample set according to an embodiment. Figure 6 As shown, the device 600 for protecting the privacy information of the image sample set includes: a first determination unit 601, configured to determine the image sample to be protected in the image sample set to be protected as a target sample, wherein the target sample includes a target image and a target label; a second determination unit 602, configured to determine an image with a label different from the target label as a selected image; an adjustment unit 603, configured to adjust the pixel value of the selected image with the goal of making the processing result of the pre-trained image recognition model for the selected image approach the processing result for the target image, to obtain an adjusted image, wherein the image recognition model is trained using the image sample set to be protected; a generation unit 604, configured to set the target label as the label of the adjusted image, to obtain a protected image sample including the adjusted image and the target label, for forming a protected image sample set.
[0068] In some optional implementations of this embodiment, the second determining unit 602 is further configured to: select an image with a label different from the target label from the to-be-protected image sample set as a selected image.
[0069] In some optional implementations of this embodiment, the above-mentioned device 600 may also include: a first model training unit (not shown in the figure), configured to respond to determining that the model structure used by the target user corresponding to the above-mentioned protected image sample set is known, use the above-mentioned image sample set to be protected, perform model training based on the above-mentioned model structure, and obtain the above-mentioned image recognition model.
[0070] In some optional implementations of the present embodiment, the above processing result is the output vector of the intermediate layer of the above image recognition model; the above adjustment unit 603 is further configured to: determine the distance between the first output vector for the above selected image and the second output vector for the above target image; and adjust the pixel value of the above selected image with the goal of minimizing the above distance.
[0071] In some optional implementations of this embodiment, the above distance includes one of the following: Euclidean distance, Manhattan distance, and infinite norm of a difference vector.
[0072] In some optional implementations of the present embodiment, adjusting the pixel value of the selected image with the goal of minimizing the distance includes: determining the gradient of the distance relative to the pixel value; and performing a predetermined number of pixel value adjustments in a direction of gradient descent with a predetermined step length.
[0073] In some optional implementations of the present embodiment, the above-mentioned device 600 also includes a second model training unit (not shown in the figure), which is configured to, in response to determining that the model structure used by the target user corresponding to the above-mentioned protected image sample set is unknown, use the above-mentioned image sample set to be protected to perform model training based on multiple preset predetermined model structures to obtain multiple image recognition models.
[0074] In some optional implementations of the present embodiment, the above processing result is the output vector of the intermediate layer of the above multiple image recognition models; the above adjustment unit 603 is further configured to: determine the weighted results of the distances between the output vectors of the above multiple image recognition models for the above selected images and the output vectors for the above target images; and adjust the pixel value of the above selected image with the goal of minimizing the above weighted results.
[0075] In some optional implementations of the present embodiment, the above-mentioned device 600 also includes: a third determination unit (not shown in the figure), configured to determine multiple selected images for the above-mentioned target image and generate multiple adjusted images; a setting unit (not shown in the figure), configured to set the labels of the above-mentioned multiple adjusted images as target labels to obtain multiple protected image samples.
[0076] In some optional implementations of this embodiment, the image recognition model includes a softmax layer, and the intermediate layer is a previous layer of the softmax layer.
[0077] According to another embodiment, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute Figure 2 The method described.
[0078] According to another embodiment, there is also provided a computing device, comprising a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, Figure 2 The method described.
[0079] Those of ordinary skill in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0080] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented by hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0081] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for protecting privacy information of an image sample set, comprising: Determine the image sample to be protected in the image sample set to be protected as a target sample, wherein the target sample includes a target image and a target label; Determine an image with a label different from the target label as a selected image; With the goal of making the processing result of the pre-trained image recognition model for the selected image close to the processing result for the target image, adjusting the pixel value of the selected image to obtain an adjusted image, wherein the image recognition model is trained using the sample set of images to be protected; The target tag is set as the tag of the adjusted image, and a protected image sample including the adjusted image and the target tag is obtained to form a protected image sample set.
2. The method according to claim 1, wherein: The determining an image with a label different from the target label as a selected image includes: An image with a label different from the target label is selected from the to-be-protected image sample set as a selected image.
3. The method according to claim 1, wherein: Before adjusting the pixel value of the selected image, the method further comprises: In response to determining that the model structure used by the target user corresponding to the protected image sample set is known, the image sample set to be protected is used to perform model training based on the model structure to obtain the image recognition model.
4. The method according to claim 3, wherein: The processing result is an output vector of the intermediate layer of the image recognition model; and adjusting the pixel value of the selected image to obtain an adjusted image includes: determining a distance between a first output vector for the selected image and a second output vector for the target image; The pixel values of the selected image are adjusted with the goal of minimizing the distance.
5. The method according to claim 4, wherein: The distance includes one of the following: Euclidean distance, Manhattan distance, and infinite norm of a difference vector.
6. The method according to claim 4, wherein: The step of adjusting the pixel value of the selected image with the goal of minimizing the distance comprises: determining a gradient of the distance relative to the pixel value; A predetermined number of pixel value adjustments are performed along the direction of gradient descent with a predetermined step length.
7. The method according to claim 1, wherein: Before adjusting the pixel value of the selected image, the method further comprises: In response to determining that the model structure used by the target user corresponding to the protected image sample set is unknown, the to-be-protected image sample set is used to perform model training based on a plurality of preset predetermined model structures to obtain a plurality of image recognition models.
8. The method according to claim 7, wherein: The processing result is an output vector of the intermediate layer of the plurality of image recognition models; and the adjusting the pixel value of the selected image to obtain the adjusted image includes: Determine weighted results of distances between output vectors of the selected image and output vectors of the target image respectively of the plurality of image recognition models; The pixel values of the selected image are adjusted with the goal of minimizing the weighted result.
9. The method according to claim 1, wherein: The method further comprises: Determining a plurality of selected images for the target image, and generating a plurality of adjusted images; The labels of the plurality of adjusted images are all set as target labels to obtain a plurality of protected image samples.
10. The method according to claim 4 or 8, wherein: The image recognition model includes a softmax layer, and the intermediate layer is a previous layer of the softmax layer.
11. A device for protecting privacy information of an image sample set, comprising: A first determining unit is configured to determine the image sample to be protected in the image sample set to be protected as a target sample, wherein the target sample includes a target image and a target label; a second determining unit, configured to determine an image with a label different from the target label as a selected image; an adjusting unit, configured to adjust the pixel values of the selected image with the goal of making the processing result of the pre-trained image recognition model for the selected image close to the processing result for the target image, so as to obtain an adjusted image, wherein the image recognition model is trained using the sample set of images to be protected; The generating unit is configured to set the target tag as the tag of the adjusted image, obtain a protected image sample including the adjusted image and the target tag, and form a protected image sample set.
12. The device according to claim 11, wherein: The second determining unit is further configured to: An image with a label different from the target label is selected from the to-be-protected image sample set as a selected image.
13. The device according to claim 11, wherein: The device also includes: The first model training unit is configured to, in response to determining that the model structure used by the target user corresponding to the protected image sample set is known, use the image sample set to be protected to perform model training based on the model structure to obtain the image recognition model.
14. The device according to claim 13, wherein: The processing result is an output vector of the intermediate layer of the image recognition model; the adjustment unit is further configured as follows: determining a distance between a first output vector for the selected image and a second output vector for the target image; The pixel values of the selected image are adjusted with the goal of minimizing the distance.
15. The device according to claim 14, wherein: The distance includes one of the following: Euclidean distance, Manhattan distance, and infinite norm of a difference vector.
16. The device according to claim 14, wherein: The step of adjusting the pixel value of the selected image with the goal of minimizing the distance comprises: determining a gradient of the distance relative to the pixel value; A predetermined number of pixel value adjustments are performed along the direction of gradient descent with a predetermined step length.
17. The device according to claim 11, wherein: The device also includes: The second model training unit is configured to, in response to determining that the model structure used by the target user corresponding to the protected image sample set is unknown, use the image sample set to be protected to perform model training based on multiple preset predetermined model structures to obtain multiple image recognition models.
18. The device according to claim 17, wherein: The processing result is the output vector of the intermediate layer of the multiple image recognition models; the adjustment unit is further configured as follows: Determine weighted results of distances between output vectors of the selected image and output vectors of the target image respectively of the plurality of image recognition models; The pixel values of the selected image are adjusted with the goal of minimizing the weighted result.
19. The device according to claim 11, wherein: The device also includes: a third determining unit, configured to determine a plurality of selected images for the target image and generate a plurality of adjusted images; A setting unit is configured to set the labels of the plurality of adjusted images as target labels to obtain a plurality of protected image samples.
20. The device according to claim 14 or 18, wherein: The image recognition model includes a softmax layer, and the intermediate layer is a previous layer of the softmax layer.
21. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 10.
22. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Image recognition method and device and computer readable storage medium
CN111079833A
Adversarial example-based method and apparatus for protecting private information and electronic device
WO2021098270A1