An adversarial sample generation method and device, and an electronic device
By generating adversarial patches for 3D object images, the problem of difficulty in generating adversarial examples suitable for 3D image recognition models in existing technologies is solved, thereby improving the robustness of the recognition model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-27
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies struggle to generate adversarial examples suitable for 3D image-based recognition models, resulting in insufficient robustness of these models.
By acquiring a 3D object image of the sample object, determining the coordinate transformation relationship between the calibration image and the 3D object image, generating adversarial patches, and adjusting the pixel positions and colors of the image to generate adversarial examples, until the recognition result is the target object.
The generated adversarial examples can effectively improve the recognition model's error probability in recognizing 3D images, train a recognition model suitable for 3D images, and enhance its robustness.
Smart Images

Figure CN116563902B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine vision technology, and in particular to an adversarial sample generation method, apparatus and electronic device. Background Technology
[0002] Due to limitations in the performance of recognition models, adding specific perturbations to an image to be recognized may cause the model to incorrectly identify objects within the image. These images with added perturbations are referred to as adversarial examples. Adversarial examples are necessary in certain applications, such as recognition model training and performance evaluation.
[0003] Since adversarial examples require the recognition model to produce incorrect recognition results, the higher the accuracy of the recognition model, the more difficult it is to generate recognition results applicable to that model. In related technologies, some image-based recognition models require the recognition of 3D images containing depth information. Because the recognition model can refer to the depth information in the image during recognition, the recognition results are more accurate. However, adversarial examples applicable to these 3D image-based recognition models are therefore more difficult to generate. Summary of the Invention
[0004] The purpose of this application is to provide an adversarial sample generation method, apparatus, and electronic device to effectively improve the robustness of the recognition model. The specific technical solution is as follows:
[0005] In a first aspect of this application, an adversarial example generation method is provided, the method comprising:
[0006] A first stereoscopic object image is obtained by capturing the sample object when a calibration image is set on the target area of the sample object, and the calibration image includes multiple calibration points;
[0007] Determine the coordinate transformation relationship between the image coordinates of each calibration point in the first stereoscopic object image and the image coordinates of each calibration point in the calibration image;
[0008] Based on the coordinate transformation relationship, generate an adversarial patch for the target object;
[0009] A second stereoscopic image of the sample object is obtained as an adversarial sample. The second stereoscopic image is obtained by taking a picture of the sample object when the adversarial patch is set on the target area of the sample object.
[0010] In one possible embodiment, the method further includes:
[0011] The second stereoscopic object image is input into a preset object recognition model to obtain the recognition result output by the preset object recognition model;
[0012] If the identification result is not the target object, return to the step of generating an adversarial patch for the target object based on the coordinate transformation relationship, until the identification result is the target object.
[0013] In one possible embodiment, generating an adversarial patch for the target object based on the coordinate transformation relationship includes:
[0014] Based on the coordinate transformation relationship, the position of each pixel in the target image of the target region of the target object is adjusted, and / or the color of each pixel in the target image is adjusted to obtain the adversarial patch.
[0015] In one possible embodiment, adjusting the position of each pixel in the target image of the target region of the target object according to the coordinate transformation relationship includes:
[0016] Based on the coordinate transformation relationship, the offset method with the highest first score is determined as the target offset method. The first score is positively correlated with the offset similarity, which represents the similarity between the first overlaid image and the third stereoscopic object image of the target object. The first overlaid image is the image obtained by overlaying the first transformed image onto the object image of the sample object. The first transformed image is obtained by transforming the offset image according to the coordinate transformation relationship. The offset image is obtained by offsetting the target image according to the offset method.
[0017] The positions of each pixel in the target image are offset according to the target offset method.
[0018] In one possible embodiment, the first score is positively correlated with smoothness, and the smoothness is negatively correlated with the offset of each pixel in the offset mode.
[0019] In one possible embodiment, adjusting the color of each pixel in the target image according to the coordinate transformation relationship includes:
[0020] Based on the coordinate transformation relationship, the color change method with the highest second score is determined as the target color change method. The color similarity is used to represent the similarity between the second overlay image and the third stereoscopic object image of the target object. The second overlay image is the image obtained by overlaying the second transformed image onto the object image of the sample object. The second transformed image is obtained by transforming the color change image according to the coordinate transformation relationship. The color change image is obtained by color-changing the target image according to the color change method.
[0021] The color of each pixel in the target image is changed according to the target color change method.
[0022] In a second aspect of this application, a method for training a recognition model is provided, the method comprising:
[0023] Obtain adversarial samples generated based on sample objects, wherein the adversarial samples are generated according to the method described above;
[0024] The adversarial sample is input into the original object recognition model to obtain the recognition result output by the original object recognition model;
[0025] Based on the degree of difference between the recognition result and the sample object, the model parameters of the original object recognition model are adjusted to obtain the target recognition model.
[0026] In a third aspect of this application, an adversarial sample generation apparatus is provided, the apparatus comprising:
[0027] The image acquisition module is used to acquire a first stereoscopic object image of the sample object when a calibration image is set on the target area of the sample object, and the calibration image includes multiple calibration points;
[0028] The calibration module is used to determine the coordinate transformation relationship between the image coordinates of each calibration point in the first stereoscopic object image and the image coordinates of each calibration point in the calibration image;
[0029] The anti-disturbance module is used to generate anti-patch for the target object based on the coordinate transformation relationship;
[0030] The sample generation module is used to acquire a second stereoscopic object image of the sample object as an adversarial sample. The second stereoscopic object image is obtained by taking a picture of the sample object when the adversarial patch is set on the target area of the sample object.
[0031] In one possible embodiment, the anti-disturbance module is further configured to input the second stereoscopic object image into a preset object recognition model to obtain the recognition result output by the preset object recognition model;
[0032] If the identification result is not the target object, return to the step of generating an adversarial patch for the target object based on the coordinate transformation relationship, until the identification result is the target object.
[0033] In one possible embodiment, the anti-perturbation module generates an anti-patch against the target object based on the coordinate transformation relationship, including:
[0034] Based on the coordinate transformation relationship, the position of each pixel in the target image of the target region of the target object is adjusted, and / or the color of each pixel in the target image is adjusted to obtain the adversarial patch.
[0035] In one possible embodiment, the anti-perturbation module adjusts the position of each pixel in the target image of the target region of the target object according to the coordinate transformation relationship, including:
[0036] Based on the coordinate transformation relationship, the offset method with the highest first score is determined as the target offset method. The first score is positively correlated with the offset similarity, which represents the similarity between the first overlaid image and the third stereoscopic object image of the target object. The first overlaid image is the image obtained by overlaying the first transformed image onto the object image of the sample object. The first transformed image is obtained by transforming the offset image according to the coordinate transformation relationship. The offset image is obtained by offsetting the target image according to the offset method.
[0037] The positions of each pixel in the target image are offset according to the target offset method.
[0038] In one possible embodiment, the first score is positively correlated with smoothness, and the smoothness is negatively correlated with the offset of each pixel in the offset mode.
[0039] In one possible embodiment, the anti-perturbation module adjusts the color of each pixel in the target image according to the coordinate transformation relationship, including:
[0040] Based on the coordinate transformation relationship, the color change method with the highest second score is determined as the target color change method. The color similarity is used to represent the similarity between the second overlay image and the third stereoscopic object image of the target object. The second overlay image is the image obtained by overlaying the second transformed image onto the object image of the sample object. The second transformed image is obtained by transforming the color change image according to the coordinate transformation relationship. The color change image is obtained by color-changing the target image according to the color change method.
[0041] The color of each pixel in the target image is changed according to the target color change method.
[0042] In a fourth aspect of this application, a recognition model training apparatus is provided, the apparatus comprising:
[0043] The sample acquisition module is used to acquire adversarial samples generated based on sample objects, wherein the adversarial samples are generated according to the above method;
[0044] The result prediction module is used to input the adversarial sample into the original object recognition model to obtain the recognition result output by the original object recognition model.
[0045] The parameter adjustment module is used to adjust the model parameters of the original object recognition model according to the degree of difference between the recognition result and the sample object, so as to obtain the target recognition model.
[0046] In a fifth aspect of this application, an electronic device is provided, comprising:
[0047] Memory, used to store computer programs;
[0048] When a processor executes a program stored in memory, it implements the steps of the method described in either the first or second aspect above.
[0049] In a fourth aspect of this application, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements any of the steps of the method described in the first aspect above.
[0050] Beneficial effects of the embodiments in this application:
[0051] The adversarial sample generation method, apparatus, and electronic device provided in this application can determine the coordinate transformation relationship between the image coordinate system of the calibration image and the image coordinate system of the first stereoscopic object image by calibrating the calibration image. The image coordinate system of the calibration image is a planar image coordinate system. Since the first stereoscopic object image has a certain stereoscopic structure, the image coordinate system of the first stereoscopic object image can reflect this stereoscopic structure. The coordinate transformation relationship between the two image coordinate systems can map the planar image into an image with this certain stereoscopic structure. Therefore, the adversarial patch generated according to this coordinate transformation can better fit the stereoscopic structure, so that the actual effect of the generated adversarial patch is consistent with the expected effect. Therefore, when the second stereoscopic object image is captured with an adversarial patch on the target area of the sample object, the probability of the recognition model making a mistake is higher, that is, the second stereoscopic object image is a more accurate adversarial sample. Therefore, by using this embodiment, adversarial samples suitable for recognition models based on three-dimensional images can be effectively trained.
[0052] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.
[0054] Figure 1 A flowchart illustrating an adversarial example generation method provided in this application embodiment;
[0055] Figure 2 A flowchart illustrating a method for adjusting the position of pixels in a target image according to an embodiment of this application;
[0056] Figure 3 A flowchart illustrating a method for adjusting the color of pixels in a target image according to an embodiment of this application;
[0057] Figure 4a A schematic diagram illustrating an adversarial example generation method for a 3D face recognition model provided in an embodiment of this application;
[0058] Figure 4b A schematic diagram illustrating the generation process of adversarial patches in the adversarial sample generation method for 3D face recognition models provided in the embodiments of this application;
[0059] Figure 5 A flowchart illustrating the recognition model training method provided in this application embodiment;
[0060] Figure 6 A schematic diagram of the adversarial sample generation device provided in an embodiment of this application;
[0061] Figure 7 A schematic diagram of the structure of the recognition model training device provided in the embodiments of this application;
[0062] Figure 8 This is a schematic diagram of the mechanism of an electronic device provided in an embodiment of this application. Detailed Implementation
[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.
[0064] To more clearly illustrate the adversarial example generation method provided in this disclosure, the following will provide an exemplary description of one possible application scenario of the adversarial example generation method provided in this disclosure. The following example is only one possible application scenario of the adversarial example generation method provided in this disclosure. In other possible embodiments, the adversarial example generation method provided in this disclosure can also be applied to other possible application scenarios. The following example does not impose any limitations on this.
[0065] In some applications, devices need to identify objects in images. For example, in facial recognition, the identity of a person needs to be determined based on the captured facial image. In autonomous driving, the vehicle's controller needs to identify vehicles, pedestrians, traffic signs, etc. Related technologies can capture images including objects and input these images into a pre-trained object recognition model to identify the objects present in the image.
[0066] Because different images have different image features, they occupy different positions in the feature space, which is defined by image features. Furthermore, since the recognition model identifies objects within images based on their image features, the mapping relationship between images and recognition results learned through machine learning can be viewed as one (or more) interfaces (or boundaries) in the feature space. Images located on one side of the interface will be identified as the first object, while images on the other side will be identified as the second object.
[0067] Assuming the first object is the first person and the second object is the second person, and assuming the first image is located on the side of the interface that would be identified as the first person, and the object in the first image is the first person, then a relatively subtle perturbation, which is perceptible to the human eye, can be superimposed on the first image to obtain the second image. Since the perturbation information is subtle to the human eye, a person observing the second image will still identify the object in the second image as the first person. However, in the feature space, the position of the second image is offset relative to the first image. If the first image is located near the interface, the second image may be located on the side of the interface that would be identified as the second person in the feature space, causing the recognition model to incorrectly identify the object in the second image as the second person.
[0068] It is evident that in some cases, images that could originally be correctly identified by the recognition model cannot be correctly identified by the recognition model after being perturbed. In the following text, images with perturbations are referred to as adversarial examples, and the perturbations are referred to as adversarial perturbations.
[0069] Some individuals can generate adversarial examples by overlaying adversarial perturbations onto specific images, causing recognition models to output incorrect recognition results. This can prevent services that rely on the recognition model's output from functioning correctly; this process is referred to as an adversarial attack. To enhance the defense capabilities of services against adversarial attacks, adversarial examples can be used to train the network model. To improve the effectiveness of adversarial training, comprehensive and accurate adversarial examples are required.
[0070] In related technologies, a printing device can be used to print an adversarial patch representing an adversarial perturbation as a physical image, and the physical image can be placed on the surface of a sample object, and an image of the sample object at this time can be captured as an adversarial sample.
[0071] Some image-based recognition models can recognize 3D images containing depth information. Since the recognition model can refer to the depth information in the image during recognition, the recognition result is more accurate and can accurately recognize adversarial examples generated in the above way. Therefore, adversarial examples generated in the above way are difficult to be effectively applied to these 3D image-based recognition models.
[0072] Based on this, embodiments of this application provide an adversarial example generation method, such as... Figure 1 As shown, it includes:
[0073] S101, when a calibration image is set on the target area of the sample object, a first stereoscopic object image of the sample object is captured, wherein the calibration image includes multiple calibration points.
[0074] S102, determine the coordinate transformation relationship between the image coordinates of each calibration point in the first stereoscopic object image and the image coordinates of each calibration point in the calibration image.
[0075] S103 generates an adversarial patch for the target object based on the coordinate transformation relationship.
[0076] S104, acquire the second stereoscopic object image as an adversarial example.
[0077] By using this embodiment, the coordinate transformation relationship between the image coordinate system of the calibration image and the image coordinate system of the first stereoscopic object image can be determined through the calibration image. The image coordinate system of the calibration image is a planar image coordinate system. Since the first stereoscopic object image has a certain three-dimensional structure, its image coordinate system can reflect this three-dimensional structure. The coordinate transformation relationship between the two image coordinate systems can map the planar image into an image with this certain three-dimensional structure. Therefore, the adversarial patch generated according to this coordinate transformation can better fit the three-dimensional structure, thus making the actual effect of the generated adversarial patch consistent with the expected effect. Therefore, the second stereoscopic object image captured when the adversarial patch is set on the target area of the sample object can successfully make the recognition model have a higher probability of misidentification, that is, the second stereoscopic object image is a more accurate adversarial sample. Therefore, by using this embodiment, adversarial samples suitable for recognition models based on three-dimensional images can be effectively trained.
[0078] The following will explain S101-S104 respectively:
[0079] In S101, depending on the application scenario, the sample object can be of different types, including but not limited to people, vehicles, etc., and the target area is the area where a certain part, part, or component of the sample object is located. For example, taking a person as the sample object, the target area can be the area where the nose is located or the area where the eyes are located. Taking a vehicle as the sample object, the target area can be the area where the car logo is located or the area where a certain number or letter of the license plate is located.
[0080] Setting a calibration image on the target area means using the calibration image to occlude the target area. The occlusion method can vary depending on the sample object. For example, taking a person as the sample object, the person can wear a 3D mask with the calibration image affixed to it. The 3D mask occludes the target area of the sample object but does not occlude other areas besides the target area. Taking a vehicle as the sample object, the calibration image can be affixed to the vehicle. In this article, occlusion refers to the obstruction of the optical path between the target area and the image acquisition device. For example, using a calibration image to occlude the target area means using the calibration image to obstruct the optical path between the target area and the image acquisition device used to capture the first 3D object image.
[0081] The calibration points in the calibration image should be as clearly distinguishable as possible from the non-calibration point areas in the calibration image. For example, in one possible implementation, the calibration image is a white-background image containing multiple black dots, which are the calibration points of the calibration image. In another possible embodiment, the calibration image is as follows: Figure 2 The image shown is a black and white grid, with the corner points of each grid serving as calibration points.
[0082] In S102, the calibration image itself is a two-dimensional planar image. Therefore, the image coordinates of each calibration point in the calibration image can be regarded as the image coordinates of each calibration point in the planar coordinate system. However, the target area of the sample object has a certain three-dimensional structure. Therefore, the calibration points set in the calibration image of the target area will also have this three-dimensional structure. Therefore, the determined coordinate transformation relationship can reflect the projection of the three-dimensional structure of the target area onto the plane.
[0083] Assume the coordinates of the same calibration point in the calibration image are The coordinates in the first stereoscopic image are The coordinate transformation relationship can be expressed as follows:
[0084]
[0085] In S103, the target object and the sample object are objects of the same category, and it is expected that the recognition model will incorrectly identify the sample object in the image as the target object. For example, suppose it is expected that the recognition model will incorrectly identify the first person as the second person, then the first person is the sample object and the second person is the target object.
[0086] Adversarial patches are obtained by superimposing perturbations on the object image of the target object. The superimposed perturbations can vary depending on the application scenario, as illustrated below, and will not be repeated here.
[0087] In S104, an adversarial patch is applied to the target area of the sample object in the second stereoscopic image. Applying an adversarial patch to the target area means using the patch to occlude the target area. For example, if the sample object is a person, this could be achieved by having the person wear a stereoscopic mask with the adversarial patch applied. The mask occludes the target area of the sample object but not other areas. If the sample object is a vehicle, the adversarial patch could be applied to the target area on the vehicle.
[0088] The second stereoscopic image can be captured by the execution subject of the adversarial sample generation method provided in this application, or it can be captured by other devices besides the execution subject and sent to the execution subject. The second stereoscopic image can also be generated by other means besides capturing, such as by image synthesis technology. This embodiment does not impose any restrictions on this.
[0089] As analyzed in S102 above, since the determined coordinate transformation relationship can reflect the projection of the three-dimensional structure of the target area onto the plane, the adversarial patch generated according to the coordinate transformation can better fit the three-dimensional structure, so that the actual effect of the generated adversarial patch is consistent with the expected effect. Therefore, when the sample object has an adversarial patch on the target area, the second three-dimensional object image captured can successfully make the recognition model have a higher probability of misidentification, that is, the second three-dimensional object image is a more accurate adversarial sample.
[0090] To ensure the generated adversarial patch is sufficiently accurate, in one possible embodiment, after obtaining the second stereoscopic image, the second stereoscopic image is input into a preset object recognition model to obtain the recognition result output by the preset object recognition model. If the recognition result is not the target object, the process returns to step S103 above to regenerate a new adversarial patch until the recognition result output by the preset object recognition model is the target object.
[0091] The following will provide an example of how to generate adversarial patches based on coordinate transformation relationships in the aforementioned S103:
[0092] In one possible embodiment, the aforementioned S103 is implemented in the following way:
[0093] Based on coordinate transformation relationships, adjust the position of each pixel in the target image of the target region of the target object, and / or adjust the color of each pixel in the target image to obtain the adversarial patch.
[0094] Specifically, adjusting the position of each pixel in the target image of the target region of the target object involves superimposing deformation perturbations on the target image, while the color of each pixel in the target image of the target region of the target object involves superimposing color perturbations on the target image. Because the color of the planar adversarial patch remains close to the color distribution of the target region of the target object after superimposing deformation and / or color perturbations, the generated modified adversarial patch exhibits stronger transferability.
[0095] The position can be adjusted in the following ways: Figure 2 As shown, it includes:
[0096] S201. Based on the coordinate transformation relationship, determine the offset method with the highest first score, and use it as the target offset method.
[0097] The first score is positively correlated with the offset similarity, which represents the similarity between the first overlaid image and the third stereoscopic object image of the target object. The first overlaid image is the image obtained by overlaying the first transformed image onto the object image of the sample object. The first transformed image is obtained by transforming the offset image according to the coordinate transformation relationship. The offset image is obtained by offsetting the target image according to the offset method.
[0098] For ease of description, let's assume the target image is denoted as δ, the offset method is represented by the function f(·), the coordinate transformation relationship is represented by the function M(·), and the object image of the sample object is denoted as X. train Let the third stereo image be denoted as T, then the aforementioned offset image is f(δ), the first transformed image is M(f(δ)), and the first superimposed image is X. train +M(f(δ)), therefore the offset similarity is expressed by the following equation:
[0099] L adv (X train ,T,f)=similarity(X train +M(f(δ)),T)
[0100] Among them, L adv (X train Let T, f) represent the offset similarity of offset method f(·), and similarity(·) be used to calculate the similarity between image features of two images. It is understood that the magnitude of similarity in this paper refers to the degree of similarity represented by similarity, rather than the numerical value of similarity. Depending on the method of calculating similarity, the degree of similarity represented by similarity and the numerical value of similarity can be positively correlated or negatively correlated. For example, in one possible embodiment, the similarity between two image features can be obtained by calculating the Euclidean distance between them. The smaller the calculated Euclidean distance, the greater the degree of similarity represented by similarity. In another possible embodiment, the similarity between two image features can also be obtained by calculating the cosine distance between them. The larger the calculated cosine distance, the greater the degree of similarity represented by similarity.
[0101] S202, offset the position of each pixel in the target image according to the target offset method.
[0102] By using this embodiment, the first stereoscopic object image can be made as similar as possible to the third stereoscopic object image after the adversarial patch is superimposed. The third stereoscopic object image is the stereoscopic object image of the target object. The more similar the first stereoscopic object image is to the third stereoscopic object image after the adversarial patch is superimposed, the higher the probability that the first stereoscopic object image will successfully cause the recognition model to misidentify after the adversarial patch is superimposed, that is, the more accurate the generated adversarial patch is.
[0103] In addition to being positively correlated with offset similarity, the aforementioned first score can also be correlated with other factors in other possible embodiments. For example, in one possible embodiment, the first score is positively correlated with smoothness, and smoothness is negatively correlated with the offset of each pixel in the offset mode.
[0104] Smoothness is determined as follows: For every two adjacent pixels, the difference between the offsets of those two pixels in the offset method is calculated, and the smoothness is determined based on the calculated difference. Smoothness is positively correlated with the degree of difference. Therefore, smoothness can be expressed by the following formula:
[0105]
[0106] Among them, L flow (f) represents the offset similarity of offset method f, where p takes values for any pixel, q takes values for any pixel adjacent to p, and Δu (P) Let Δu be the horizontal offset of p in offset mode f. (q) Let Δv be the horizontal offset of q in offset mode f. (p) Δv is the vertical offset of p in offset mode f. (q) This represents the vertical offset of q in offset mode f.
[0107] The offset method f is identified by a 2*H*W dimensional matrix, where H is the pixel height of the planar adversarial patch and W is the pixel width of the planar adversarial patch. Each value in offset method f represents the horizontal or vertical offset of the corresponding pixel in the planar adversarial patch.
[0108] In this embodiment, the determined target offset can be expressed as follows:
[0109]
[0110] Among them, f * For target offset mode, argmax f L adv (X train T, f) refers to the state that enables L adv (X train The largest offset method f is T, f).
[0111] The ways to adjust the color are as follows Figure 3 As shown, it includes:
[0112] S301. Based on the coordinate transformation relationship, determine the second highest-scoring color change method as the target color change method.
[0113] Color similarity is used to represent the similarity between the second overlay image and the third stereoscopic object image of the target object. The second overlay image is the image obtained by overlaying the second transformed image onto the object image of the sample object. The second transformed image is obtained by transforming the color change image according to the coordinate transformation relationship. The color change image is obtained by changing the color of the target image according to the color change method.
[0114] For ease of description, let's assume the target image is denoted as δ. v And the offset method is based on the function t hsv2rgb The coordinate transformation relationship is represented by the function M(·), and the object image of the sample object is denoted as X. train Let the third stereo image be denoted as T, then the aforementioned color change image is t. hsv2rgb (δ v The first transformed image is M(t). hsv2rgb (δ v The first superimposed image is X. train +M(t hsv2rgb (δ v Therefore, color similarity is represented by the following formula:
[0115] L adv (X train ,T,δ)=similarity(X train +M(t hsv2rgb (δ v )), T)
[0116] Among them, L adv (X train T, δ) represent the brightness variation mode t hsv2rgb (·) color similarity.
[0117] It's understandable that if the pixel positions of the target image have been adjusted before adjusting the pixel colors, then when adjusting the pixel colors, the target image refers to the target image whose pixel positions have already been adjusted. Conversely, if the pixel positions of the target image have not been adjusted before adjusting the pixel colors, then when adjusting the pixel colors, the target image refers to the original target image.
[0118] S302, according to the target color change method, change the color of each pixel in the target image.
[0119] It is understood that color includes hue, saturation, and brightness, and adjusting color can refer to adjusting one or more of hue, saturation, and brightness. For example, in one possible embodiment, only brightness is adjusted without adjusting hue and saturation; in another possible embodiment, only brightness and hue are adjusted without adjusting saturation; and in yet another possible embodiment, hue, saturation, and brightness are adjusted.
[0120] For embodiments that adjust only brightness without adjusting hue, in order to avoid changing the color of a pixel when adjusting its brightness, the pixel values of each pixel in the planar adversarial patch can be mapped to HSV (a color space), and only the V (luminance) component of the pixel value of each pixel can be adjusted. The adjusted planar adversarial patch can then be mapped to RGB (another color space) to achieve the change of the color of each pixel in the planar adversarial patch.
[0121] To more clearly illustrate the adversarial example generation method provided in the embodiments of this application, the following will provide an exemplary description in conjunction with specific application scenarios:
[0122] Suppose we need to generate an adversarial patch to cause a face recognition model to incorrectly identify a sample person as a target person. In this example, the sample person is the aforementioned sample object, and the target person is the aforementioned target object. Then, we generate a 3D mask with a similar facial structure to the target person's face.
[0123] See Figure 4a Assuming the target area mentioned above is the area where the nose is located, then the following settings are made on the area where the nose is located in the 3D mask: Figure 4a The calibration image shown in this example represents the aforementioned calibration points, where each grid vertex is a calibration point. The sample personnel were photographed while wearing the stereoscopic mask, resulting in the image shown below. Figure 4a The first 3D object image shown. It is understandable that... Figure 4a The image shown only contains the image information of the first stereoscopic object image, which also includes depth information. A coordinate transformation relationship is established based on the image coordinates of each grid vertex in the first stereoscopic object image and the coordinates of each grid vertex in the calibration image.
[0124] Obtain an image of the area containing the nose of the target person, as the target image. Then, based on coordinate transformation, adjust the position of each pixel in the target image, and based on the coordinate transformation, adjust the color of each pixel in the repositioned target image, resulting in the image shown below. Figure 4a The example shows an adversarial patch. The process of generating an adversarial patch is as follows: Figure 4b As shown, for information on adjusting the position and color, please refer to the previous sections. Figure 2 , Figure 3 The relevant explanations will not be repeated here.
[0125] The adversarial patch was printed as a physical image and placed on the area where the nose of the 3D mask would be located. The position of the adversarial patch should coincide as closely as possible with the position of the previously calibrated image. The sample personnel were photographed while wearing the 3D mask with the adversarial patch, and the results were as follows: Figure 4a The image of the second 3D object shown.
[0126] After generating adversarial examples, the generated adversarial examples can be used to train the recognition model, or the performance of the recognition model can be tested.
[0127] By generating adversarial examples using the aforementioned adversarial example generation method and training the recognition model using the generated adversarial examples, the system's security performance is enhanced when applied to 3D face recognition. This enhances the system's ability to resist 3D face recognition attacks and improves the security of the face recognition system.
[0128] The training process of the recognition model will be explained below. See [link / reference] Figure 5 , Figure 5 The diagram shown is a flowchart of a recognition model training method provided in this application embodiment. This recognition model training method can be applied to any electronic device with recognition model training capabilities. Furthermore, the executing entity of the recognition model training method provided in this application embodiment and the executing entity of the aforementioned adversarial example generation method can be different electronic devices or the same electronic device; this embodiment does not impose any limitations in this regard. Figure 5 As shown, it includes:
[0129] S501, Obtain adversarial examples generated based on sample objects.
[0130] Among them, the adversarial sample is generated according to any of the aforementioned adversarial sample generation methods.
[0131] S502, input the adversarial sample into the original object recognition model, and obtain the recognition result output by the original object recognition model.
[0132] The original object recognition model can be either an untrained recognition model or a trained recognition model. Since the adversarial sample is a three-dimensional object image, the original object recognition model should be a recognition model that recognizes three-dimensional images.
[0133] S503, based on the degree of difference between the recognition result and the sample object, adjust the model parameters of the original object recognition model to obtain the target recognition model.
[0134] Since adversarial examples are 3D images of the target object, if the original object recognition model can accurately identify the adversarial example, the recognition result should be the target object. However, because adversarial examples contain adversarial perturbations, the original recognition model has a certain probability of incorrectly identifying the adversarial example as the target object, resulting in a certain degree of difference between the recognition result and the target object. Adjusting the model parameters of the original object recognition model according to this degree of difference can make the resulting target recognition model more accurately identify the adversarial example as the target object, that is, avoid the target recognition model being interfered with by adversarial perturbations. Therefore, this embodiment can generate a target recognition model that can effectively resist adversarial interference through adversarial training, thereby achieving more accurate recognition.
[0135] The original recognition model can be a trained recognition model or an untrained recognition model. For example, in one possible embodiment, the original recognition model is a recognition model trained with non-adversarial samples. In another possible embodiment, the original recognition model is an untrained recognition model designed based on user experience.
[0136] The adjustment method can vary depending on the application scenario. For example, a loss function can be constructed based on the degree of difference between the recognition result and the sample object, and the model parameters of the original object recognition model can be adjusted in the direction of gradient descent of the loss function until the convergence of the original object recognition model reaches a preset convergence threshold, or the number of adjustments reaches a preset number threshold. The original object recognition model at this time is then used as the target recognition model.
[0137] See Figure 6 , Figure 6 The diagram shown is a structural schematic of an adversarial sample generation device provided in an embodiment of this application, which may include:
[0138] Image acquisition module 601 is used to acquire a first stereoscopic object image of the sample object when a calibration image is set on the target area of the sample object, the calibration image including multiple calibration points;
[0139] The calibration module 602 is used to determine the coordinate transformation relationship between the image coordinates of each calibration point in the first stereoscopic object image and the image coordinates of each calibration point in the calibration image;
[0140] The anti-disturbance module 603 is used to generate an anti-patch for the target object based on the coordinate transformation relationship;
[0141] The sample generation module 604 is used to acquire a second stereoscopic image of the sample object as an adversarial sample, wherein the target region of the sample object in the second stereoscopic image is provided with the adversarial patch.
[0142] In one possible embodiment, the anti-disturbance module is further configured to input the second stereoscopic object image into a preset object recognition model to obtain the recognition result output by the preset object recognition model;
[0143] If the identification result is not the target object, return to the step of generating an adversarial patch for the target object based on the coordinate transformation relationship, until the identification result is the target object.
[0144] In one possible embodiment, the anti-perturbation module generates an anti-patch against the target object based on the coordinate transformation relationship, including:
[0145] Based on the coordinate transformation relationship, the position of each pixel in the target image of the target region of the target object is adjusted, and / or the color of each pixel in the target image is adjusted to obtain the adversarial patch.
[0146] In one possible embodiment, the anti-perturbation module adjusts the position of each pixel in the target image of the target region of the target object according to the coordinate transformation relationship, including:
[0147] Based on the coordinate transformation relationship, the offset method with the highest first score is determined as the target offset method. The first score is positively correlated with the offset similarity, which represents the similarity between the first overlaid image and the third stereoscopic object image of the target object. The first overlaid image is the image obtained by overlaying the first transformed image onto the object image of the sample object. The first transformed image is obtained by transforming the offset image according to the coordinate transformation relationship. The offset image is obtained by offsetting the target image according to the offset method.
[0148] The positions of each pixel in the target image are offset according to the target offset method.
[0149] In one possible embodiment, the first score is positively correlated with smoothness, and the smoothness is negatively correlated with the offset of each pixel in the offset mode.
[0150] In one possible embodiment, the anti-perturbation module adjusts the color of each pixel in the target image according to the coordinate transformation relationship, including:
[0151] Based on the coordinate transformation relationship, the color change method with the highest second score is determined as the target color change method. The color similarity is used to represent the similarity between the second overlay image and the third stereoscopic object image of the target object. The second overlay image is the image obtained by overlaying the second transformed image onto the object image of the sample object. The second transformed image is obtained by transforming the color change image according to the coordinate transformation relationship. The color change image is obtained by color-changing the target image according to the color change method.
[0152] The color of each pixel in the target image is changed according to the target color change method.
[0153] See Figure 7 , Figure 7 The diagram shown is a structural schematic of a recognition model training device provided in an embodiment of this application, which may include:
[0154] The sample acquisition module 701 is used to acquire adversarial samples generated based on sample objects, wherein the adversarial samples are generated according to any of the adversarial sample generation methods described above.
[0155] The result prediction module 702 is used to input the adversarial sample into the original object recognition model to obtain the recognition result output by the original object recognition model;
[0156] The parameter adjustment module 703 is used to adjust the model parameters of the original object recognition model according to the degree of difference between the recognition result and the sample object, so as to obtain the target recognition model.
[0157] This application also provides an electronic device, such as... Figure 8 As shown, it includes:
[0158] Memory 801 is used to store computer programs;
[0159] When processor 802 executes a program stored in memory 801, it performs the following steps:
[0160] A first stereoscopic object image is obtained by capturing the sample object when a calibration image is set on the target area of the sample object, and the calibration image includes multiple calibration points;
[0161] Determine the coordinate transformation relationship between the image coordinates of each calibration point in the first stereoscopic object image and the image coordinates of each calibration point in the calibration image;
[0162] Based on the coordinate transformation relationship, generate an adversarial patch for the target object;
[0163] A second stereoscopic object image obtained from the sample object is used as an adversarial sample, wherein the target region of the sample object in the second stereoscopic image is provided with the adversarial patch.
[0164] Alternatively, the following steps can be implemented:
[0165] Obtain adversarial samples generated based on sample objects, wherein the adversarial samples are generated according to any of the adversarial sample generation methods described above;
[0166] The adversarial sample is input into the original object recognition model to obtain the recognition result output by the original object recognition model;
[0167] Based on the degree of difference between the recognition result and the sample object, the model parameters of the original object recognition model are adjusted to obtain the target recognition model.
[0168] The memory mentioned in the aforementioned electronic device may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0169] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0170] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described adversarial example generation methods or recognition model training methods.
[0171] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the adversarial sample generation methods or recognition model training methods in the above embodiments.
[0172] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0173] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0174] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0175] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.
Claims
1. An adversarial sample generation method, characterized in that, The method comprises: acquiring a first stereoscopic object image of a sample object, wherein a calibration image is arranged on a target region of the sample object, the calibration image comprises a plurality of calibration points, wherein the sample object is a sample person, and the calibration image arranged on the target region of the sample object means that the sample person wears a stereoscopic mask with the calibration image pasted thereon, the stereoscopic mask covers the target region of the sample object, and other regions except the target region are not covered; determining a coordinate conversion relationship between image coordinates of each of the calibration points in the first stereoscopic object image and image coordinates of each of the calibration points in the calibration image; generating an adversarial patch for a target object according to the coordinate conversion relationship, wherein the adversarial patch is obtained by superimposing a disturbance on an object image of the target object; acquiring a second stereoscopic object image of the sample object as an adversarial sample, wherein the target region of the sample object in the second stereoscopic object image is provided with the adversarial patch.
2. The method of claim 1, wherein, The method further comprises: inputting the second stereoscopic object image into a preset object recognition model to obtain an identification result output by the preset object recognition model; if the identification result is not the target object, returning to execute the step of generating an adversarial patch for a target object according to the coordinate conversion relationship until the identification result is the target object.
3. The method of claim 1, wherein, The step of generating an adversarial patch for a target object according to the coordinate conversion relationship comprises: adjusting positions of each pixel point in a target image of a target region of the target object and / or adjusting colors of each pixel point in the target image according to the coordinate conversion relationship to obtain an adversarial patch.
4. The method of claim 3, wherein, The step of adjusting positions of each pixel point in a target image of a target region of the target object according to the coordinate conversion relationship comprises: determining a first highest-score offset mode as a target offset mode according to the coordinate conversion relationship, wherein the first score is positively correlated with offset similarity, the offset similarity is used to represent similarity between a first superimposed image and a third stereoscopic object image of the target object, the first superimposed image is obtained by superimposing a first conversion image on an object image of the sample object, the first conversion image is obtained by converting an offset image according to the coordinate conversion relationship, and the offset image is obtained by offsetting the target image according to the offset mode; offsetting positions of each pixel point in the target image according to the target offset mode.
5. The method of claim 4, wherein, The first score is positively correlated with smoothness, and the smoothness is negatively correlated with offset amounts of each pixel point in the offset mode.
6. The method of claim 3, wherein, The step of adjusting colors of each pixel point in the target image according to the coordinate conversion relationship comprises: According to the coordinate conversion relationship, a second highest score color change mode is determined as a target color change mode, wherein the color similarity is used to represent the similarity between a second superimposed image and a third stereoscopic object image of the target object, the second superimposed image is an image obtained by superimposing a second conversion image on an object image of the sample object, the second conversion image is obtained by converting a color change image according to the coordinate conversion relationship, and the color change image is obtained by performing color change on the target image according to the color change mode; According to the target color change mode, the color of each pixel point in the target image is changed.
7. A face recognition method, characterized by, The method comprises: An adversarial sample generated based on a sample object is obtained, and the adversarial sample is generated according to the method in any one of claims 1-6; The adversarial sample is input into an original object recognition model to obtain a recognition result output by the original object recognition model; According to the difference between the recognition result and the sample object, the model parameters of the original object recognition model are adjusted to obtain a target recognition model; Three-dimensional face recognition is performed by using the target recognition model.
8. An adversarial sample generation apparatus, comprising: The device comprises: An image acquisition module is configured to acquire a first stereoscopic object image of a sample object with a calibration image arranged on a target region of the sample object, the calibration image comprising a plurality of calibration points, wherein the sample object is a sample person, and the calibration image arranged on the target region means that the sample person wears a stereoscopic mask with the calibration image pasted thereon, the stereoscopic mask covers the target region of the sample object, and does not cover other regions except the target region; A calibration module is configured to determine a coordinate conversion relationship between image coordinates of each calibration point in the first stereoscopic object image and image coordinates of each calibration point in the calibration image; An adversarial disturbance module is configured to generate an adversarial patch for a target object according to the coordinate conversion relationship, wherein the adversarial patch is obtained by superimposing a disturbance on an object image of the target object; A sample generation module is configured to acquire a second stereoscopic object image of the sample object as an adversarial sample, and the target region of the sample object in the second stereoscopic object image is provided with the adversarial patch. 9.A recognition model training apparatus, characterized by comprising: The device comprises: A sample acquisition module is configured to acquire an adversarial sample generated based on a sample object, and the adversarial sample is generated according to the method in any one of claims 1-6; A result prediction module is configured to input the adversarial sample into an original object recognition model to obtain a recognition result output by the original object recognition model; A parameter adjustment module is configured to adjust model parameters of the original object recognition model according to the difference between the recognition result and the sample object to obtain a target recognition model.
10. An electronic device, comprising: It comprises: A memory is configured to store a computer program; A processor is configured to execute the program stored in the memory to implement the method steps in any one of claims 1-6 or 7.
Citation Information
Patent Citations
Adversarial patch generation method and device
CN111626925A
Image processing method, device and equipment
CN113436279A