Disturbance image creating method and device
By optimizing the similarity loss of face recognition results and the information loss from compressing digital pixels to projected pixels, a perturbed image is generated, which solves the problem of information loss when converting the projection light source from the digital domain to the physical domain and improves the effectiveness and accuracy of the perturbation image.
Patent Information
- Application Number
- CN202411932510.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-09-23
AI Technical Summary
In the face recognition optical projection countermeasure technology based on the projector-camera system, there is a large information loss when the digital image is converted into projected pixels, which affects the effectiveness and accuracy of the perturbed image.
A perturbed image is generated by optimizing the similarity loss of face recognition results and the information loss of digital pixel compression to projected pixels. The effectiveness and accuracy of the perturbed image are improved by using mask image synthesis and gradient vector update.
It improves the effectiveness and accuracy of perturbed images in face recognition systems, reduces information loss, and enhances the effectiveness of adversarial attacks.
Smart Images

Figure CN120689911A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of neural network technology, and in particular to a method and device for creating a disturbed image. Background Art
[0002] Currently, in face recognition optical projection adversarial technology based on projector-camera systems, the perturbed image calculation process utilizes a transformation matrix that optimizes the color and position differences between the digital image and the projected image. During actual adversarial attack testing, the resolution difference between the projected pixels of the projector light source and the digital image results in significant information loss when the digital image is compressed and converted to projected pixels, significantly impacting the effectiveness and accuracy of the perturbed image. Summary of the Invention
[0003] The embodiments of the present application provide a method and apparatus for creating a disturbed image. Based on the optimization of similarity loss in face recognition results, the method and apparatus solve the problem of information loss in converting the digital domain to the physical domain of the projection light source by compressing digital pixels to projected pixels. This helps to improve the effectiveness and accuracy of the device in creating a disturbed image in the digital domain.
[0004] In a first aspect, an embodiment of the present application provides a method for creating a disturbed image, the method comprising:
[0005] A first perturbed image is generated based on a first image of a first object; a mask image is used to synthesize the first perturbed image with a second image of a second object to obtain a third image, where the second object is different from the first object; face recognition is performed on the third image and the first image to obtain a face recognition result, and similarity loss optimization is performed on the face recognition result to obtain a first gradient vector; digital pixel compression is performed on the first perturbed image to projected pixels, and information loss optimization is performed on the projected pixels to obtain a second gradient vector; and the first perturbed image is updated using the first gradient vector and the second gradient vector to obtain a second perturbed image.
[0006] In a second aspect, an embodiment of the present application provides a disturbance image creation device, the device comprising:
[0007] a generating unit, configured to generate a first disturbed image according to a first image of a first object;
[0008] a synthesis unit, configured to synthesize the first perturbed image with a second image of a second object using a mask image to obtain a third image, where the second object is different from the first object;
[0009] a first optimization unit, configured to perform face recognition on the third image and the first image to obtain a face recognition result, and perform similarity loss optimization using the face recognition result to obtain a first gradient vector;
[0010] a second optimization unit, configured to perform digital pixel compression on the first disturbed image to projected pixels, and perform information loss optimization on the projected pixels to obtain a second gradient vector;
[0011] An updating unit is configured to update the first disturbed image using the first gradient vector and the second gradient vector to obtain a second disturbed image.
[0012] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0013] A memory, a processor, and a disturbed image creation program stored in the memory and executable on the processor, wherein the disturbed image creation program is configured to implement part or all of the steps described in any method of the first aspect.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a disturbance image creation program is stored. When the disturbance image creation program is executed by a processor, some or all of the steps described in any method in the first aspect are implemented.
[0015] In a fifth aspect, embodiments of the present application provide a computer program product, wherein the computer program product includes a computer program operable to cause a computer to perform some or all of the steps described in any method of the first aspect of the embodiments of the present application. The computer program product may be a software installation package.
[0016] By implementing the embodiments of the present application, the electronic device first generates a first perturbed image based on the first image of the first object; uses a mask image to synthesize the first perturbed image onto the second image of the second object to obtain a third image, where the second object is different from the first object; then performs face recognition on the third image and the first image to obtain a face recognition result, and uses the face recognition result to perform similarity loss optimization to obtain a first gradient vector; then performs digital pixel compression on the first perturbed image to projected pixels, and performs information loss optimization on the projected pixels to obtain a second gradient vector; finally, the first perturbed image is updated using the first gradient vector and the second gradient vector to obtain a second perturbed image. On the basis of optimizing the similarity loss of the face recognition result, the information loss optimization of the digital pixel compression to the projected pixel is used to solve the problem of information loss in converting the digital domain of the projection light source to the physical domain, which is beneficial to improving the effectiveness and accuracy of the device in creating perturbed images in the digital domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.
[0018] Figure 1 Schematic diagram of the architecture of a disturbance image creation system provided in an embodiment of the present application;
[0019] Figure 2 This is a flowchart of a disturbed image creation method provided by an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of a scenario for obtaining a second image of a second object provided by an embodiment of the present application;
[0021] Figure 4 is a schematic diagram of a mask image provided in an embodiment of the present application;
[0022] Figure 5 This is a schematic diagram of an attack test report provided in an embodiment of the present application;
[0023] Figure 6 Schematic diagram of the structure of a disturbance image creation device provided in an embodiment of the present application;
[0024] Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work should fall within the scope of protection of the present invention.
[0026] The terms "first," "second," and "third," etc. in the specification, claims, and drawings of this application are used to distinguish between different objects, not to describe a particular order. In addition, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0027] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0028] Currently, in face recognition optical projection adversarial technology based on projector-camera systems, the perturbed image calculation process utilizes a transformation matrix that optimizes the color and position differences between the digital image and the projected image. During actual adversarial attack testing, the resolution difference between the projected pixels of the projector light source and the digital image results in significant information loss when the digital image is compressed and converted to projected pixels, significantly impacting the effectiveness and accuracy of the perturbed image.
[0029] In response to the above-mentioned problems, an embodiment of the present application provides a method and device for creating a disturbed image. The electronic device first generates a first disturbed image based on a first image of a first object; uses a mask image to synthesize the first disturbed image into a second image of a second object to obtain a third image, where the second object is different from the first object; then uses the third image and the first image to optimize the similarity loss of the face recognition result of the first disturbed image to obtain a first gradient vector; then optimizes the information loss of the digital pixel compression to the projection pixel of the first disturbed image to obtain a second gradient vector; finally, uses the first gradient vector and the second gradient vector to update the first disturbed image to obtain a second disturbed image. On the basis of optimizing the similarity loss of the face recognition result, the information loss of the digital pixel compression to the projection pixel is optimized to solve the problem of information loss in converting the projection light source from the digital domain to the physical domain, which is beneficial to improving the effectiveness and accuracy of the device in creating the disturbed image in the digital domain.
[0030] The disturbed image creation method and device provided in the embodiments of the present application can be applied to Figure 1 The perturbation image creation system shown in Figure 1 , Figure 1This is a schematic diagram of the architecture of a disturbance image creation system 100 provided in an embodiment of the present application. The disturbance image creation system 100 includes an electronic device 101, a projection component 102, and a photographing component 103. The electronic device 101 can communicate with the projection component 102 and the photographing component 103 through a network. The electronic device 101 can be a computer or mobile device for processing a large number of computing tasks and storing data, such as a smart phone, a tablet computer, etc. In this solution, the electronic device 101 is responsible for generating a first disturbance image based on a first image of a first object, synthesizing the first disturbance image with a second image of a second object using a mask image to obtain a third image, and then calculating a first gradient vector and a second gradient vector using multiple images. Finally, the first gradient vector and the second gradient vector are used to update the first disturbance image to obtain a second disturbance image. The projection component 102 and the photographing component 103 can be integrated into the electronic device 101, or integrated into other identical or different devices, and are not limited here.
[0031] Projection assembly 102 is a component or device capable of projecting images or videos onto a screen or other flat surface, such as a laser projector or micro-projector. Projection assembly 102 primarily consists of a light source, an optical system (including lenses, reflectors, etc.), and an image display element. In this embodiment, projection assembly 102 is responsible for projecting the resulting image onto a specific surface (e.g., a specific subject's face).
[0032] Camera component 103 is a component or device used to capture still images or record dynamic videos. Examples include digital cameras and mobile phone cameras. Camera component 103 primarily consists of a lens, an image sensor, an analog-to-digital converter, and a control circuit. In this embodiment, camera component 103 is responsible for capturing a first image of a first object and a second image of a second object. Both the first and second images contain facial images.
[0033] Based on this, the present application provides a method and device for creating a disturbed image, which is described in detail below with reference to the accompanying drawings.
[0034] See also Figure 2 , Figure 2 This is a flowchart of a disturbed image creation method provided by an embodiment of the present application, which is applied to Figure 1 The electronic equipment shown, such as Figure 2 As shown, the method includes the following steps:
[0035] S201: The electronic device generates a first disturbed image according to a first image of a first object.
[0036] Among them, the first image of the first object can be directly obtained from the database by the electronic device 101 of the disturbance image creation system 100. The first object is an object that has been authenticated by the face recognition model and serves as the attacked object in the face recognition model test scenario. The first image corresponding to the first object is a face image that can be recognized by the face recognition model. The first image can be used in the future to cooperate with the disturbance image to perform adversarial attack testing, thereby obtaining the test results of the face recognition model on the created disturbance image.
[0037] Specifically, the step of generating a first disturbed image according to a first image of a first object includes:
[0038] Step 1: preprocess the first image to obtain a preprocessed first image y, wherein the preprocessing operation on the first image includes but is not limited to a normalization operation, a filtering operation, and a histogram equalization operation.
[0039] Step 2: Obtain the grayscale image of the preprocessed first image Converting a color face image to a grayscale image can reduce the amount of data while retaining the basic contour and texture information of the image. Specifically, the grayscale conversion method can be a weighted average method, and the calculation formula is as follows:
[0040] Gray=0.299R+0.587G+0.114B;
[0041] Among them, R, G, and B are the values of the red, green, and blue channels of the color image respectively, and Gray is the grayscale value.
[0042] Step 3: Generate a random moiré image Ψ based on the moiré pattern characteristics of the projected image, wherein the steps of generating the moiré image may be: obtain projection data of the image in different directions, where the projection data may be a set of data or a vector; perform statistical analysis on the projection data in each direction, for example, calculate statistical characteristics such as the mean value, standard deviation, and frequency distribution in different directions; generate a random number sequence or probability distribution function in each direction based on the obtained statistical characteristics; sample in each direction based on the random number sequence or probability distribution function and combine the sampling results to obtain a moiré image; alternatively, the moiré image may be obtained by adding noise to the current attacked face image, for example, decomposing the selected image and converting it into grayscale or other color space, and generating the moiré image by adding random noise.
[0043] Step 4: Grayscale image Convert to RGB image and fuse the RGB image with the Moore image Ψ to get the image Among them, the RGB image is fused with the Moore image Ψ to obtain the image The formula is as follows:
[0044]
[0045] and The subscript t represents the current iteration number, and t = 0 in the initial state.
[0046] Among them, θ is the transparency coefficient. Specifically, the transparency coefficient is a parameter used to adjust the image superposition effect. The transparency coefficient determines the relationship between the moiré image and the image. The degree of transparency during overlay. Specifically, the transparency coefficient usually ranges from 0 to 1. When the transparency coefficient is 0, it indicates complete transparency. The image with a transparency coefficient of 0 will not have any effect on the underlying image during the overlay process. When the transparency coefficient is 1, it indicates complete opacity. The image with a transparency coefficient of 1 will be fully displayed in its original state in the final composite image, and will cover other underlying elements.
[0047] Step 5: Based on the image and a random brightness factor γ to simulate the illumination change, we get the image is the first perturbation image. Specifically, according to the image and random brightness factor γ to obtain the first perturbation image The formula is as follows:
[0048]
[0049] in, and The subscript t indicates the current number of iterations. In the initial state, t = 0. is the first perturbation image, and γ is the brightness factor.
[0050] The brightness factor γ is a parameter used to adjust the brightness of the image. By increasing or decreasing the brightness factor, the image can be visually presented with different levels of brightness and darkness. The brightness factor usually ranges from -1 to 1. When the value is 0, it means no change is made. When it is a negative value, it means reducing the brightness. When it is a positive value, it means increasing the brightness.
[0051] In a possible implementation, before step S202 , the method further includes: acquiring the second image of the second object at a preset position, where a distance between the preset position and the position of the projection light source is less than a preset distance.
[0052] Among them, the second object is the attack object photographed in the face recognition model testing scene, and the created perturbation image is used to project on the face of the second object, so as to perform an attack test on the face recognition model. The second image of the second object can be obtained by photographing the photographing component 103 of the perturbation image creation system 100. Preferably, the photographing component 103 is a smart phone with an image acquisition device. The smart phone photographs the face of the second object through its own image acquisition device to obtain the second image.
[0053] The second object is located at a preset position, and the distance between the preset position and the position of the projection light source does not exceed the preset distance. Figure 1 The optical lens of the projection assembly 102, for example, when the preset distance is 5 meters, the distance between the preset position of the second object and the position of the optical lens of the projection assembly 102 does not exceed 5 meters, and the preset distance can be set and changed by the user on the electronic device.
[0054] Among them, when using the image acquisition device to capture the face of the second object, it can be slowly moved at a position close to the projection light source to obtain multiple second images of the second object captured from multiple angles. The specific movement method can be translational movement, rotational movement, spiral movement, etc., and the movement method of the test equipment is not restricted here.
[0055] Specifically, see Figure 3 , Figure 3 is a schematic diagram of a scenario for obtaining a second image of a second object provided by an embodiment of the present application, such as Figure 3 As shown, the preset position of the second object is close to the position of the projection light source, and the image acquisition device is close to the position of the projection light source. After the image acquisition device aims its own image acquisition device at the face of the second object, it ensures that the face of the second object is completely within the viewfinder of the image acquisition device. At this time, the image acquisition device is slowly rotated, and the image acquisition device is used to capture the face of the second object during the rotation process, thereby obtaining multiple second images captured from multiple angles.
[0056] It can be seen that in this example, the second image that the electronic device can obtain is obtained by photographing the second object at the preset position, and the distance between the preset position and the position of the projection light source is less than the preset distance. This position constraint characteristic with a very small distance difference can effectively reduce the test error caused by the lack of position calibration due to the camera and projector being set far apart in the existing test system. Therefore, the accuracy that can only be achieved by the strong position calibration of the existing test system can be achieved without position calibration, which is conducive to reducing the algorithm complexity of the test system and improving the scenario applicability of the test system.
[0057] S202: The electronic device uses a mask image to synthesize the first disturbance image with a second image of a second object to obtain a third image, where the second object is different from the first object.
[0058] The mask image includes the to-be-occluded area with a mask of 1 and the non-occluded area with a mask of 0. Users can determine the values of the mask image at different positions based on the key facial features required for face recognition.
[0059] The formula for synthesizing the first perturbed image into the second image of the second object using the mask image to obtain the third image is as follows:
[0060]
[0061] Among them, x i ′ is the third image, is the first perturbation image, t represents the current iteration number, in the initial state t=0, x i is the second image of the second object, M is the mask image, and k is the number of the second images of the second object taken.
[0062] Specifically, see Figure 4 , Figure 4 is a schematic diagram of a mask image provided in an embodiment of the present application, such as Figure 4 As shown, the mask image has the same shape and size as the second image of the second object. The mask of the black area in the mask image is 1, and the mask of the white area is 0. The areas with the mask of 1 in the mask image are the eye area, the mouth area, and the nose area.
[0063] S203: The electronic device performs face recognition on the third image and the first image to obtain a face recognition result, and performs similarity loss optimization using the face recognition result to obtain a first gradient vector.
[0064] In one possible implementation, face recognition is performed on the third image and the first image to obtain a face recognition result, and similarity loss optimization is performed using the face recognition result to obtain a first gradient vector, including: using the third image and the model to obtain a face recognition prediction result; using the first image and the model to obtain a true face recognition result; and performing similarity loss optimization on the first perturbation image between the face recognition prediction result and the true face recognition result to obtain the first gradient vector.
[0065] Among them, the first image is the image that has passed the authentication of the face recognition model, and the third image is the image after adding disturbance. When inputting the third image and the first image into the model, it is necessary to ensure that the format, resolution, etc. of the third image and the first image meet the model input requirements.
[0066] Among them, the model here can be a face recognition model, and the face recognition model can be a model based on convolutional neural network, a traditional machine learning model or a hybrid model. Taking the convolutional neural network model as an example, the face recognition model includes convolutional layers (Convolutional Layers), pooling layers (Pooling Layers) and fully-connected layers (Fully-Connected Layers).
[0067] Specifically, the convolutional layer is the core component of CNN, which extracts local features of the image by sliding convolution kernels (also called filters) across the image. For face recognition, the convolutional layer can extract features such as edges and textures of the face. Different convolution kernels can extract different features, such as horizontal edges and vertical edges.
[0068] The main function of the pooling layer is to downsample the convolutional feature map to reduce the amount of data. Common pooling methods include max pooling and average pooling. Max pooling selects the maximum value within a small area as the output, while average pooling calculates the average value within the area. In face recognition, the pooling layer helps to highlight key facial features, such as the strongest features around the eyes and mouth.
[0069] After the convolutional and pooling layers, the fully connected layer flattens the feature maps from these layers into a one-dimensional vector. This is then fed into the fully connected layer. Each neuron in this layer is connected to all neurons in the previous layer, allowing the fully connected layer to integrate the various previously extracted features for classification. In a face recognition model, the fully connected layer determines whether the face belongs to a recognized person based on the previously extracted facial feature vector. The output can be a probability value (e.g., the probability that the input image is a recognized face) or a category label (e.g., 0 for an unrecognized face, 1 for a recognized face).
[0070] The face recognition result is optimized by similarity loss using the following gradient function to obtain a first gradient vector:
[0071]
[0072] Among them, g2(i) is the first gradient vector, is the gradient operator, which represents the variable Find the gradient, that is, calculate the function about The partial derivative of Function is the second loss function, The two variables in the function brackets are the input of the function, namely the third image x′ i , the first image y, where is the first perturbation image.
[0073] Specifically, the loss function of the above formula is split as follows:
[0074]
[0075] in,
[0076] in, is the cosine similarity loss function.
[0077] Furthermore, the loss function L cos Relative to the perturbation image The expression of the gradient function G(j) is inferred as follows:
[0078] Known:
[0079]
[0080] Among them, for each item in G(j), the chain rule is applied to split the calculation and the complete expression of G(j) is obtained as follows:
[0081]
[0082] It can be seen that in this example, the electronic device can optimize the first perturbed image based on the similarity loss between the predicted result and the true result to obtain the first gradient vector. This operation can adjust the relevant parameters based on the actual recognition difference feedback, so that the subsequent adjustment of the first perturbed image through the first gradient vector can continuously improve in the direction of narrowing the difference between the predicted and true results, thereby improving the accuracy of the generated perturbed image.
[0083] S204: The electronic device performs digital pixel compression on the first disturbed image to convert it into projection pixels, and performs information loss optimization on the projection pixels to obtain a second gradient vector.
[0084] In a possible implementation, performing digital pixel compression on the first perturbed image to projected pixels, and performing information loss optimization on the projected pixels to obtain a second gradient vector includes:
[0085] The first disturbed image is divided into blocks to obtain first disturbed image blocks corresponding to the projection pixels; and color depth difference constraint processing is performed on pixels within the first disturbed image block to obtain a second gradient vector.
[0086] Among them, image segmentation can divide an image into multiple smaller image blocks, so as to perform refined processing on each smaller image block. The image segmentation method can be uniform segmentation according to a fixed size, for example, the image is segmented into square areas with a fixed pixel value on each side, or non-uniform segmentation can be performed according to the specific feature distribution and pixel coordinate relationship in the image.
[0087] Among them, color depth is used to indicate the color richness that pixels in an image can represent. Color depth difference constraint processing can be achieved by setting a color depth threshold range. When the color depth within an image block exceeds the threshold range, the color depth within the image block is adjusted to return to the threshold range through a specific algorithm.
[0088] It can be seen that in this example, the electronic device can obtain the second gradient vector by dividing the first disturbed image into blocks and processing the color depth differences of the pixels within the blocks, which helps to subsequently refine the disturbed image according to the second gradient vector, so that each part of the disturbed image is accurately matched with the projected pixels, reducing the feature deviation caused by color anomalies, and improving the accuracy of the disturbed image.
[0089] In a possible implementation, performing color depth difference constraint processing on pixels within a block of the first disturbed image block to obtain a second disturbed image gradient value includes:
[0090] Color depth difference processing is performed on adjacent pixels in a block of the first disturbed image block to obtain a first calculation result; and gradient calculation is performed on the first disturbed image using the first calculation result to obtain a second gradient vector.
[0091] Specifically, dividing the first disturbed image into blocks to obtain first disturbed image blocks corresponding to the projection pixels includes:
[0092] The method further comprises determining a size of a minimum digital pixel block capable of covering at least a single projection pixel based on a resolution relationship between the projection pixel and the digital pixels of the first disturbed image; and performing an average block processing on the first disturbed image based on the size of the minimum digital pixel block to obtain a first disturbed image block corresponding to the projection pixel.
[0093] The resolution determines the image's fineness and pixel distribution density, among other characteristics. The projection pixel represents the relevant characteristics of the pixel in the projection image. The first perturbed image has its own digital pixel situation. The digital pixels of the first perturbed image can at least cover the minimum range of a single projection pixel, which is the minimum digital pixel block.
[0094] The process of determining the minimum digital pixel block includes: firstly determining the pixel density ratios of the projection pixels and the first perturbation image digital pixels in the horizontal and vertical directions respectively; then converting the pixel density ratio relationship into pixels to determine the size of the minimum digital pixel block; finally, the size of the minimum digital pixel block is determined to be a rectangular area.
[0095] After the size of the minimum digital pixel block is determined, the first disturbed image may be evenly divided according to the size of the minimum digital pixel block to obtain a plurality of first disturbed image blocks corresponding to the projection pixels.
[0096] Specifically, performing gradient calculation on the first disturbed image using the first calculation result to obtain a second gradient vector includes:
[0097] The gradient value of the first loss function Lres relative to the color depth of each pixel of the first perturbed image block is calculated using the following first gradient function g1(k) to obtain a second gradient vector:
[0098]
[0099] in,
[0100] Where P is the number of the first perturbed image blocks, H is the height of the first perturbed image blocks, W is the width of the first perturbed image blocks, sign(y b,j ) is the sign function, y b,j is the first calculation result, is the color depth of the jth pixel of the bth first perturbation image block, is the color depth of the j-1th pixel adjacent to the jth pixel in the bth first perturbation image block, is the color depth of the kth pixel of the first perturbation image.
[0101] Among them, the first loss function L res The calculation formula is as follows:
[0102]
[0103] It can be seen that in this example, the electronic device can obtain the second gradient vector by performing color depth difference processing on adjacent pixels in the first disturbed image block and performing gradient calculation on the first disturbed image, which is conducive to further optimizing the color change of pixels in the disturbed image according to the second gradient vector and improving the accuracy of the disturbed image.
[0104] In one possible implementation, the method further includes:
[0105] Normalization is performed on the second gradient vector to obtain a processed second gradient vector.
[0106] Specifically, an L1 norm calculation is performed on the total gradient function to obtain an L1 norm calculation result; and the second gradient vector is normalized using the L1 norm calculation result to obtain a processed second gradient vector.
[0107] Among them, the total gradient function is the gradient function of the total loss function relative to the first perturbation image, and the total loss function is the first loss function Lres, the second loss function L cos , the third loss function L smooth The harmony.
[0108] Among them, the loss function L smooth is the sum of the absolute values of all adjacent pixel differences in the perturbed image during the current iteration, and the loss function L smooth The expression is as follows:
[0109]
[0110] Among them, M and N are the total number of rows and columns of the perturbed image, respectively, and x i,j is the pixel at row i and column j.
[0111] It can be seen that in this example, the electronic device can obtain a processed second gradient vector by normalizing the second gradient vector, and can map the numerical values of each element of the gradient vector to a unified standard range, eliminating differences caused by factors such as different magnitudes and scales, thereby facilitating subsequent optimization of the disturbed image based on the first gradient vector and the second gradient vector.
[0112] S205: The electronic device updates the first disturbed image by using the first gradient vector and the second gradient vector to obtain a second disturbed image.
[0113] The formula for updating the first perturbed image to obtain the second perturbed image is as follows:
[0114]
[0115] Among them, α is the learning rate, is the first perturbation image, is the second perturbation image, g is obtained according to the first gradient vector and the second gradient vector, g can be the sum of the first gradient vector and the second gradient vector, g1 is the L1 norm of g, and the calculation formula of α is as follows:
[0116] α=∈ / t;
[0117] Among them, ∈ is a hyperparameter that controls the intensity of iterative perturbation, and t is the current number of iterations.
[0118] Among them, in the process of multiple iterations, the parameters can be continuously updated by the gradient descent method, so that the perturbation image obtained in the current iteration is optimized in the direction of minimizing the total loss function. For example, the steps of updating the parameters using the projected gradient descent method are: initialize the parameters θ0; for each iteration t, calculate the objective function at the current parameters θ tThe gradient at , then perform gradient descent update, and finally project the updated parameters back to the constraint set.
[0119] It can be seen that in this example, the electronic device can solve the problem of information loss in the conversion of digital domain to physical domain of projection light source by optimizing the similarity loss of face recognition results while compressing digital pixels to projection pixels, which is beneficial to improving the effectiveness and accuracy of the device in creating disturbed images in the digital domain.
[0120] In a possible implementation, after updating the first perturbed image by using the first gradient vector and the second gradient vector to obtain a second perturbed image, the method further includes: performing a color clipping operation on the second perturbed image by using a color interval to obtain an updated second perturbed image.
[0121] Among them, reasonable color ranges can be statistically calculated by analyzing a large number of face images. The color range includes the red component value range, the green component value range, and the blue component value range in the RGB color space. For example, the red component value range, the green component value range, and the blue component value range in the RGB color space are [180, 255], [0, 50], and [0, 80], respectively.
[0122] Among them, color clipping is an image processing operation that filters and modifies the pixel colors in the image by setting a specific color interval in a color space (such as common RGB, CMYK, HSV and other color spaces). For example, in the RGB color space, the color of each pixel is represented by the values of the red, green, and blue channels, and the value range is usually 0-255. At this time, the red channel value is set to be greater than 180 and less than 255, the green channel value is less than 50, and the blue channel value is less than 80 as the color interval to be clipped. Then, when traversing the image pixels, the pixels that do not meet this interval will be processed. Among them, the processing method can be to directly replace the colors of these pixels with other colors, or to reassign them according to the colors of the surrounding pixels.
[0123] Among them, the Clip function can be used for color clipping. Clip is a clipping function used to limit a given value to a specified range. The clipping function usually consists of two parameters: the value to be restricted and the specified range. For example, for Clip(x,min_value,max_value), x represents the value to be restricted, min_value represents the minimum value allowed, and max_value represents the maximum value allowed. If x is less than min_value, min_value is returned; if x is greater than max_value, max_value is returned; otherwise, x itself is returned.
[0124] The projector will change the pixel color value after projecting the input image. For each color channel, the mapping relationship from the input image to the projected image satisfies the linear relationship only in the middle area, and the two ends are nonlinear. For each color channel, the user can estimate a color interval [L, H] in the middle area (that is, satisfying the linear relationship) and change the perturbed image Crop to the [L, H] interval.
[0125] It can be seen that in this example, the electronic device can constrain the solution process of the disturbed image to the linear part by cropping the disturbed image to the interval where the mapping relationship between the digital image and the projection image of the projection light source is a linear relationship, thereby reducing the difference between the digital image solution and the projected image.
[0126] In a possible implementation, before performing a color cropping operation on the second disturbed image using a preset color interval to obtain an updated second disturbed image, the method further includes: obtaining a random disturbance value; and randomly perturbing the color interval using the random disturbance value to obtain the disturbed color interval.
[0127] Specifically, a random perturbation value v is randomly sampled in a given interval [-τ, τ], and the value of [L, H] is adjusted to [Lv, H+v].
[0128] For example, the color interval is [11, 45], the given interval is [1, 3], and the random perturbation value is 2, then 11-2=9, 45+2=47, and the adjusted color linear conversion interval is determined to be [9, 47].
[0129] It can be seen that in this example, the electronic device can introduce a random perturbation value to correct the color interval because the color interval is an estimated value obtained based on a large amount of data statistics and is not necessarily accurate. This ensures that the perturbation image obtained by subsequent color cropping based on the color interval falls within the linear mapping interval of the digital image to the projection light source.
[0130] In one possible implementation, after updating the first perturbed image using the first gradient vector and the second gradient vector to obtain a second perturbed image, the method further includes: acquiring a fourth image of the second object under the projection light source corresponding to the second perturbed image; invoking the model to process the fourth image to obtain a recognition result; and determining a test result of the model based on the second object and the recognition result.
[0131] After obtaining at least one test result, a test report can be generated based on the test result and displayed on the display interface of the electronic device, for example, see Figure 5 , Figure 5 This is a schematic diagram of an attack test report provided by an embodiment of the present application. Figure 5 As shown, the current number of tested samples is 2, the number of samples to be tested is 3, the number of tested samples that have passed the face recognition model verification is 0, and the pass rate of the face recognition model is 0.
[0132] When the pass rate of the face recognition model exceeds the preset pass rate threshold, a prompt message is displayed on the display interface of the electronic device, and the display area of the model pass rate in the test report is highlighted.
[0133] The display area of the test report is provided with a stop function icon, and the user can click the stop function icon to stop the creation of the disturbance image or the test of the face recognition model.
[0134] It can be seen that in this example, the electronic device can obtain the fourth image of the second object under a specific projection light source, and then call the model processing to obtain the recognition result and determine the test result of the model accordingly, so as to effectively test the use effect of the created disturbed image in the corresponding lighting scene, and facilitate targeted improvement of the disturbed image according to the test results.
[0135] See also Figure 6 , Figure 6 is a structural diagram of a disturbance image creation device provided in an embodiment of the present application, such as Figure 6 As shown, the disturbance image creation device 600 includes:
[0136] A generating unit 601 is configured to generate a first disturbed image according to a first image of a first object;
[0137] a synthesis unit 602 configured to synthesize the first disturbed image with a second image of a second object using a mask image to obtain a third image, where the second object is different from the first object;
[0138] A first optimization unit 603 is configured to perform face recognition on the third image and the first image to obtain a face recognition result, and perform similarity loss optimization using the face recognition result to obtain a first gradient vector;
[0139] A second optimization unit 604 is configured to perform digital pixel compression on the first disturbed image to projected pixels, and perform information loss optimization on the projected pixels to obtain a second gradient vector;
[0140] The updating unit 605 is configured to update the first perturbed image using the first gradient vector and the second gradient vector to obtain a second perturbed image.
[0141] In one possible implementation, in terms of performing digital pixel compression on the first perturbed image to projected pixels and performing information loss optimization on the projected pixels to obtain the second gradient vector, the second optimization unit 604 is specifically configured to: divide the first perturbed image into blocks to obtain first perturbed image blocks corresponding to the projected pixels; and perform color depth difference constraint processing on pixels within the first perturbed image blocks to obtain the second gradient vector.
[0142] In a possible implementation, in terms of performing color depth difference constraint processing on pixels within a block of the first perturbed image block to obtain a second perturbed image gradient value, the second optimization unit 604 is specifically configured to: perform color depth difference processing on adjacent pixels within the block of the first perturbed image block to obtain a first calculation result; and perform gradient calculation on the first perturbed image using the first calculation result to obtain a second gradient vector.
[0143] In a possible implementation, in terms of performing gradient calculation on the first perturbed image using the first calculation result to obtain the second gradient vector, the second optimization unit 604 is specifically configured to: use the following first gradient function g1(k) to calculate the gradient value of the first loss function Lres relative to the color depth of each pixel of the first perturbed image block to obtain the second gradient vector:
[0144]
[0145] in,
[0146] Where P is the number of the first perturbed image blocks, H is the height of the first perturbed image blocks, W is the width of the first perturbed image blocks, sign(y b,j ) is the sign function, y b,j is the first calculation result, is the color depth of the jth pixel of the bth first perturbation image block, is the color depth of the j-1th pixel adjacent to the jth pixel in the bth first perturbation image block, x adv (k) is the color depth of the kth pixel of the first perturbed image.
[0147] Among them, the first loss function L res The calculation formula is as follows:
[0148]
[0149] In one possible implementation, in terms of dividing the first disturbed image into blocks to obtain first disturbed image blocks corresponding to the projection pixels, the second optimization unit 604 is specifically configured to: determine, based on a resolution relationship between the projection pixels and digital pixels of the first disturbed image, a size of a minimum digital pixel block capable of covering at least a single projection pixel; and perform average blocking processing on the first disturbed image based on the size of the minimum digital pixel block to obtain first disturbed image blocks corresponding to the projection pixels.
[0150] In a possible implementation, the second optimization unit 604 is further configured to: perform normalization processing on the second gradient vector to obtain a processed second gradient vector.
[0151] In a possible implementation, after updating the first perturbed image by using the first gradient vector and the second gradient vector to obtain a second perturbed image, the updating unit 605 is further configured to: perform a color clipping operation on the second perturbed image by using a color interval to obtain an updated second perturbed image.
[0152] In a possible implementation, before performing a color cropping operation on the second disturbed image using a preset color interval to obtain an updated second disturbed image, the updating unit 605 is further configured to: obtain a random disturbance value; and randomly perturb the color interval using the random disturbance value to obtain the disturbed color interval.
[0153] In one possible implementation, after updating the first perturbed image using the first gradient vector and the second gradient vector to obtain a second perturbed image, the updating unit 605 is further configured to: obtain a fourth image of the second object under the projection light source corresponding to the second perturbed image; call the model to process the fourth image to obtain a recognition result; and determine a test result of the model based on the second object and the recognition result.
[0154] It is worth noting that the specific functional implementation of the disturbance image creation device 600 can be found in the above Figure 2In the description of the disturbed image creation method shown in FIG. 1 , for example, the generation unit 601 is used to implement the relevant content of execution S201, the synthesis unit 602 is used to implement the relevant content of execution S202, the first optimization unit 603 is used to implement the relevant content of S203, the second optimization unit 604 is used to implement the relevant content of S204, and the update unit 605 is used to implement the relevant content of S205. Each unit or module in the disturbed image creation device 600 can be individually or completely merged into one or more other units or modules to form a structure, or one (or more) of the units or modules can be further divided into multiple functionally smaller units or modules to form a structure, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above-mentioned units or modules are divided according to logical functions. In actual applications, the functions of one unit (or module) are implemented by multiple units (or modules), or the functions of multiple units (or modules) are implemented by one unit (or module).
[0155] According to the description of the above method embodiment and related device embodiment, please refer to Figure 7 , Figure 7 This is a structural diagram of an electronic device provided in an embodiment of the present application. Figure 7 The electronic device 700 shown includes a processor 701 , a memory 702 , a communication interface 703 , and a bus 704 . The processor 701 , the memory 702 , and the communication interface 703 are communicatively connected to each other via the bus 704 .
[0156] Optionally, the memory 702 is a ROM, a static storage device, a dynamic storage device or a RAM.
[0157] The memory 702 can store executable program codes. When the executable program codes stored in the memory 702 are executed by the processor 701, the processor 701 and the communication interface 703 are used to execute the program codes. Figure 2 The various steps of the disturbed image creation method of the illustrated embodiment.
[0158] The processor 701 adopts a general CPU, a microprocessor, an application-specific integrated circuit ASIC, a GPU or one or more integrated circuits to execute relevant programs to perform the disturbed image creation method of the method embodiment of the present application.
[0159] Processor 701 can also be an integrated circuit chip with signal processing capabilities. During implementation, each step of the disturbed image creation method of the present application can be completed by hardware integrated logic circuits or software instructions in processor 701. Optionally, processor 701 is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The processor can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The optional software modules are located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media well-known in the art. The storage medium is located in the memory 702. The processor 701 reads the information in the memory 702 and, in combination with its hardware, completes the functions required to be performed by the modules included in the disturbance image creation device 600 in an embodiment of the present application, or performs the disturbance image creation method in the method embodiment of the present application.
[0160] The communication interface 703 uses, for example but not limited to, a transceiver and other transceiver-related devices.
[0161] The bus 704 may include a path for transmitting information between various components of the electronic device 700 (eg, the memory 702 , the processor 701 , and the communication interface 703 ).
[0162] It should be noted that although Figure 7 The electronic device 700 shown only shows a memory, a processor, and a communication interface. However, in the specific implementation process, those skilled in the art should understand that the electronic device 700 also includes other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the electronic device 700 may also include hardware devices that implement other additional functions. In addition, those skilled in the art should understand that the electronic device 700 may also include only the devices necessary to implement the embodiments of the present application, and does not necessarily include Figure 7 All devices shown in .
[0163] An embodiment of the present application provides a computer-readable storage medium storing a computer program for electronic data exchange. The computer program includes execution instructions for executing some or all of the steps of any one of the disturbed image creation methods described in the above-mentioned disturbed image creation method embodiments. The computer includes an electronic terminal device.
[0164] An embodiment of the present application provides a computer program product, wherein the computer program product includes a computer program, and the computer program is operable to cause a computer to perform part or all of the steps of any of the disturbance image creation methods described in the above method embodiments. The computer program product can be a software installation package.
[0165] It should be noted that for any of the aforementioned embodiments of the disturbed image creation method, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by this application.
[0166] The embodiments of the present application are introduced in detail above. Specific examples are used herein to illustrate the principles and implementation methods of a disturbed image creation method and device of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those skilled in the art, based on the idea of a disturbed image creation method and device of the present application, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present application.
[0167] The present application is described with reference to the flowcharts and / or block diagrams of the methods, hardware products, and computer program products of the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable perturbation image creation device to produce a machine, so that the instructions executed by the processor of the computer or other programmable perturbation image creation device generate instructions for implementing the process Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0168] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable perturbation image creation device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, the instruction device being implemented in the process Figure 1 a process or multiple processes and / or boxes Figure 1The memory may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0169] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps. The fact that certain measures are recited in different dependent claims does not mean that these measures cannot be combined to produce good results.
[0170] Those skilled in the art will appreciate that all or part of the steps in the various methods of any of the above-mentioned methods for creating a disturbed image can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0171] It is understandable that any product that is controlled or configured to execute the processing method of the flowchart described in an embodiment of a disturbed image creation method of the present application, such as the device and computer program product in the above flowchart, falls within the scope of the related products described in the present application.
[0172] Obviously, those skilled in the art may make various modifications and variations to the disturbed image creation method and apparatus provided herein without departing from the spirit and scope of this application. Thus, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application is intended to encompass such modifications and variations.
Claims
1. A disturbed image creation method, characterized in that: include: generating a first disturbed image based on the first image of the first object; synthesizing the first perturbed image with a second image of a second object using a mask image to obtain a third image, where the second object is different from the first object; performing face recognition on the third image and the first image to obtain a face recognition result, and performing similarity loss optimization using the face recognition result to obtain a first gradient vector; Performing digital pixel compression on the first disturbed image to projected pixels, and performing information loss optimization on the projected pixels to obtain a second gradient vector; The first perturbed image is updated using the first gradient vector and the second gradient vector to obtain a second perturbed image.
2. The method according to claim 1, wherein The step of digitally compressing the first disturbed image into projected pixels and performing information loss optimization on the projected pixels to obtain a second gradient vector includes: Dividing the first disturbed image into blocks to obtain first disturbed image blocks corresponding to the projection pixels; Color depth difference constraint processing is performed on pixels within the first disturbed image block to obtain a second gradient vector.
3. The method according to claim 2, wherein The performing color depth difference constraint processing on pixels within the first disturbed image block to obtain a second disturbed image gradient value includes: Performing color depth difference processing on adjacent pixels in a block of the first disturbed image block to obtain a first calculation result; Performing gradient calculation on the first disturbed image using the first calculation result to obtain a second gradient vector.
4. The method according to claim 3, wherein The step of performing gradient calculation on the first disturbed image using the first calculation result to obtain a second gradient vector includes: The gradient value of the first loss function Lres relative to the color depth of each pixel of the first perturbed image block is calculated using the following first gradient function g1(k) to obtain a second gradient vector: in, Wherein, P is the number of the first perturbed image blocks, H is the height of the first perturbed image blocks, W is the width of the first perturbed image blocks, sign(y b,j ) is the symbol function, u b,j is the first calculation result, is the color depth of the jth pixel of the bth first perturbed image block, is the color depth of the j-1th pixel adjacent to the jth pixel in the bth first perturbed image block, x adv (k) is the color depth of the k-th pixel of the first perturbation image, and the calculation formula of the first loss function Lres is:
5. The method according to claim 2, wherein The step of dividing the first disturbed image into blocks to obtain first disturbed image blocks corresponding to the projection pixels includes: determining a size of a minimum digital pixel block capable of at least covering a single projection pixel according to a resolution relationship between the projection pixel and the digital pixels of the first disturbed image; The first disturbed image is averagely divided into blocks according to the size of the minimum digital pixel block to obtain a first disturbed image block corresponding to the projection pixel.
6. The method according to any one of claims 2 to 5, characterized in that The method further comprises: Normalization is performed on the second gradient vector to obtain a processed second gradient vector.
7. The method according to claim 1, wherein After updating the first disturbed image by using the first gradient vector and the second gradient vector to obtain a second disturbed image, the method further includes: A color clipping operation is performed on the second disturbed image using the color interval to obtain an updated second disturbed image.
8. The method according to claim 7, wherein Before performing a color clipping operation on the second disturbed image using a preset color range to obtain an updated second disturbed image, the method further includes: Get random perturbation value; The color interval is randomly disturbed by using the random disturbance value to obtain the disturbed color interval.
9. The method according to claim 1, wherein Before synthesizing the first disturbed image into the second image of the second object using the mask image to obtain the third image, the method further includes: The second image of the second object at a preset position is acquired, where a distance between the preset position and a position of a projection light source is less than a preset distance.
10. The method according to claim 1, wherein After updating the first disturbed image by using the first gradient vector and the second gradient vector to obtain a second disturbed image, the method further includes: Acquire a fourth image of the second object under the projection light source corresponding to the second disturbed image; Calling the model to process the fourth image to obtain a recognition result; A test result of the model is determined according to the second object and the recognition result.
11. A disturbance image creation device, characterized in that: include: a generating unit, configured to generate a first disturbed image according to a first image of a first object; a synthesis unit, configured to synthesize the first perturbed image with a second image of a second object using a mask image to obtain a third image, where the second object is different from the first object; a first optimization unit, configured to perform face recognition on the third image and the first image to obtain a face recognition result, and perform similarity loss optimization using the face recognition result to obtain a first gradient vector; a second optimization unit, configured to perform digital pixel compression on the first disturbed image to projected pixels, and perform information loss optimization on the projected pixels to obtain a second gradient vector; An updating unit is configured to update the first disturbed image using the first gradient vector and the second gradient vector to obtain a second disturbed image.
12. An electronic device, characterized in that: The electronic device comprises: A memory, a processor, and an executable program code stored in the memory and capable of running on the processor, wherein the processor executes the steps of any one of the methods according to claims 1 to 10 when executing the executable program code.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores executable program code, which includes execution instructions for executing the steps of any one of the methods according to claims 1-10.
14. A computer program product, characterized in that The computer program product includes a computer program, and the computer program is used to enable a computer to execute the steps in any one of the methods 1-10.