Image processing method and device, equipment and medium
By transferring knowledge to multiple image task models, and updating image adversarial samples using the cumulative value of feature similarity, the problem of limited application scope of white box adversarial attack methods is solved, and the wide applicability and evaluation effect of the image recognition model are improved.
Patent Information
- Application Number
- CN202410016130.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
The scope of application of the adversarial samples generated by the existing white box-based adversarial attack methods is limited, and the performance of different recognition models cannot be effectively evaluated, resulting in frequent misidentification of face-scanning products.
By transferring knowledge to M image task models, an image migration model is generated, image adversarial samples are updated using the accumulated feature similarity values, and a test sample pair suitable for multiple recognition models is generated to evaluate the performance of the image recognition model.
It improves the applicability of image adversarial samples and the evaluation effect of recognition models, enhances the attack migration of adversarial samples, and improves the security evaluation effect of image recognition models.
Smart Images

Figure CN120260093A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an image processing method, apparatus, device, and medium. Background Art
[0002] In scenarios where face recognition technology is widely applied, deep neural networks are generally used as the core algorithms of face recognition products. However, in the face of noise samples generated by some special means, incorrect prediction results will be output, which may lead to misidentification in face recognition products, bringing huge negative impacts to face recognition-related services. In order to defend against vulnerabilities in deep neural networks, the performance of the deep neural networks in face recognition products can be evaluated.
[0003] Currently, the performance evaluation scheme of deep neural networks usually adopts a white-box based adversarial attack method. In the white-box based adversarial attack method, the model parameters and the backpropagated gradient information of a single recognition model (for example, recognition model A) can be obtained, and the pixel values of the original sample are updated step by step or iteratively according to the backpropagated gradient information, so that the sample becomes aggressive. However, the adversarial samples generated by the white-box based adversarial attack method are only applicable to the aforementioned recognition model A. If the generated adversarial samples are applied to other recognition models, the model evaluation effect will be very poor. Therefore, the applicable range of the adversarial samples generated by the white-box based adversarial attack method is very limited. Summary of the Invention
[0004] Embodiments of this application provide an image processing method, apparatus, device, and medium, which can improve the applicability of image adversarial samples and enhance the evaluation effect of recognition models.
[0005] One aspect of the embodiments of this application provides an image processing method, including:
[0006] Obtain M image transfer models; the M image transfer models are recognition models obtained by knowledge transfer of M image task models, and the M image task models are pre-trained models applied to different image tasks. One image task model corresponds to one image transfer model, and M is a positive integer;
[0007] Obtain a first image and a second image, and input the first image and the second image into the i-th image transfer model among the M image transfer models, and obtain a first recognition feature corresponding to the first image and a second recognition feature corresponding to the second image through the i-th image transfer model; different objects are included in the first image and the second image, and i is a positive integer less than or equal to M;
[0008] According to the first recognition feature and the second recognition feature, obtain the feature similarity associated with the i-th image migration model, and add up the feature similarities associated with the M image migration models to obtain the cumulative feature similarity value;
[0009] Update the first image according to the cumulative feature similarity value to generate an image adversarial sample, and combine the image adversarial sample and the second image into a test sample pair; the test sample pair is used to evaluate the performance of the image recognition model in the business application.
[0010] One aspect of the embodiments of the present application provides an image processing device, including:
[0011] A migration model acquisition module, configured to acquire M image migration models; the M image migration models are recognition models obtained by performing knowledge migration on M image task models, the M image task models are pre-trained models applied to different image tasks, one image task model corresponds to one image migration model, and M is a positive integer;
[0012] An image recognition module, configured to acquire a first image and a second image, input the first image and the second image into the i-th image migration model among the M image migration models, and obtain a first recognition feature corresponding to the first image and a second recognition feature corresponding to the second image through the i-th image migration model; different objects are included in the first image and the second image, and i is a positive integer less than or equal to M;
[0013] A similarity acquisition module, configured to obtain the feature similarity associated with the i-th image migration model according to the first recognition feature and the second recognition feature, and add up the feature similarities associated with the M image migration models to obtain the cumulative feature similarity value;
[0014] An adversarial sample generation module, configured to update the first image according to the cumulative feature similarity value to generate an image adversarial sample, and combine the image adversarial sample and the second image into a test sample pair; the test sample pair is used to evaluate the performance of the image recognition model in the business application.
[0015] Among them, the migration model acquisition module is specifically configured to:
[0016] Acquire M image task models and M initial recognition models; one image task model corresponds to one initial recognition model;
[0017] Acquire a sample image, input the sample image into the M image task models, and obtain first sample recognition features of the sample image in each image task model;
[0018] Input the sample image into the M initial recognition models, and obtain second sample recognition features of the sample image in each initial recognition model;
[0019] According to the first sample recognition features of the sample image in each image task model and the second sample recognition features of the sample image in each initial recognition model, correct the network parameters of the M initial recognition models, and determine the M initial recognition models containing the corrected network parameters as M image transfer models.
[0020] Among them, the transfer model acquisition module acquires the sample image, including:
[0021] Acquire the original image set used to train the M initial recognition models, perform object detection on the original images in the original image set, and obtain the object location information corresponding to the original images in the original image set;
[0022] According to the object location information, perform segmentation processing on the original images in the original image set to obtain object region images, and adjust the size of the object region images to obtain object region images with a fixed size;
[0023] Add the object region images with a fixed size to the sample data set, and determine any one of the object region images in the sample data set as the sample image.
[0024] Among them, the transfer model acquisition module inputs the sample image into the M initial recognition models to obtain the second sample recognition features of the sample image in each initial recognition model, including:
[0025] Input the sample image into the i-th initial recognition model among the M initial recognition models; the i-th initial recognition model includes N convolutional components, and N is a positive integer;
[0026] Obtain the input features of the j-th convolutional component among the N convolutional components; when j is 1, the input features of the j-th convolutional component are the sample image, and j is a positive integer less than N;
[0027] According to one or more convolutional layers in the i-th convolutional component, perform a convolutional operation on the input features of the j-th convolutional component to obtain candidate convolutional features;
[0028] According to the weight vector corresponding to the normalization layer in the j-th convolutional component, perform normalization processing on the candidate convolutional features to obtain normalized features;
[0029] Combine the normalized features and the input features of the j-th convolutional component to obtain the output features of the j-th convolutional component, and use the output features of the j-th convolutional component as the input features of the (j + 1)-th convolutional component; the j-th convolutional component is connected to the (j + 1)-th convolutional component;
[0030] Determine the output features of the N-th convolutional component as the second sample recognition features of the sample image in the i-th initial recognition model.
[0031] Among them, the migration model acquisition module corrects the network parameters of the M initial recognition models according to the first sample recognition features of the sample image in each image task model and the second sample recognition features of the sample image in each initial recognition model, and determines the M initial recognition models including the corrected network parameters as M image migration models, including:
[0032] Classify the second sample recognition features of the sample image in the i-th initial recognition model to obtain the sample recognition result corresponding to the i-th initial recognition model;
[0033] Determine the error between the sample recognition result corresponding to the i-th initial recognition model and the label information corresponding to the sample image as the classification loss corresponding to the i-th initial recognition model;
[0034] Determine the sample similarity distance between the i-th initial recognition model and the M image task models according to the second sample recognition features of the sample image in the i-th initial recognition model and the first sample recognition features of the sample image in each image task model;
[0035] Determine the joint loss corresponding to the i-th initial recognition model according to the classification loss corresponding to the i-th initial recognition model and the sample similarity distance between the i-th initial recognition model and the M image task models;
[0036] Iteratively train the network parameters of the i-th initial recognition model according to the joint loss corresponding to the i-th initial recognition model, and stop training until the joint loss corresponding to the i-th initial recognition model meets the training end condition, and determine the i-th initial recognition model at the end of training as the i-th image migration model.
[0037] Among them, the migration model acquisition module determines the joint loss corresponding to the i-th initial recognition model according to the classification loss corresponding to the i-th initial recognition model and the sample similarity distance between the i-th initial recognition model and the M image task models, including:
[0038] Obtain the task weights corresponding to the M image task models, and perform weighted summation on the sample similarity distance between the i-th initial recognition model and the M image task models according to the task weights corresponding to the M image task models to obtain the distillation loss corresponding to the i-th initial recognition model;
[0039] Combine the classification loss corresponding to the i-th initial recognition model and the distillation loss corresponding to the i-th initial recognition model as the joint loss corresponding to the i-th initial recognition model.
[0040] Among them, the adversarial sample generation module updates the first image according to the feature similarity cumulative value to generate image adversarial samples, including:
[0041] Determine the adversarial noise corresponding to the first image according to the cumulative value of feature similarity, add the first image and the adversarial noise to obtain a candidate adversarial sample;
[0042] Obtain the candidate recognition features of the candidate adversarial sample in the i-th image migration model, and obtain the similarity distance between the candidate adversarial sample and the second image in the i-th image migration model according to the candidate recognition features and the second recognition features corresponding to the second image;
[0043] Accumulate the similarity distances between the candidate adversarial sample and the second image in M image migration models to obtain a joint adversarial loss, and perform minimization optimization processing on the joint adversarial loss to obtain the image adversarial sample corresponding to the first image.
[0044] Wherein, the device further includes:
[0045] An image feature extraction module, configured to input the image adversarial sample and the second image in the test sample pair into an image recognition model in a service application, and output, through the image recognition model, the adversarial recognition features corresponding to the image adversarial sample and the object recognition features corresponding to the second image;
[0046] A test result determination module, configured to obtain the image similarity between the adversarial recognition features and the object recognition features, and determine the test result of the image adversarial sample on the image recognition model according to the image similarity.
[0047] Wherein, the test result determination module determines the test result of the image adversarial sample on the image recognition model according to the image similarity, including:
[0048] If the image similarity is greater than the recognition threshold, it is determined that the test result of the image adversarial sample on the image recognition model is an attack success;
[0049] If the image similarity is less than or equal to the recognition threshold, it is determined that the test result of the image adversarial sample on the image recognition model is an attack failure.
[0050] Wherein, the number of test sample pairs is T, and T is an integer greater than 1;
[0051] The device further includes:
[0052] A model evaluation module, configured to obtain the test results of the image adversarial samples included in the T test sample pairs on the image recognition model, and count the attack success rate corresponding to the image recognition model among the test results corresponding to the T test sample pairs;
[0053] The model evaluation module is further configured to determine the model evaluation result corresponding to the image recognition model according to the attack success rate.
[0054] One aspect of the embodiments of the present application provides a computer device, including a memory and a processor. The memory is connected to the processor. The memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method provided in the above-mentioned aspect of the embodiments of the present application.
[0055] One aspect of the embodiments of the present application provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. The computer program is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method provided in the above-mentioned aspect of the embodiments of the present application.
[0056] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes the method provided in the above-mentioned aspect.
[0057] In the embodiments of the present application, M image migration models are obtained by knowledge migration of M image task models. The M image task models are pre-trained models applied to different image tasks, and one image task model corresponds to one image migration model. For an image pair composed of a first image and a second image, the first recognition feature corresponding to the first image and the second recognition feature corresponding to the second image can be obtained through the i-th image migration model among the M image migration models. At this time, the similarity between the first recognition feature and the second recognition feature can be used as the feature similarity associated with the i-th image migration model. The feature similarities associated with the M image migration models are added up to obtain a cumulative feature similarity value. The first image is updated according to the cumulative feature similarity value, and finally an image adversarial sample for attacking the image recognition model in the business application is obtained. The image adversarial sample at this time can contain the knowledge migrated from the M image task models, which can enhance the attack transferability of the image adversarial sample, and further improve the applicability of the image adversarial sample. When using the image adversarial sample to attack the image recognition model in the business application, the evaluation effect of the image recognition model can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0059] Figure 1It is a schematic structural diagram of a network architecture provided by an embodiment of the present application;
[0060] Figure 2 It is a schematic flowchart of an image processing method provided by an embodiment of the present application Figure 1 ;
[0061] Figure 3 It is a schematic diagram of the generation of an image adversarial example provided by an embodiment of the present application;
[0062] Figure 4 It is a schematic flowchart of an image processing method provided by an embodiment of the present application Figure 2 ;
[0063] Figure 5 It is a schematic diagram of a knowledge transfer process provided by an embodiment of the present application;
[0064] Figure 6 It is a schematic diagram of the performance evaluation of an image recognition model provided by an embodiment of the present application;
[0065] Figure 7 It is a schematic structural diagram of an image processing device provided by an embodiment of the present application;
[0066] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0067] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0068] For ease of understanding, the basic technical concepts related to the embodiments of the present application will be described first:
[0069] Computer Vision Technology (CV): Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as target recognition, positioning, and measurement, and further performs graphics processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0070] The embodiments of this application specifically relate to the image recognition technology under computer vision technology, and specifically propose an evaluation scheme for an image recognition model based on multi-task knowledge transfer. By using the knowledge transfer of different image tasks, image adversarial samples for attacking the image recognition model in business applications are generated. Using these image adversarial samples, the performance (specifically, the security of the image recognition model) of the image recognition model in different business applications can be evaluated, the applicability of the image adversarial samples can be improved, and the evaluation effect of the image recognition model can be enhanced.
[0071] Please refer to Figure 1 , Figure 1 FIG. is a schematic structural diagram of a network architecture provided by the embodiments of this application. The network architecture may include a server 10d and a terminal cluster. The terminal cluster may include one or more terminal devices, and the number of terminal devices included in the terminal cluster is not limited here. As Figure 1 shown, the terminal cluster may specifically include terminal devices 10a, 10b, and 10c, etc.; all terminal devices in the terminal cluster (for example, may include terminal devices 10a, 10b, and 10c, etc.) can be network-connected to the server 10d, so that each terminal device can perform data interaction with the server 10d through this network connection.
[0072] The terminal devices in the terminal cluster may include electronic devices such as smart phones, tablet computers, laptop computers, palm computers, mobile internet devices (MIDs), wearable devices (such as smart watches, smart bracelets, etc.), intelligent voice interaction devices, smart home appliances (such as smart TVs, etc.), vehicle-mounted devices, and aircraft. The type of terminal device is not limited in this application. It can be understood that, as Figure 1 each terminal device in the terminal cluster shown can install a business application. When the business application runs on each terminal device, it can perform data interaction with the Figure 1 server 10d shown above respectively. Among them, the business applications running on each terminal device can correspond to an independent client or an embedded sub-client integrated in a certain client. This application does not make a limitation on this.
[0073] Among them, the business application may specifically include, but is not limited to: browsers, vehicle-mounted applications, smart home applications, entertainment applications (such as game applications), multimedia applications (such as video applications, short video applications), conference applications, and social applications, etc., applications with image recognition functions. Among them, if the terminal device included in the terminal cluster is a vehicle-mounted device, then the vehicle-mounted device can be an intelligent terminal in the intelligent transportation scenario, and the business application running in the vehicle-mounted device can be called a vehicle-mounted application.
[0074] Among them, the server 10d can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The type of the server is not limited in this application.
[0075] It can be understood that Figure 1 one or more business applications can be installed in each of the terminal devices shown. One or more image recognition models for implementing different image tasks can be deployed in each business application. The image recognition model here can be any deep neural network model. The network structure of the image recognition model is not limited in this application. Among them, the business applications involved in the embodiments of this application may include, but are not limited to: mobile phone face unlocking, face login of application software, remote face verification, face recognition access control system, offline face payment, automatic face clearance, etc. business applications. In order to test the security of the image recognition models in the above business applications, the embodiments of this application can actively test the image recognition models in the above business applications by means of adversarial attacks to evaluate the vulnerabilities of the image recognition models.
[0076] Please refer to Figure 2 , Figure 2 , which is a schematic flowchart of an image processing method provided by an embodiment of the present application Figure 1 ; It can be understood that the image processing method can be executed by a computer device, and the computer device can be a server (such as Figure 1 the server 10d shown), or can be a terminal device (such as Figure 1 any terminal device in the terminal cluster shown), and the present application does not limit this. As Figure 2 shown, the image processing method can include the following steps S101 to step S104:
[0077] Step S101, obtain M image migration models; The M image migration models are recognition models obtained after knowledge migration of M image task models. The M image task models are pre-trained models applied to different image tasks. One image task model corresponds to one image migration model, and M is a positive integer.
[0078] In the embodiment of the present application, the M image task models can be pre-trained models applied to different image tasks, that is, network models that have been trained. Here, M represents the number of image tasks used in the knowledge migration process, and M is a positive integer. For example, M can take values of 1, 2,...; For each of the M image tasks, a trained image task model can be obtained. Among them, the M image tasks here can include but are not limited to: face attribute recognition, face segmentation, expression recognition, expression conversion, attribute editing and other tasks, and the M image task models can have different network structures; Optionally, the M image recognition models can be network models with the same network structure and different network parameters, and the present application does not limit the network structure of the M image recognition models.
[0079] In the process of knowledge migration, knowledge migration can be performed on the M image task models to obtain image migration models corresponding to the M image tasks respectively, that is, the M image migration models. Here, the image migration model can be a recognition model after knowledge migration. In the embodiment of the present application, one image task can correspond to one image task model and one image migration model; That is to say, one image task model can correspond to one image migration model, and the M image migration models can be used as proxy models for the image recognition models deployed in subsequent attack service applications. Among them, the specific implementation process of the knowledge migration process can be referred to in the subsequent content.
[0080] Among them, the M image migration models may refer to recognition models with the same network structure and different network parameters. For example, the initialization models of the M image migration models before training may be the same network model. After model training in the knowledge migration process, M image migration models can be obtained, and the M image migration models obtained thereby have different network parameters. Among them, the M image migration models can be deep neural networks with any network structure. For example, the network structure of each image migration model may include, but is not limited to: Convolutional Neural Networks (CNN), Feedforward Neural Network, Long Short-Term Memory (LSTM), Recurrent Convolutional Neural Network (RCNN), Attention Mechanism, Variational Autoencoder (VAE), a deformation of any of the above network structures, or a combination of any two or more of the above network structures, etc. The embodiments of the present application do not limit the network structure of the image migration models.
[0081] In step S102, a first image and a second image are obtained, and the first image and the second image are input into the i-th image migration model among the M image migration models. Through the i-th image migration model, a first recognition feature corresponding to the first image and a second recognition feature corresponding to the second image are obtained; different objects are included in the first image and the second image, and i is a positive integer less than or equal to M.
[0082] In the embodiments of the present application, when the knowledge migration process ends, M image migration models can be obtained. The inputs of these image migration models are two image data, namely the first image (which can also be called the attack image and can be denoted as x a ) and the second image (which can also be called the attacked image and can be denoted as x v ). Here, the first image and the second image can be RGB images (R: red, G: green, B: blue) pictures, or can be HSL (H: hue, S: saturation, L: lightness) images, or can be HSV (H: hue, S: saturation, V: value) images, or can be grayscale images, etc. The embodiments of the present application do not limit the types of the first image and the second image. Among them, different objects are included in the first image and the second image. For example, when both the first image and the second image are face images, the first image and the second image can be face images of different objects. For example, the first image is a face image of object A, and the second image is a face image of object B.
[0083] The first image x a and the second image x v It can be input into M image migration models, and the first image x is calculated in each image migration model a The corresponding first recognition feature, and the second image x v The corresponding second identification feature. If the M image migration models have the same network structure, the first image x a and the second image x v The forward calculation in each image transfer model is similar. a and the second image x v The forward calculation process in any image migration model included in the M image migration models (for example, the i-th image migration model, i is a positive integer less than or equal to M) is described. a and the second image x v Input into the i-th image migration model, and obtain the first image x through the i-th image migration model a The first recognition feature in the i-th image transfer model, and the second image x v The second identifying feature in the i-th image transfer model.
[0084] For example, assuming that the network structure of the i-th image transfer model is a convolutional neural network, and the i-th image transfer model includes N convolutional components, the N convolutional components can be residually connected in each image transfer model, or can be chained, and this application does not limit this. Wherein, N is a positive integer, such as N can be 1, 2, .... The following description takes the N convolutional components in the i-th image transfer model as an example of residual connection. The first image x a After being input into the i-th image transfer model, it is first input into the first convolutional component in the i-th image transfer model, and the first image x is converted according to one or more convolutional layers in the first convolutional component. a Perform convolution operation to obtain the candidate convolution features in the first convolution component. According to the weight vector corresponding to the normalization layer in the first convolution component, the candidate convolution features in the first convolution component are normalized to obtain the normalized features in the first convolution component. Compare the normalization layer in the first convolution component with the first image x a The output features of the first convolution component are combined to obtain the output features of the first convolution component. The output features of the first convolution component can be used as the input features of the second convolution group in the i-th image migration model. Similarly, the output features of the second convolution component can be calculated until the output features of the N-th convolution component in the i-th image migration model are calculated. Then, the output features of the N-th convolution component can be determined as the first image xa The first recognition feature in the i-th image migration model. By adopting the same calculation process as described above, the second image x can be obtained. v The second recognition feature in the i-th image migration model.
[0085] Similarly, the first image x can be obtained respectively. a The first recognition feature in each image migration model, that is, the first image x a There can be M corresponding first recognition features. The second image x can be obtained respectively. v The second recognition feature in each image migration model, that is, the second image x v There can be M corresponding second recognition features. For example, the first recognition feature in the embodiments of the present application may refer to the face recognition feature calculated after the first image x a is input into each image migration model; the second recognition feature may refer to the face recognition feature calculated after the second image x v is input into each image migration model.
[0086] Step S103: According to the first recognition feature and the second recognition feature, obtain the feature similarity associated with the i-th image migration model, and add up the feature similarities associated with the M image migration models to obtain a cumulative feature similarity value.
[0087] Specifically, the first recognition feature calculated by the first image x a in the i-th image migration model can be denoted as the i-th first recognition feature, and the second recognition feature calculated by the second image x v in the i-th image migration model can be denoted as the i-th second recognition feature. The feature similarity between the i-th first recognition feature and the i-th second recognition feature can be calculated, and the feature similarity at this time can be referred to as the feature similarity associated with the i-th image migration model.
[0088] In an embodiment of the present application, a similarity algorithm may be used to calculate the feature similarity between two identification features. The similarity algorithm may include, but is not limited to: cosine similarity, Euclidean distance, Manhattan distance, and Jaccard Similarity Coefficient, which is not limited in the present application. Through any of the above similarity algorithms, the feature similarity between the first identification feature and the second identification feature output by each image migration model can be obtained, that is, the feature similarities respectively associated with the M image migration models can be calculated. Furthermore, the feature similarities respectively associated with the M image migration models can be accumulated, that is, the M feature similarities are accumulated to obtain a feature similarity cumulative value.
[0089] Step S104, updating the first image according to the feature similarity cumulative value, generating an image adversarial sample, and combining the image adversarial sample and the second image into a test sample pair; the test sample pair is used to evaluate the performance of the image recognition model in the business application.
[0090] Specifically, after calculating the feature similarity cumulative value, the first image x can be calculated by gradient backpropagation (also called error backpropagation). a The corresponding adversarial noise σ can be used to update the first image so that the feature similarity between the updated first image and the second image is maximized, and finally an image adversarial sample is obtained. Among them, the image adversarial sample refers to the first image x a The new image obtained by making a small, targeted modification causes the first image that was originally correctly identified to be misclassified. Before generating the image adversarial sample, the first image x a Adding initial disturbance or noise, the initial disturbance or noise at this time can be called the initialized adversarial noise; the initialized adversarial noise can be zero, or can be randomly initialized, or initialized in other ways, which is not limited in this application. It can be understood that the image adversarial sample can refer to a new image obtained by adding the first image to the adversarial noise; that is, the update process of the first image can essentially be understood as an iterative update process of the adversarial noise.
[0091] In one or more embodiments, the adversarial noise corresponding to the first image may be determined in a gradient backpropagation manner according to the cumulative feature similarity value, and the first image and the adversarial noise are added to obtain a candidate adversarial sample; the candidate recognition feature of the candidate adversarial sample in the i-th image transfer model is obtained, and according to the candidate recognition feature and the second recognition feature corresponding to the second image, the similarity distance between the candidate adversarial sample and the second image in the i-th image transfer model is obtained; the similarity distances between the candidate adversarial sample and the second image in the M image transfer models are accumulated to obtain a joint adversarial loss, and the joint adversarial loss is minimized and optimized to obtain the image adversarial sample corresponding to the first image.
[0092] Among them, the joint adversarial loss can be shown as in formula (1):
[0093] L adv =∑ i∈M (1 - sim(F i (x a +σ)+F i (x v ))) (1)
[0094] Among them, L adv in formula (1) represents the joint adversarial loss, σ represents the adversarial noise calculated by the gradient backpropagation method, x a +σ represents the candidate adversarial sample, F i (·) represents the i-th image transfer model, sim represents the cosine similarity; sim(F i (x a +σ)+F i (x v ) represents the feature similarity between the candidate adversarial sample x a +σ and the second image x v in the i-th image transfer model; among them, F i (x a +σ) represents the candidate recognition feature of the candidate adversarial sample x a +σ in the i-th image transfer model, and the adversarial noise σ at this time can refer to the adversarial noise calculated after the cumulative feature similarity value is input into the i-th image transfer model in a gradient backpropagation manner; F i (x v ) represents the second recognition feature of the second image x v in the i-th image transfer model. 1 - sim(F i (x a +σ)+F i (x v )) represents the candidate adversarial sample x a +σ and the second image x vThe feature distance in the i-th image transfer model; the feature distance can be the result of subtracting the feature similarity from 1. The smaller the similarity distance, the greater the feature similarity. a +σ and the second image x v By accumulating the similar distances in each image transfer model, we can get the joint adversarial loss L adv .
[0095] By using the joint adversarial loss L in the above formula (1) adv After minimization and optimization, we can get the final image adversarial sample x adv , the way to optimize the joint adversarial loss can be shown as formula (2):
[0096]
[0097] Among them, L in formula (2) adv (F i (x a +σ)+F i (x v )) represents the joint adversarial loss L in the above formula (1) adv ; Represents the first image x a After iterative update, the candidate adversarial sample β represents the p-norm upper bound used to constrain the adversarial noise σ. Formula (2) can be expressed as: through continuous iterative update, a joint adversarial loss L is calculated. adv When the minimum value is reached (candidate adversarial sample), or it can be understood as calculating a joint adversarial loss L adv The adversarial noise σ reaches the minimum value; at this time, the joint adversarial loss can be Determined as image adversarial sample x adv .
[0098] Furthermore, the image adversarial sample x adv and the second image x v As a test sample pair, the test sample pair can be used to attack an image recognition model in any business application to evaluate the security of the image recognition model in the business application. It can be understood that a large number of test sample pairs can be generated for the image recognition model in the business application through the above steps S102 to S104.
[0099] See also Figure 3 , Figure 3It is a schematic diagram for generating an image adversarial sample provided by an embodiment of the present application. Taking M = 2 as an example, the embodiment of the present application describes the generation process of the image adversarial sample. As Figure 3 shown, after the knowledge transfer process ends, two image transfer models can be obtained, denoted as image transfer model 21a and image transfer model 22a respectively. After obtaining the first image 20a and the second image 20b, the first image 20a and the second image 20b can be input into the image transfer model 21a. Through this image transfer model 21a, the first recognition feature 21b corresponding to the first image 20a and the second recognition feature 21c corresponding to the second image 20b can be calculated; furthermore, the feature similarity 21d between the first recognition feature 21b and the second recognition feature 21c can be calculated. Similarly, the first image 20a and the second image 20b can be input into the image transfer model 22a. Through this image transfer model 22a, the first recognition feature 22b corresponding to the first image 20a and the second recognition feature 22c corresponding to the second image 20b can be calculated; furthermore, the feature similarity 22d between the first recognition feature 22b and the second recognition feature 22c can be calculated.
[0100] By adding the feature similarity 21d and the feature similarity 22d, a cumulative feature similarity value 23a can be obtained. The cumulative feature similarity value 23a is used to calculate the adversarial noise in the form of gradient backpropagation. Furthermore, through continuous iterative updates using the above formulas (1) and (2), an adversarial noise 23b that minimizes the joint adversarial loss L adv can be obtained. At this time, the adversarial noise 23b is added to the first image 20a to generate the final image adversarial sample 23c.
[0101] In an embodiment of the present application, M image migration models are obtained by performing knowledge migration on M image task models. The M image task models are pre-trained models applied in different image tasks, and one image task model corresponds to one image migration model. For an image pair consisting of a first image and a second image, the first recognition feature corresponding to the first image and the second recognition feature corresponding to the second image can be obtained by the i-th image migration model in the M image migration models; at this time, the similarity between the first recognition feature and the second recognition feature can be used as the feature similarity associated with the i-th image migration model. The feature similarities associated with the M image migration models are added to obtain a feature similarity cumulative value; the first image is updated according to the feature similarity cumulative value, and finally an image adversarial sample for attacking the image recognition model in the business application is obtained. At this time, the image adversarial sample can contain knowledge migrated from the M image task models, which can enhance the attack transferability of the image adversarial sample, and thus improve the applicability of the image adversarial sample; when the image adversarial sample is used to attack the image recognition model in the business application, the evaluation effect of the image recognition model can be improved.
[0102] See also Figure 4 , Figure 4 This is a schematic diagram of an image processing method provided in an embodiment of the present application. Figure 2 It can be understood that the image processing method can be executed by a computer device, which can be a server or a terminal device, and this application does not limit this. Figure 4 As shown, the image processing method may include the following steps S201 to S209:
[0103] Step S201, obtaining M image task models and M initial recognition models; the M image task models refer to pre-trained models applied in different image tasks, and one image task model corresponds to one initial recognition model.
[0104] In the embodiment of the present application, the M image task models may refer to pre-trained models applied in different image tasks, that is, network models that have been trained. The relevant description of the M image task models can be found in the aforementioned Figure 2 The relevant description in step S101 of the corresponding embodiment will not be repeated here. The M initial recognition models can be network models with the same network structure and initialized in the same way. For example, the M initial recognition models can be convolutional neural networks with the same structure. Optionally, the M initial recognition models can also be network models with different network structures, and the present application does not limit the network structure of the M initial recognition models. One image task model can correspond to one initial recognition model.
[0105] Step S202: Obtain a sample image, input the sample image into M image task models, and obtain the first sample recognition features of the sample image in each image task model.
[0106] Specifically, obtain the original image set for training M initial recognition models, perform object detection on the original images in the original image set to obtain the object location information corresponding to the original images in the original image set; according to the object location information, perform segmentation processing on the original images in the original image set to obtain object region images, adjust the sizes of the object region images to obtain object region images with a fixed size; add the object region images with a fixed size to the sample data set, and determine any one of the object region images in the sample data set as the sample image. Among them, the images in the sample data set can all be referred to as sample images, and all the sample images have a fixed size. Optionally, the sample data set can be any currently open-source face image database, or can be other image databases, and this application does not make any limitations in this regard.
[0107] For example, if the M image task models are face device models applied in M face recognition tasks, then the original images in the above original image set can be face images. The face images in the original image set can be registered using a face registration tool. For example, the facial feature coordinates in the face images included in the original image set can be detected through certain algorithms, and the facial feature coordinates at this time are the detected object location information; through the facial feature coordinates returned by the face registration tool, the face part is segmented from the face image to obtain a face region image (object region image); furthermore, the size of the segmented face region image can be adjusted to a fixed size, such as 112*112*3. Here, 112*112 represents the width and height of the face region image after size adjustment, and 3 represents the number of channels of the face region image after size adjustment. At this time, the face region image is a three-channel color image.
[0108] Input the sample image into the M image task models. In each image task model, the first sample recognition features corresponding to the sample image can be calculated; in other words, each image task model can perform calculations on the input sample image to obtain a first sample recognition feature. The number of the first sample recognition features here is M.
[0109] Step S203: Input the sample image into the M initial recognition models, and obtain the second sample recognition features of the sample image in each initial recognition model.
[0110] Specifically, the sample image can be input into M initial recognition models, and the second sample recognition feature corresponding to the sample image can be calculated in each initial recognition model; in other words, each initial recognition model can calculate the input sample image to obtain a second sample recognition feature, and the number of the second sample recognition features here is M.
[0111] For example, assume that the network structure of the i-th initial recognition model among the M initial recognition models is a convolutional neural network, and the i-th image migration model includes N convolutional components, and these N convolutional components are connected in a residual connection manner; then the sample image can be input into the i-th initial recognition model among the M initial recognition models; the i-th initial recognition model includes N convolutional components, where N is a positive integer; obtain the input feature of the j-th convolutional component among the N convolutional components; when j is 1, the input feature of the j-th convolutional component is the sample image, and j is a positive integer less than N; perform a convolutional operation on the input feature of the j-th convolutional component according to one or more convolutional layers in the i-th convolutional component to obtain a candidate convolutional feature; perform a normalization process on the candidate convolutional feature according to the weight vector corresponding to the normalization layer in the j-th convolutional component to obtain a normalized feature; combine the normalized feature and the input feature of the j-th convolutional component to obtain the output feature of the j-th convolutional component, and use the output feature of the j-th convolutional component as the input feature of the (j + 1)-th convolutional component; the j-th convolutional component is connected to the (j + 1)-th convolutional component; determine the output feature of the N-th convolutional component as the second sample recognition feature of the sample image in the i-th initial recognition model. Among them, the calculation process of the second sample recognition feature can refer to the relevant description of the first recognition feature in step S102 of the corresponding embodiment described above, and details will not be repeated here. Figure 2 For the relevant description of the first recognition feature in step S102 of the corresponding embodiment, details will not be repeated here.
[0112] Step S204: According to the first sample recognition features of the sample image in each image task model and the second sample recognition features of the sample image in each initial recognition model, correct the network parameters of the M initial recognition models, and determine the M initial recognition models including the corrected network parameters as M image migration models.
[0113] Specifically, the second sample recognition features of the sample image in each initial recognition model can be classified to obtain the sample recognition result of the sample image in each initial recognition model. At this time, the sample recognition result is the object recognition result obtained after each initial recognition model performs object recognition processing on the sample image. It can be understood that each sample image in the sample dataset can carry a label information, and this label information can be used to characterize the category of the object contained in the sample image. For example, when the sample image is an image containing a real face, after each initial recognition model performs face recognition processing on the sample image, the face recognition result corresponding to the face contained in the sample image can be obtained. Here, the face recognition result can be used to characterize which object the real face contained in the sample image belongs to.
[0114] According to the sample recognition result of the sample image in each initial recognition model and the label information corresponding to the sample image, calculate the classification loss corresponding to each initial recognition model. Among them, the object recognition processing processes of the sample image in each initial recognition model are similar. Hereinafter, any one of the M initial recognition models (such as the i-th initial recognition model, where i is a positive integer less than or equal to M, and M is an integer greater than 1) will be taken as an example for description. The second sample recognition features of the sample image in the i-th initial recognition model can be classified to obtain the sample recognition result corresponding to the i-th initial recognition model; the error between the sample recognition result corresponding to the i-th initial recognition model and the label information corresponding to the sample image is determined as the classification loss corresponding to the i-th initial recognition model.
[0115] During the knowledge transfer process, the sample data in the sample dataset can be input into the M image task models and the M initial recognition models simultaneously, and a classification loss and a distillation loss are calculated for each initial recognition model; the classification loss of an initial recognition model can be used to characterize the error between the sample recognition result output by the initial recognition model and the label information of the sample image; the distillation loss of an initial recognition model can be calculated through the sample similarity distance between the second sample recognition features output by the initial recognition model and the first sample recognition features output by each image task model. Among them, the sample similarity distance can be determined by calculating the feature similarity between the second sample recognition features and the first sample recognition features. For example, the sample similarity distance can be the difference between the value 1 and the feature similarity; for example, when the feature similarity is the cosine similarity calculated using the cosine similarity algorithm, then the sample similarity distance here can be the cosine distance.
[0116] Among them, the calculation process of the distillation loss corresponding to each of the M initial recognition models is similar. Below, any one of the M initial recognition models (the i-th initial recognition model) is taken as an example for description. The sample similarity distance between the i-th initial recognition model and the M image task models can be determined according to the second sample recognition feature of the sample image in the i-th initial recognition model and the first sample recognition feature of the sample image in each image task model; according to the classification loss corresponding to the i-th initial recognition model and the sample similarity distance between the i-th initial recognition model and the M image task models, the joint loss corresponding to the i-th initial recognition model is determined.
[0117] Optionally, the calculation process of the joint loss corresponding to the i-th initial recognition model may include: obtaining the task weights corresponding to the M image task models, and performing a weighted sum on the sample similarity distance between the i-th initial recognition model and the M image task models according to the task weights corresponding to the M image task models to obtain the distillation loss corresponding to the i-th initial recognition model; combining the classification loss corresponding to the i-th initial recognition model and the distillation loss corresponding to the i-th initial recognition model into the joint loss corresponding to the i-th initial recognition model.
[0118] Among them, the joint loss corresponding to the i-th initial recognition model can be shown as in formula (3):
[0119]
[0120] Among them, L tr,comb represents the joint loss corresponding to the i-th initial recognition model, L FR (F i ′) represents the classification loss corresponding to the i-th initial recognition model, d(F i ′, g i ) represents the sample similarity distance between the i-th initial recognition model and the i-th image task model; among them, F i ′ represents the second sample recognition feature calculated from the sample image in the i-th initial recognition model; g i represents the first sample recognition feature calculated from the sample image in the i-th image task model. For example, d(F′ i , g i ) can be expressed as d(F′ i , g i ) = 1 - sim(F′ i , g i ), where sim(F′ i , g i) represents the cosine similarity between the second sample recognition feature calculated from the sample image in the i-th initial recognition model and the first sample recognition feature calculated from the sample image in the i-th image task model. represents the task weight corresponding to the i-th image task model during the knowledge transfer process, and this task weight can be predefined; λ represents the constraint parameter for restricting the distillation loss. represents the distillation loss corresponding to the i-th initial recognition model. Optionally, when M takes the value of 1, it means that only the knowledge in one image task model needs to be transferred during the knowledge transfer process. At this time, the task weight corresponding to this image task model can be set to 1, and only one initial recognition model needs to be trained during the knowledge transfer process. Then, the above formula (3) can be simplified to: L tr,comb = L FR (F′) + λd(F′, g); where L FR (F′) represents the classification loss of the initial recognition model, and d(F′, g) represents the distillation loss of the initial recognition model.
[0121] Further, the network parameters of the i-th initial recognition model can be iteratively trained according to the joint loss corresponding to the i-th initial recognition model. When the joint loss corresponding to the i-th initial recognition model satisfies the training end condition, the training is stopped, and the i-th initial recognition model at the end of the training is determined as the i-th image transfer model. Here, the training end condition can be that the number of training times of the i-th initial recognition model reaches the maximum number of iterations, or it can be that the joint loss reaches the convergence condition. This application does not limit the setting of the training condition. When the i-th initial recognition model reaches the training end condition, the training of the i-th initial recognition model can be stopped, and the network parameters at the time of stopping the training are saved, and the network parameters saved here are used as the i-th image transfer model.
[0122] In the embodiments of this application, the network parameters of the i-th initial recognition model can be jointly optimized through the above formula (3). In this way, the knowledge in the M image task models can be transferred to the currently trained i-th initial recognition model. The update process of the network parameters of the i-th initial recognition model can be shown as the following formula (4):
[0123]
[0124] where, represents the network parameters of the i-th initial recognition model, i represents the current knowledge transfer from the i-th image task model; p represents the p-th update, and α1 represents the learning rate; represents calculating the joint loss L tr,comb under the network parameters of the (p - 1)-th round, and for the network parameters of the (p - 1)-th round Derivation is performed. Through the above formulas (3) and (4), the i-th image migration model can be trained. Similarly, by adopting the same training method, M trained image migration models can be obtained.
[0125] Please refer to Figure 5 , Figure 5 which is a schematic diagram of a knowledge migration process provided by an embodiment of the present application; for ease of understanding, the embodiment of the present application is described by taking M = 1 as an example. As Figure 5 shown, an image task model 30b can be obtained, and an initial recognition model 30c for performing knowledge migration can be constructed. After obtaining the sample image 30a from the sample dataset, the sample image 30a can be input into the image task model 30b and the initial recognition model 30c simultaneously. By calculating the sample image 30a through the image task model 30b, the first sample recognition feature 30d can be obtained. Similarly, the second sample recognition feature 30e can be obtained by calculating the sample image 30a through the initial recognition model 30c.
[0126] Calculate the sample similarity distance 30h between the first sample recognition feature 30d and the second sample recognition feature 30e. Here, the sample similarity distance 30h can be used as the distillation loss corresponding to the initial recognition model 30c; according to the error between the second sample recognition feature 30e and the label information 30f, calculate the classification loss 30g corresponding to the initial recognition model 30c. Further, the sample similarity distance 30h and the classification loss 30g can be combined into a joint loss 30i, and the network parameters of the initial recognition model 30c can be optimized according to the joint loss 30i. After the knowledge migration is completed, an image migration model can be obtained.
[0127] Step S205: Obtain a first image and a second image, and input the first image and the second image into the i-th image migration model among the M image migration models. Through the i-th image migration model, obtain the first recognition feature corresponding to the first image and the second recognition feature corresponding to the second image; different objects are included in the first image and the second image, and i is a positive integer less than or equal to M.
[0128] Step S206: According to the first recognition feature and the second recognition feature, obtain the feature similarity associated with the i-th image migration model, and add up the feature similarities associated with the M image migration models to obtain a cumulative feature similarity value.
[0129] Step S207: Update the first image according to the cumulative feature similarity value to generate an image adversarial sample, and combine the image adversarial sample and the second image into a test sample pair.
[0130] Among them, the specific implementation process of steps S205 to S207 can be referred to the foregoing Figure 2For the relevant descriptions in steps S102 to S104 of the corresponding embodiment, no further elaboration will be provided here.
[0131] Step S208: Input the image adversarial sample and the second image in the test sample pair into the image recognition model in the business application, and output the adversarial recognition features corresponding to the image adversarial sample and the object recognition features corresponding to the second image through the image recognition model.
[0132] Specifically, a pair of a first image and a second image containing different objects can be obtained (for example, the first image can be a face image containing object A, and the second image can be a face image containing object B). Through the M image migration models after knowledge transfer, the first image can be iteratively updated multiple times, and finally the image adversarial sample corresponding to the first image can be obtained. The generated image adversarial sample and the second image are used as a pair of test sample pairs for attacking the image recognition model in the business application. In the embodiments of the present application, the purpose of using the M image migration models to update the first image is to maximize the feature similarity between the updated first image and the second image, that is, the purpose of using each image migration model to update the first image is to make the finally generated image adversarial sample as similar as possible to the second image. However, in essence, the object contained in the image adversarial sample (which is the same object as the object contained in the first image) and the object contained in the second image are not the same object.
[0133] For each test sample pair, the image adversarial sample and the second image in a test sample pair can be input into the image recognition model in the business application, and the image recognition model can be a black-box network model; and the aforementioned image migration model used to generate the image adversarial sample can be used as a white-box network model for testing the black-box network model in the business application. Among them, the black-box network model can refer to a network model that only knows the model input and the model output, but does not know the internal structure of the model. In the embodiments of the present application, it can refer to the image recognition model that has been launched in the business application. The white-box network model is a network model that knows the model input, the model output, and the internal structure of the model, such as the M image migration models obtained at the end of the knowledge transfer process in the embodiments of the present application. Optionally, the M image task models used for knowledge transfer can be black-box network models, or can be white-box network models. The present application does not make any limitations in this regard.
[0134] Among them, the business application can be any business application that requires model testing. For example, the business application can be any face-swiping related product that requires face verification. At this time, the M image migration models after knowledge migration, as well as the models to be tested in the business application, can all be face recognition models; the business applications at this time can be related applications such as face unlocking, face login, face payment, face clearance, and face access control. Optionally, the business application can be any related product that requires human pose verification. At this time, the M image migration models after knowledge migration, as well as the models to be tested in the business application, can all be pose recognition models; the business applications at this time can be related applications such as human pose payment, human pose clearance, human pose access control, and human pose verification. In the image recognition model of the business application, object recognition processing can be performed on the image adversarial sample to obtain an adversarial recognition feature; object recognition processing can be performed on the second image to obtain an object recognition feature.
[0135] Step S209: Obtain the image similarity between the adversarial recognition feature and the object recognition feature, and determine the test result of the image adversarial sample on the image recognition model according to the image similarity.
[0136] Specifically, for each image adversarial sample and the second image in the test sample, the image similarity between the adversarial recognition feature and the object recognition feature can be obtained. If the image similarity is greater than the recognition threshold, it can be determined that the test result of the image adversarial sample on the image recognition model is a successful attack; a successful attack can indicate that the image recognition model in the business application recognizes the objects in the image adversarial sample and the second image as the same object, that is, the image recognition model misidentifies the image adversarial sample, and the image adversarial sample successfully attacks the image recognition model in the business application; if the image similarity is less than or equal to the recognition threshold, it is determined that the test result of the image adversarial sample on the image recognition model is a failed attack; a failed attack can indicate that the image recognition model in the business application recognizes the objects in the image adversarial sample and the second image as different objects, that is, the image recognition model correctly identifies the image adversarial sample, and the image adversarial sample fails to attack the image recognition model in the business application. Among them, the recognition threshold here can be preset according to the specific requirements of the actual application scenario. For example, it can be set to 0.8, 0.9, etc. This application does not make any limitations in this regard.
[0137] In one or more embodiments, a large number of test sample pairs can be generated for an image recognition model in a business application. For example, the number of test samples can be denoted as T, where T is an integer greater than 1. Obtain the test results of the image adversarial samples included in the T test sample pairs on the image recognition model. Among the test results corresponding to the T test sample pairs, count the attack success rate of the image recognition model. Determine the model evaluation result corresponding to the image recognition model according to the attack success rate. Among them, the higher the attack success rate, the worse the performance (security) of the tested image recognition model, and the lower the attack success rate, the better the performance (security) of the tested image recognition model. Optionally, an attack threshold can be set in advance. If the attack success rate corresponding to the image recognition model tested in the business application is greater than the attack threshold, it can be determined that the model evaluation result corresponding to the tested image recognition model is model anomaly. If the attack success rate corresponding to the image recognition model tested in the business application is less than or equal to the attack threshold, it can be determined that the model evaluation result corresponding to the tested image recognition model is model secure. Among them, the attack threshold here can be set in advance according to the specific requirements of the actual application scenario. For example, it can be set to 0.6, 0.7, etc., and this application does not make any limitations in this regard.
[0138] Please refer to Figure 6 , Figure 6 which is a schematic diagram of the performance evaluation of an image recognition model provided by an embodiment of this application. As Figure 6 shown, if it is necessary to test an image recognition model in a certain business application, a model test page 40b of this business application can be displayed in a terminal device 40a (which can be understood as a computer device). In this model test page 40b, multiple test sample pairs for testing the above image recognition model can be displayed, such as test sample pair 41a (which can include image adversarial sample 42a and second image 42b), test sample pair 41b (which can include image adversarial sample 42c and second image 42d), etc. In this model test page 40b, a sample test pair for testing the image recognition model in this business application can be selected. It can be understood that in this model test page 40b, one test sample pair can be selected each time for model testing, or multiple test sample pairs can be selected in batches for model testing, or all test sample pairs can be imported into the image recognition model of the business application at one time. This application does not make any limitations in this regard.
[0139] For example, as Figure 6As shown in the figure, on the model test page 40b, a test sample pair 41a is selected to test the image recognition model in the business application. After the test sample object 41a is selected, a trigger operation can be performed on the "Enter Model Test" control 40c on the model test page 40b. At this time, the terminal device 40a can input the image adversarial sample 42a and the second image 42b included in the test sample pair 41 into the image recognition model 43a in the business application. Through the image recognition model 43a, the adversarial recognition feature 43b corresponding to the image adversarial sample 42a and the object recognition feature 43c corresponding to the second image 42b can be output; furthermore, the similarity between the adversarial recognition feature 43b and the object recognition feature 43c can be calculated. At this time, the similarity can refer to the image similarity 43d between the image adversarial sample 42a and the second image 42b. Among them, the calculation method of the image similarity can be the same as the calculation method of the aforementioned feature similarity, which will not be described here. By comparing the image similarity 43d with a preset recognition threshold 43e, the test result 43f of the test sample pair 41a for the image recognition model 43a can be obtained. If the image similarity 43d is greater than the recognition threshold 43e, it means that the test result 43f of the image adversarial sample 42a for the image recognition model 43a is an attack success; if the image similarity 43d is less than or equal to the recognition threshold 43e, it means that the test result 43f of the image adversarial sample 42a for the image recognition model 43a is an attack failure.
[0140] In the embodiments of the present application, M image transfer models are obtained by performing knowledge transfer on M image task models. The M image task models are pre-trained models applied to different image tasks, and one image task model corresponds to one image transfer model, so that the feature differences between different image tasks can be eliminated. For an image pair composed of a first image and a second image, the first recognition feature corresponding to the first image and the second recognition feature corresponding to the second image can be obtained through the i-th image transfer model among the M image transfer models; at this time, the similarity between the first recognition feature and the second recognition feature can be used as the feature similarity associated with the i-th image transfer model. Add up the feature similarities associated with the M image transfer models to obtain a cumulative feature similarity value; update the first image according to the cumulative feature similarity value, and finally obtain an image adversarial sample for attacking the image recognition model in the business application. At this time, the image adversarial sample can contain the knowledge transferred from the M image task models, which can improve the optimization efficiency of the image adversarial sample, enhance the attack transferability and generalization of the image adversarial sample, and further improve the applicability of the image adversarial sample; when using this image adversarial sample to attack the image recognition model in the business application, the evaluation effect of the image recognition model can be improved.
[0141] It is understandable that in the specific embodiments of the present application, images of human body parts such as the user's face image may be involved. When the above embodiments of the present application are applied to specific products or technologies, permission or consent from relevant institutions or departments, or the user himself / herself is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in the relevant regions.
[0142] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an image processing device provided by an embodiment of the present application. As Figure 7 shown, the image processing device 1 includes: a migration model acquisition module 11, an image recognition module 12, a similarity acquisition module 13, and an adversarial sample generation module 14;
[0143] The migration model acquisition module 11 is used to acquire M image migration models; the M image migration models are recognition models obtained by performing knowledge migration on M image task models, and the M image task models are pre-trained models applied to different image tasks. One image task model corresponds to one image migration model, and M is a positive integer;
[0144] The image recognition module 12 is used to acquire a first image and a second image, input the first image and the second image into the i-th image migration model among the M image migration models, and obtain a first recognition feature corresponding to the first image and a second recognition feature corresponding to the second image through the i-th image migration model; different objects are included in the first image and the second image, and i is a positive integer less than or equal to M;
[0145] The similarity acquisition module 13 is used to obtain the feature similarity associated with the i-th image migration model according to the first recognition feature and the second recognition feature, and add up the feature similarities associated with the M image migration models to obtain a cumulative feature similarity value;
[0146] The adversarial sample generation module 14 is used to update the first image according to the cumulative feature similarity value to generate an image adversarial sample, and combine the image adversarial sample and the second image into a test sample pair; the test sample pair is used to evaluate the performance of the image recognition model in the business application.
[0147] In one or more embodiments, the migration model acquisition module 11 is specifically used for:
[0148] acquire M image task models and M initial recognition models; one image task model corresponds to one initial recognition model;
[0149] acquire a sample image, input the sample image into the M image task models, and obtain first sample recognition features of the sample image in each image task model;
[0150] Input the sample image into M initial recognition models to obtain the second sample recognition features of the sample image in each initial recognition model;
[0151] According to the first sample recognition features of the sample image in each image task model and the second sample recognition features of the sample image in each initial recognition model, correct the network parameters of the M initial recognition models, and determine the M initial recognition models containing the corrected network parameters as M image transfer models.
[0152] In one or more embodiments, the transfer model acquisition module 11 acquires a sample image, including:
[0153] Acquire the original image set used to train the M initial recognition models, perform object detection on the original images in the original image set, and obtain the object location information corresponding to the original images in the original image set;
[0154] According to the object location information, perform segmentation processing on the original images in the original image set to obtain object region images, and resize the object region images to obtain object region images with a fixed size;
[0155] Add the object region images with a fixed size to the sample data set, and determine any one of the object region images in the sample data set as the sample image.
[0156] In one or more embodiments, the transfer model acquisition module 11 inputs the sample image into the M initial recognition models to obtain the second sample recognition features of the sample image in each initial recognition model, including:
[0157] Input the sample image into the i-th initial recognition model among the M initial recognition models; the i-th initial recognition model includes N convolutional components, and N is a positive integer;
[0158] Obtain the input features of the j-th convolutional component among the N convolutional components; when j is 1, the input features of the j-th convolutional component are the sample image, and j is a positive integer less than N;
[0159] Perform a convolution operation on the input features of the j-th convolutional component according to one or more convolutional layers in the i-th convolutional component to obtain candidate convolutional features;
[0160] Perform normalization processing on the candidate convolutional features according to the weight vector corresponding to the normalization layer in the j-th convolutional component to obtain normalized features;
[0161] Combine the normalized features and the input features of the j-th convolutional component to obtain the output features of the j-th convolutional component, and use the output features of the j-th convolutional component as the input features of the (j + 1)-th convolutional component; the j-th convolutional component is connected to the (j + 1)-th convolutional component;
[0162] Determine the output features of the N-th convolutional component as the second sample recognition features of the sample image in the i-th initial recognition model.
[0163] In one or more embodiments, the transfer model acquisition module 11 corrects the network parameters of the M initial recognition models according to the first sample recognition features of the sample image in each image task model and the second sample recognition features of the sample image in each initial recognition model, and determines the M initial recognition models including the corrected network parameters as M image transfer models, including:
[0164] Classify the second sample recognition features of the sample image in the i-th initial recognition model to obtain the sample recognition result corresponding to the i-th initial recognition model;
[0165] Determine the error between the sample recognition result corresponding to the i-th initial recognition model and the label information corresponding to the sample image as the classification loss corresponding to the i-th initial recognition model;
[0166] Determine the sample similarity distance between the i-th initial recognition model and the M image task models according to the second sample recognition features of the sample image in the i-th initial recognition model and the first sample recognition features of the sample image in each image task model;
[0167] Determine the joint loss corresponding to the i-th initial recognition model according to the classification loss corresponding to the i-th initial recognition model and the sample similarity distance between the i-th initial recognition model and the M image task models;
[0168] Iteratively train the network parameters of the i-th initial recognition model according to the joint loss corresponding to the i-th initial recognition model, and stop training until the joint loss corresponding to the i-th initial recognition model meets the training end condition, and determine the i-th initial recognition model at the end of training as the i-th image transfer model.
[0169] In one or more embodiments, the transfer model acquisition module 11 determines the joint loss corresponding to the i-th initial recognition model according to the classification loss corresponding to the i-th initial recognition model and the sample similarity distance between the i-th initial recognition model and the M image task models, including:
[0170] Obtain the task weights corresponding to M image task models. According to the task weights corresponding to the M image task models, perform a weighted sum of the sample similarity distances between the i-th initial recognition model and the M image task models to obtain the distillation loss corresponding to the i-th initial recognition model;
[0171] Combine the classification loss corresponding to the i-th initial recognition model and the distillation loss corresponding to the i-th initial recognition model into the joint loss corresponding to the i-th initial recognition model.
[0172] In one or more embodiments, the adversarial sample generation module 14 updates the first image according to the feature similarity cumulative value to generate an image adversarial sample, including:
[0173] Determine the adversarial noise corresponding to the first image according to the feature similarity cumulative value, and add the first image and the adversarial noise to obtain a candidate adversarial sample;
[0174] Obtain the candidate recognition features of the candidate adversarial sample in the i-th image migration model, and according to the candidate recognition features and the second recognition features corresponding to the second image, obtain the similarity distance between the candidate adversarial sample and the second image in the i-th image migration model;
[0175] Accumulate the similarity distances between the candidate adversarial sample and the second image in the M image migration models to obtain a joint adversarial loss, and perform a minimization optimization process on the joint adversarial loss to obtain the image adversarial sample corresponding to the first image.
[0176] In one or more embodiments, the image processing device 1 further includes: an image feature extraction module 15 and a test result determination module 16;
[0177] The image feature extraction module 15 is configured to input the image adversarial sample and the second image in the test sample pair into the image recognition model in the business application, and output the adversarial recognition features corresponding to the image adversarial sample and the object recognition features corresponding to the second image through the image recognition model;
[0178] The test result determination module 16 is configured to obtain the image similarity between the adversarial recognition features and the object recognition features, and determine the test result of the image adversarial sample on the image recognition model according to the image similarity.
[0179] In one or more embodiments, the test result determination module 16 determines the test result of the image adversarial sample on the image recognition model according to the image similarity, including:
[0180] If the image similarity is greater than the recognition threshold, determine that the test result of the image adversarial sample on the image recognition model is an attack success;
[0181] If the image similarity is less than or equal to the recognition threshold, it is determined that the test result of the image adversarial sample against the image recognition model is an attack failure.
[0182] In one or more embodiments, the number of test sample pairs is T, and T is an integer greater than 1;
[0183] The image processing apparatus 1 further includes: a model evaluation module 17;
[0184] The model evaluation module 17 is configured to obtain the test results of the image adversarial samples included in the T test sample pairs against the image recognition model, and count the attack success rate corresponding to the image recognition model among the test results corresponding to the T test sample pairs;
[0185] The model evaluation module 17 is further configured to determine the model evaluation result corresponding to the image recognition model according to the attack success rate.
[0186] According to an embodiment of the present application, the steps involved in the image processing method described above Figure 2 can be executed by each module in the image processing apparatus 1 shown Figure 7 For example, Figure 2 the step S101 shown can be executed by Figure 7 the transfer model acquisition module 11 shown, Figure 2 the step S102 shown can be executed by Figure 7 the image recognition module 12 shown, Figure 2 the step S103 shown can be executed by Figure 7 the similarity acquisition module 13 shown, Figure 2 the step S104 shown can be executed by Figure 7 the adversarial sample generation module 14 shown, etc.
[0187] According to an embodiment of the present application, Figure 7 each module in the image processing apparatus 1 shown can be separately or all combined into one or several units to form, or a certain one (or some) of the units can be further split into at least two smaller sub-units in terms of function, and the same operation can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In actual applications, the function of one module can also be realized by at least two units, or the functions of at least two modules can be realized by one unit. In other embodiments of the present application, the image processing apparatus 1 may also include other units. In actual applications, these functions can also be assisted by other units and can be realized by the cooperation of at least two units.
[0188] In an embodiment of the present application, M image migration models are obtained by performing knowledge migration on M image task models. The M image task models are pre-trained models applied in different image tasks, and one image task model corresponds to one image migration model. For an image pair consisting of a first image and a second image, the first recognition feature corresponding to the first image and the second recognition feature corresponding to the second image can be obtained by the i-th image migration model in the M image migration models; at this time, the similarity between the first recognition feature and the second recognition feature can be used as the feature similarity associated with the i-th image migration model. The feature similarities associated with the M image migration models are added to obtain a feature similarity cumulative value; the first image is updated according to the feature similarity cumulative value, and finally an image adversarial sample for attacking the image recognition model in the business application is obtained. At this time, the image adversarial sample can contain knowledge migrated from the M image task models, which can enhance the attack transferability of the image adversarial sample, and thus improve the applicability of the image adversarial sample; when the image adversarial sample is used to attack the image recognition model in the business application, the evaluation effect of the image recognition model can be improved.
[0189] See also Figure 8 , Figure 8 Schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 8 As shown, the computer device 1000 may be a terminal device, for example, Figure 1 The terminal device 10a in the corresponding embodiment may also be a server, for example, Figure 1 The server 10d in the corresponding embodiment will not be limited here. For ease of understanding, this application takes a computer device as an example of a terminal device. The computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or it may be a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1005 may optionally also be at least one storage device located away from the aforementioned processor 1001. As Figure 8 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a user interface module, and a device control application program.
[0190] Among them, the network interface 1004 in the computer device 1000 can also provide network communication functions, and the optional user interface 1003 can also include a display and a keyboard. In Figure 8 In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0191] Obtain M image migration models; the M image migration models are recognition models obtained after knowledge migration of M image task models. The M image task models are pre-trained models applied to different image tasks. One image task model corresponds to one image migration model, and M is a positive integer;
[0192] Obtain a first image and a second image, input the first image and the second image into the i-th image migration model among the M image migration models, and obtain a first recognition feature corresponding to the first image and a second recognition feature corresponding to the second image through the i-th image migration model; different objects are included in the first image and the second image, and i is a positive integer less than or equal to M;
[0193] According to the first recognition feature and the second recognition feature, obtain the feature similarity associated with the i-th image migration model, and add up the feature similarities associated with the M image migration models to obtain a cumulative feature similarity value;
[0194] Update the first image according to the cumulative feature similarity value to generate an image adversarial sample, and combine the image adversarial sample and the second image into a test sample pair; the test sample pair is used to evaluate the performance of the image recognition model in the business application.
[0195] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the image processing method in any one of the foregoing Figure 2 、 Figure 4 embodiments, and can also execute the description of the image processing device 1 in the corresponding embodiment of the foregoing Figure 7 , which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.
[0196] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer program executed by the foregoing image processing device 1 or image processing device 2, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the foregoing Figure 2 、 Figure 4The description of the image processing method in any of the embodiments will not be repeated here. In addition, the beneficial effects of using the same method will not be described again. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, the program instructions may be deployed to be executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. The multiple computer devices distributed at multiple locations and interconnected through a communication network may form a blockchain system.
[0197] In addition, it should be noted that: The embodiments of this application also provide a computer program product or a computer program. The computer program product or the computer program may include computer instructions, and the computer instructions may be stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions, so that the computer device executes the Figure 2 , Figure 4 description of the image processing method in any of the embodiments, so it will not be repeated here. In addition, the beneficial effects of using the same method will not be described again. For the technical details not disclosed in the embodiments of the computer program product or the computer program involved in this application, please refer to the description of the method embodiments of this application.
[0198] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of this application are used to distinguish different media contents, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or modules, but may optionally further include steps or modules not listed, or may optionally further include other step units inherent to these processes, methods, devices, products, or equipment.
[0199] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of function in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0200] The methods and related devices provided by the embodiments of this application are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of this application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.
[0201] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0202] The above-disclosed are only the preferred embodiments of the present application. Certainly, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. An image processing method, characterized in that, Including: Obtain M image transfer models; the M image transfer models are recognition models obtained by performing knowledge transfer on M image task models, where the M image task models are pre-trained models applied to different image tasks, and one image task model corresponds to one image transfer model, and M is a positive integer; Obtain a first image and a second image, input the first image and the second image into the i-th image transfer model among the M image transfer models, and obtain a first recognition feature corresponding to the first image and a second recognition feature corresponding to the second image through the i-th image transfer model; different objects are included in the first image and the second image, and i is a positive integer less than or equal to M; According to the first recognition feature and the second recognition feature, obtain the feature similarity associated with the i-th image transfer model, and add the feature similarities associated with the M image transfer models to obtain a cumulative feature similarity value; Update the first image according to the cumulative feature similarity value to generate an image adversarial sample, and combine the image adversarial sample and the second image into a test sample pair; the test sample pair is used to evaluate the performance of the image recognition model in the business application.
2. The method according to claim 1, wherein The obtaining of the M image transfer models includes: Obtain M image task models and M initial recognition models; one image task model corresponds to one initial recognition model; Obtain a sample image, input the sample image into the M image task models, and obtain first sample recognition features of the sample image in each image task model; Input the sample image into the M initial recognition models, and obtain second sample recognition features of the sample image in each initial recognition model; According to the first sample recognition features of the sample image in each image task model and the second sample recognition features of the sample image in each initial recognition model, correct the network parameters of the M initial recognition models, and determine the M initial recognition models including the corrected network parameters as M image transfer models.
3. The method according to claim 2, characterized in that, The obtaining of the sample image includes: Obtain an original image set for training the M initial recognition models, perform object detection on the original images in the original image set, and obtain object location information corresponding to the original images in the original image set; According to the object location information, perform segmentation processing on the original images in the original image set to obtain object region images, and adjust the sizes of the object region images to obtain object region images with a fixed size; Add the object region images with a fixed size to the sample data set, and determine any one of the object region images in the sample data set as the sample image.
4. The method according to claim 2, wherein The inputting of the sample image into the M initial recognition models to obtain the second sample recognition features of the sample image in each initial recognition model includes: Input the sample image into the i-th initial recognition model among the M initial recognition models; the i-th initial recognition model includes N convolutional components, and N is a positive integer; Obtain the input features of the j-th convolutional component among the N convolutional components; when j is 1, the input features of the j-th convolutional component are the sample image, and when j is a positive integer less than N; Perform a convolution operation on the input features of the j-th convolutional component according to one or more convolutional layers in the i-th convolutional component to obtain candidate convolutional features; Perform normalization processing on the candidate convolutional features according to the weight vector corresponding to the normalization layer in the j-th convolutional component to obtain normalized features; Combine the normalized features and the input features of the j-th convolutional component to obtain the output features of the j-th convolutional component, and use the output features of the j-th convolutional component as the input features of the (j + 1)-th convolutional component; the j-th convolutional component is connected to the (j + 1)-th convolutional component; Determine the output features of the N-th convolutional component as the second sample recognition features of the sample image in the i-th initial recognition model.
5. The method according to claim 2, wherein The method of correcting the network parameters of the M initial recognition models according to the first sample recognition features of the sample image in each image task model and the second sample recognition features of the sample image in each initial recognition model, and determining the M initial recognition models including the corrected network parameters as M image transfer models includes: Classify the second sample recognition features of the sample image in the i-th initial recognition model to obtain the sample recognition result corresponding to the i-th initial recognition model; Determine the error between the sample recognition result corresponding to the i-th initial recognition model and the label information corresponding to the sample image as the classification loss corresponding to the i-th initial recognition model; Determine the sample similarity distance between the i-th initial recognition model and the M image task models according to the second sample recognition features of the sample image in the i-th initial recognition model and the first sample recognition features of the sample image in each image task model; Determine the joint loss corresponding to the i-th initial recognition model according to the classification loss corresponding to the i-th initial recognition model and the sample similarity distance between the i-th initial recognition model and the M image task models; Iteratively train the network parameters of the i-th initial recognition model according to the joint loss corresponding to the i-th initial recognition model until the joint loss corresponding to the i-th initial recognition model meets the training end condition, then stop training, and determine the i-th initial recognition model at the end of training as the i-th image transfer model.
6. The method according to claim 4, characterized in that, The method of determining the joint loss corresponding to the i-th initial recognition model according to the classification loss corresponding to the i-th initial recognition model and the sample similarity distance between the i-th initial recognition model and the M image task models includes: Obtain the task weights corresponding to the M image task models, and based on the task weights corresponding to the M image task models, perform a weighted sum of the sample similarity distances between the i-th initial recognition model and the M image task models to obtain the distillation loss corresponding to the i-th initial recognition model; Combine the classification loss corresponding to the i-th initial recognition model and the distillation loss corresponding to the i-th initial recognition model into the joint loss corresponding to the i-th initial recognition model.
7. The method according to claim 1, characterized in that, The updating the first image according to the feature similarity cumulative value to generate an image adversarial sample includes: Determine the adversarial noise corresponding to the first image according to the feature similarity cumulative value, and add the first image and the adversarial noise to obtain a candidate adversarial sample; Obtain the candidate recognition features of the candidate adversarial sample in the i-th image transfer model, and based on the candidate recognition features and the second recognition features corresponding to the second image, obtain the similarity distance between the candidate adversarial sample and the second image in the i-th image transfer model; Accumulate the similarity distances between the candidate adversarial sample and the second image in the M image transfer models to obtain a joint adversarial loss, and perform a minimization optimization process on the joint adversarial loss to obtain the image adversarial sample corresponding to the first image.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Input the image adversarial sample and the second image in the test sample pair into the image recognition model in the business application, and output, through the image recognition model, the adversarial recognition features corresponding to the image adversarial sample and the object recognition features corresponding to the second image; Obtain the image similarity between the adversarial recognition features and the object recognition features, and determine the test result of the image adversarial sample on the image recognition model according to the image similarity.
9. The method according to claim 8, wherein The determining the test result of the image adversarial sample on the image recognition model according to the image similarity includes: If the image similarity is greater than the recognition threshold, determine that the test result of the image adversarial sample on the image recognition model is an attack success; If the image similarity is less than or equal to the recognition threshold, determine that the test result of the image adversarial sample on the image recognition model is an attack failure.
10. The method according to claim 9, wherein The number of the test sample pairs is T, and T is an integer greater than 1; The method further includes: Obtain the test results of the image adversarial samples included in the T test sample pairs on the image recognition model, and count the attack success rate corresponding to the image recognition model among the test results corresponding to the T test sample pairs; Determine the model evaluation result corresponding to the image recognition model according to the attack success rate.
11. An image processing apparatus, characterized in that, including: A transfer model acquisition module, configured to acquire M image transfer models; the M image transfer models are recognition models obtained by performing knowledge transfer on M image task models, the M image task models are pre-trained models applied to different image tasks, one image task model corresponds to one image transfer model, and M is a positive integer; An image recognition module, configured to obtain a first image and a second image, input the first image and the second image into the i-th image migration model among the M image migration models, and obtain a first recognition feature corresponding to the first image and a second recognition feature corresponding to the second image through the i-th image migration model; different objects are included in the first image and the second image, and i is a positive integer less than or equal to M; A similarity acquisition module, configured to obtain the feature similarity associated with the i-th image migration model according to the first recognition feature and the second recognition feature, and add the feature similarities associated with the M image migration models to obtain a cumulative feature similarity value; An adversarial sample generation module, configured to update the first image according to the cumulative feature similarity value to generate an image adversarial sample, and combine the image adversarial sample and the second image into a test sample pair; the test sample pair is used to evaluate the performance of the image recognition model in a business application.
12. A computer device, characterized in that, Comprising a memory and a processor; The memory is connected to the processor, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1 to 10.
14. A computer program product, characterized in that, Comprising computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.