Image classification model training method and device, equipment and storage medium
By constructing an equivalent model of the ISP process and generating adversarial RAW-RGB data, the image classification model is trained. This solves the problem of ISP attacks ignoring RGB images in existing technologies and improves the recognition accuracy and robustness of the model.
Patent Information
- Application Number
- CN202510937217.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-26
AI Technical Summary
Existing image scaling attack methods only modify RGB images and ignore the important ISP process, resulting in image classification models being unable to accurately identify images when attacked by ISP, and existing training methods are unable to effectively improve the accuracy of the model.
By constructing an ISP process equivalent model based on the encoder-decoder architecture and training with RAW-RGB data pairs, adversarial RAW data is generated and converted into adversarial RGB data for training the image classification model to simulate the ISP attack process.
The recognition accuracy of the image classification model is improved, the model's robustness to ISP attacks is enhanced, the richness of training samples is increased, and the classification accuracy of the model is improved.
Smart Images

Figure CN120707991A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and artificial intelligence security technology, and in particular to an image classification model training method, device, equipment and storage medium. Background Art
[0002] With the widespread adoption of deep learning technology, an increasing number of visual applications rely on the output of deep neural networks (DNNs) to make decisions. However, numerous studies have shown that DNNs are susceptible to various attacks, leading to erroneous output. This poses a potential threat to the deployment of DNNs in security-sensitive areas. Real-world visual applications often include a preprocessing step called image scaling to reduce images to a fixed size. This process is not secure and is susceptible to image scaling attacks, where subtle noise is added to the original image to create an attack image. This scaled attack image is then fed into an image classification model, causing the model to misclassify the image. Therefore, when training image classification models, attack images are often used to train the model, forcing it to correctly classify the image.
[0003] Currently, the specific process of image scaling attack is as follows: Figure 1 As shown, image scaling attack methods only modify RGB (Red, Green, Blue) images, ignoring the crucial Image Signal Processing (ISP) process. Consequently, attack images generated by existing image attack methods cannot fully train image classification models. Therefore, the challenge is to generate more accurate attack images to improve the accuracy of image classification model training. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an image classification model training method, apparatus, device, and storage medium that can improve the accuracy of image classification model training. The specific solution is as follows:
[0005] In a first aspect, the present application discloses an image classification model training method, comprising:
[0006] Obtain the original RAW data input and the original RGB image output during the ISP process of the visual application to obtain the corresponding original RAW-RGB data pairs;
[0007] Constructing an initial ISP process equivalent model based on a preset encoder-decoder architecture, and training the initial ISP process equivalent model using the original RAW-RGB data to obtain a target ISP process equivalent model;
[0008] The original RAW data is processed using a preset optimization target to obtain adversarial RAW data, and adversarial RGB data corresponding to the adversarial RAW data is generated using the target ISP process equivalent model, so that the adversarial RGB data can be used to perform model training on the image classification model to be trained.
[0009] Optionally, acquiring the original RAW data input and the original RGB image output during the visual application ISP process to obtain corresponding original RAW-RGB data pairs includes:
[0010] Original RAW data is read from a preset data set, and the original RAW data is input into a visual application OpenISP process to obtain an output original RGB image, and the original RAW data and the original RGB image are determined as an original RAW-RGB data pair.
[0011] Optionally, constructing an equivalent model of the initial ISP process based on a preset encoder-decoder architecture includes:
[0012] A convolutional neural network model is constructed based on a preset encoder-decoder architecture, and a residual block is added to the convolutional neural network model to obtain an equivalent model of the initial ISP process.
[0013] Optionally, the training the initial ISP process equivalent model using the original RAW-RGB data to obtain a target ISP process equivalent model includes:
[0014] The original RAW-RGB data pair is input into the initial ISP process equivalent model, and the initial ISP process equivalent model is trained using a preset supervised learning method combined with a target loss function to obtain a target ISP process equivalent model.
[0015] Optionally, before the initial ISP process equivalent model is trained using a preset supervised learning method in combination with a target loss function to obtain a target ISP process equivalent model, the method further includes:
[0016] Constructing a content loss function based on the image distance difference between the equivalent model of the initial ISP process and the image distance difference of the visual application ISP process;
[0017] Constructing a structural similarity loss function based on the image color deviation of the initial ISP process equivalent model and the visual application ISP process;
[0018] The target loss function is generated according to the content loss function and the structural similarity loss function.
[0019] Optionally, processing the original RAW data using a preset optimization target to obtain adversarial RAW data, and generating adversarial RGB data corresponding to the adversarial RAW data using the target ISP process equivalent model, includes:
[0020] Determining a preset optimization target corresponding to the original RAW data, and processing the preset optimization target using a gradient optimization method to obtain adversarial RAW data;
[0021] The adversarial RAW data is input into the target ISP process equivalent model to generate adversarial RGB data corresponding to the adversarial RAW data.
[0022] Optionally, after processing the original RAW data using a preset optimization target to obtain countermeasure RAW data, the method further includes:
[0023] Inputting the adversarial RAW data into the visual application ISP process to obtain a target attack image;
[0024] Scaling the target attack image to obtain a corresponding image to be detected, and inputting the image to be detected into an image classification detection model to obtain an image classification result;
[0025] If the image classification result is inconsistent with the actual image classification of the adversarial RAW data, the adversarial RGB data is determined as an interference image, so as to use the interference image to perform model training on the image classification model to be trained.
[0026] In a second aspect, the present application discloses an image classification model training device, comprising:
[0027] The data pair acquisition module is used to acquire the original RAW data input and the original RGB image output during the ISP process of the visual application to obtain the corresponding original RAW-RGB data pairs;
[0028] An equivalent model training module is used to construct an initial ISP process equivalent model based on a preset encoder-decoder architecture, and train the initial ISP process equivalent model using the original RAW-RGB data to obtain a target ISP process equivalent model;
[0029] A model training module is used to process the original RAW data using a preset optimization target to obtain adversarial RAW data, and to generate adversarial RGB data corresponding to the adversarial RAW data using the target ISP process equivalent model, so as to use the adversarial RGB data to perform model training on the image classification model to be trained.
[0030] In a third aspect, the present application discloses an electronic device, comprising:
[0031] Memory, used to store computer programs;
[0032] A processor is used to execute the computer program to implement the aforementioned image classification model training method.
[0033] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program, which implements the aforementioned image classification model training method when executed by a processor.
[0034] As can be seen, in this application, the original RAW data input and the original RGB image output of the vision application ISP process are obtained to obtain corresponding original RAW-RGB data pairs; an initial ISP process equivalent model is constructed based on a preset encoder-decoder architecture, and the initial ISP process equivalent model is trained using the original RAW-RGB data pair to obtain a target ISP process equivalent model; the original RAW data is processed using a preset optimization target to obtain adversarial RAW data, and adversarial RGB data corresponding to the adversarial RAW data is generated using the target ISP process equivalent model, so that the adversarial RGB data can be used to train the image classification model to be trained. In other words, by constructing an equivalent model and training it using the RAW-RGB data pair to approximate the ISP process, the RAW data is then modified based on certain optimization information so that the attack image generated by the ISP process can be transformed into another image after scaling, that is, an image (the attack image) that can mislead the image classification model. The image classification model is then trained using these generated attack images, which can improve the recognition accuracy of the image classification model training and thus the accuracy of the model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0036] Figure 1 This is a schematic diagram of an image scaling attack disclosed in this application;
[0037] Figure 2 This is a flow chart of an image classification model training method disclosed in this application;
[0038] Figure 3 A schematic diagram of a specific equivalent model structure disclosed in this application;
[0039] Figure 4 This is a structural diagram of an image classification model training device disclosed in this application;
[0040] Figure 5 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0042] Real-world visual applications all have a built-in preprocessing process for image scaling operations to scale images to a fixed size. This process is not secure and is vulnerable to image scaling attacks, which involves adding subtle noise to the original image to generate an attack image. The attack image is scaled and then input into the image classification model, causing the model to output an incorrect classification result. Therefore, when training an image classification model, such attack images are usually added to the image samples to enrich the samples and improve the accuracy of image classification model training. However, the image scaling attack method only modifies the RGB image and ignores the important ISP process. Once the image is attacked during the ISP process, the image classification model will not be able to accurately recognize it. Therefore, this application will specifically introduce a method for training an image classification model. By generating images that are attacked during the ISP process and then performing model training, an image classification model with a higher classification recognition rate is obtained.
[0043] See also Figure 2 As shown, the embodiment of the present application discloses an image classification model training method, comprising:
[0044] Step S11: Acquire the original RAW data input and the original RGB image output during the visual application ISP process to obtain corresponding original RAW-RGB data pairs.
[0045] In this embodiment, the original RAW data input and the original RGB image output in the visual application ISP process are obtained to obtain the corresponding original RAW-RGB data pair, including: reading the original RAW data from a preset data set, and inputting the original RAW data into the visual application OpenISP process to obtain the output original RGB image, and determining the original RAW data and the original RGB image as an original RAW-RGB data pair. That is, in order to enable the constructed equivalent model to approximate the ISP process, it is necessary to obtain the RAW-RGB data pair consisting of the RAW (Raw Data) data input to the ISP process and the RGB image output. Specifically, the ISP process is set to the OpenISP process, and the RAW data in the Zurich RAW to RGB dataset is input into the OpenISP process. OpenISP outputs the corresponding RGB image to obtain the RAW-RGB data pair for training the equivalent model.
[0046] Step S12: construct an initial ISP process equivalent model based on a preset encoder-decoder architecture, and use the original RAW-RGB data to train the initial ISP process equivalent model to obtain a target ISP process equivalent model.
[0047] In this embodiment, the initial ISP process equivalent model is constructed based on the preset encoder-decoder architecture, including: constructing a convolutional neural network model based on the preset encoder-decoder architecture, and adding a residual block to the convolutional neural network model to obtain an initial ISP process equivalent model. The ISP process built into the visual device usually includes multiple complex modules, and there is a problem of difficulty in obtaining gradient information. By constructing an equivalent model based on a convolutional neural network to approximate the ISP process of the visual device, the problem of difficulty in obtaining gradient information can be solved. The structure of the equivalent model is as follows Figure 3 As shown in Figure 2, the equivalent model uses an encoder-decoder architecture, where the encoder extracts features from the RAW data and the decoder reconstructs the corresponding RGB image. To ensure efficient feature extraction and image reconstruction, the equivalent model uses a residual block, a skip-connection structure consisting of multiple convolutional layers.
[0048] In this embodiment, the training of the initial ISP process equivalent model using the original RAW-RGB data pair to obtain a target ISP process equivalent model includes: inputting the original RAW-RGB data pair into the initial ISP process equivalent model, and training the initial ISP process equivalent model using a preset supervised learning method in combination with a target loss function to obtain a target ISP process equivalent model. Using the RAW-RGB data pair, the equivalent model is trained using a supervised training method in combination with the loss function of the equivalent model, so that the equivalent model can approximate the RAW-RGB data pair.
[0049] In this embodiment, before the preset supervised learning method is used to train the initial ISP process equivalent model in combination with the target loss function to obtain the target ISP process equivalent model, the method further includes: constructing a content loss function based on the image distance difference between the initial ISP process equivalent model and the image in the visual application ISP process; constructing a structural similarity loss function based on the image color deviation between the initial ISP process equivalent model and the image in the visual application ISP process; and generating the target loss function based on the content loss function and the structural similarity loss function. The target loss function can be expressed as:
[0050] ;
[0051] in, is the content loss function, is the structural similarity loss function, It is an adjustable parameter.
[0052] During the training of the ISP process equivalent model, the content loss applied to the RAW-RGB data used for training is As shown below:
[0053] ;
[0054] In order to avoid color deviation in the reconstructed RGB image, structural similarity (SSIM) is used to enhance its dynamic range. The SSIM calculation process of image a and image b is as follows:
[0055] ;
[0056] in, 、 Represent the mean and standard deviation of image a respectively, 、 represent the mean and standard deviation of image b respectively, represents the covariance of image a and image b, and Indicates an adjustable parameter.
[0057] RGB image converted by ISP process With the equivalent model Reconstructed RGB image SSIM loss As shown below:
[0058] ;
[0059] in, and Larger values indicate equivalent models Reconstructed RGB image The better the quality.
[0060] Step S13: Process the original RAW data using a preset optimization target to obtain adversarial RAW data, and use the target ISP process equivalent model to generate adversarial RGB data corresponding to the adversarial RAW data, so as to use the adversarial RGB data to perform model training on the image classification model to be trained.
[0061] In this embodiment, the original RAW data is processed using a preset optimization target to obtain adversarial RAW data, and adversarial RGB data corresponding to the adversarial RAW data is generated using the target ISP process equivalent model, including: determining the preset optimization target corresponding to the original RAW data, and processing the preset optimization target using a gradient optimization method to obtain adversarial RAW data; inputting the adversarial RAW data into the target ISP process equivalent model to generate adversarial RGB data corresponding to the adversarial RAW data. Specifically, for the ISP process equivalent model, the adversarial RAW data is generated using a gradient descent method based on the optimization target for generating the adversarial RAW data. If the generated adversarial RAW data is converted into an RGB image through an equivalent model. Among them, the optimization target of the adversarial RAW data can be expressed as:
[0062] ;
[0063] in, It is the original RAW data. yes Converted through the ISP process The original image after It is against RAW data. yes After the equivalent model The converted attack image, O is the output image of the attack image A after the scaling operation, T is the target image, and Respectively represent the classification results of the image classification model for the target image T and the output image O. c is an adjustable parameter, 、 As shown in formula (6):
[0064] ;
[0065] In formula (6), m, n, c are attack images The data dimension, 、 、 is the data dimension of the output image O. Adapting gradient optimization to minimize the optimization target can generate adversarial RAW data , The converted attack image is similar to the original image S, After scaling, the output image O is obtained. The output image O is similar to the target image T. After the output image O is input into the image classification model, the classification result is the same as the target image T. After generating the corresponding adversarial RGB data, the adversarial RGB data is used to train the image classification model to be trained. This can increase the richness of the samples in the model training and thus improve the classification accuracy of the trained model.
[0066] In addition, after the original RAW data is processed using a preset optimization target to obtain adversarial RAW data, the method further includes: inputting the adversarial RAW data into the visual application ISP process to obtain a target attack image; scaling the target attack image to obtain a corresponding image to be detected, and inputting the image to be detected into an image classification detection model to obtain an image classification result; if the image classification result is inconsistent with the actual image classification of the adversarial RAW data, the adversarial RGB data is determined to be an interference image, so that the interference image can be used to train the image classification model to be trained. The RGB image is input into the image classification detection model, and the model outputs an incorrect classification result. The generated adversarial RAW data successfully attacks the ISP process equivalent model. After generating the adversarial RAW data for the ISP process equivalent model, it is input into the ISP process of the visual device. The adversarial RAW data is converted into an attack image through the ISP process, and the attack image is scaled and input into the image classification model. If the model outputs an incorrect classification result, the adversarial RAW data successfully attacks the ISP process of the visual device. In this way, by detecting the generated target attack image, we can ensure that the generated adversarial RGB data can include the ISP attack process, increase the richness of samples in model training, and thus improve the accuracy of model classification after training.
[0067] It should be noted that the dataset used to train the ISP process equivalent model is the Zurich RAW to RGB dataset. The target images were randomly selected from the Animals10 dataset. Each target image had its long side resized to 224 and its short side scaled proportionally. After the adversarial RAW data was transformed and scaled through the ISP process to obtain the output image, the classification model was used to classify the output image. The image classification test model used in the experiment was the VGG-19 model trained on the Animals10 data, using the Nearest and Bilinear scaling operations.
[0068] As can be seen, in this embodiment, the original RAW data input and the original RGB image output during the vision application ISP process are obtained to obtain corresponding original RAW-RGB data pairs; an initial ISP process equivalent model is constructed based on a preset encoder-decoder architecture, and the initial ISP process equivalent model is trained using the original RAW-RGB data pairs to obtain a target ISP process equivalent model; the original RAW data is processed using a preset optimization target to obtain adversarial RAW data, and adversarial RGB data corresponding to the adversarial RAW data is generated using the target ISP process equivalent model, so that the adversarial RGB data can be used to train the image classification model to be trained. Specifically, by constructing an equivalent model and training it using the RAW-RGB data pairs to approximate the ISP process, the RAW data is then modified based on certain optimization information so that the attack image generated by the ISP process can be transformed into another image after scaling, that is, an image that can mislead the image classification model (the attack image). Furthermore, by using these generated attack images to train the image classification model, the recognition accuracy of the image classification model training can be improved, thereby enhancing the accuracy of the model training.
[0069] refer to Figure 4 The present application also discloses an image classification model training device, including:
[0070] The data pair acquisition module 11 is used to acquire the original RAW data input and the original RGB image output during the visual application ISP process to obtain the corresponding original RAW-RGB data pairs;
[0071] an equivalent model training module 12, configured to construct an initial ISP process equivalent model based on a preset encoder-decoder architecture, and train the initial ISP process equivalent model using the original RAW-RGB data to obtain a target ISP process equivalent model;
[0072] The model training module 13 is used to process the original RAW data using a preset optimization target to obtain adversarial RAW data, and use the target ISP process equivalent model to generate adversarial RGB data corresponding to the adversarial RAW data, so as to use the adversarial RGB data to perform model training on the image classification model to be trained.
[0073] Derivation of the visible + weight 1 effect
[0074] In some specific embodiments, the data pair acquisition module 11 may specifically include:
[0075] A data pair determination unit is configured to read original RAW data from a preset data set, input the original RAW data into the visual application OpenISP process to obtain an output original RGB image, and determine the original RAW data and the original RGB image as an original RAW-RGB data pair.
[0076] In some specific embodiments, the equivalent model training module 12 may specifically include:
[0077] The initial model determination unit is used to construct a convolutional neural network model based on a preset encoder-decoder architecture, and add a residual block to the convolutional neural network model to obtain an equivalent model of the initial ISP process.
[0078] In some specific embodiments, the equivalent model training module 12 may specifically include:
[0079] An equivalent model training unit is used to input the original RAW-RGB data pair into the initial ISP process equivalent model, and train the initial ISP process equivalent model using a preset supervised learning method combined with a target loss function to obtain a target ISP process equivalent model.
[0080] In some specific embodiments, the image classification model training device may further include:
[0081] A content loss function construction module, configured to construct a content loss function based on the image distance difference between the equivalent model of the initial ISP process and the image distance difference between the ISP process and the vision application process;
[0082] A structural similarity loss function construction module, configured to construct a structural similarity loss function based on the image color deviation between the initial ISP process equivalent model and the visual application ISP process;
[0083] The target loss function determination module is used to generate the target loss function according to the content loss function and the structural similarity loss function.
[0084] In some specific embodiments, the model training module 13 may specifically include:
[0085] an adversarial RAW data optimization unit, configured to determine a preset optimization target corresponding to the original RAW data, and process the preset optimization target using a gradient optimization method to obtain adversarial RAW data;
[0086] The adversarial RGB data generating unit is used to input the adversarial RAW data into the target ISP process equivalent model to generate adversarial RGB data corresponding to the adversarial RAW data.
[0087] In some specific embodiments, the image classification model training device may further include:
[0088] a target attack image determination module, configured to input the adversarial RAW data into the visual application ISP process to obtain a target attack image;
[0089] An image classification module is used to scale the target attack image to obtain a corresponding image to be detected, and input the image to be detected into an image classification detection model to obtain an image classification result;
[0090] A model training module is used to determine the adversarial RGB data as an interference image if the image classification result is inconsistent with the actual image classification of the adversarial RAW data, so as to use the interference image to perform model training on the image classification model to be trained.
[0091] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0092] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps of the image classification model training method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0093] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0094] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0095] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of implementing the image classification model training method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 may further include a computer program capable of implementing other specific tasks.
[0096] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned image classification model training method. The specific steps of this method can be referred to the corresponding content disclosed in the aforementioned embodiments and will not be repeated here.
[0097] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0098] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0099] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0100] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0101] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for training an image classification model, characterized in that: include: Obtain the original RAW data input and the original RGB image output during the ISP process of the visual application to obtain the corresponding original RAW-RGB data pairs; Constructing an initial ISP process equivalent model based on a preset encoder-decoder architecture, and training the initial ISP process equivalent model using the original RAW-RGB data to obtain a target ISP process equivalent model; The original RAW data is processed using a preset optimization target to obtain adversarial RAW data, and adversarial RGB data corresponding to the adversarial RAW data is generated using the target ISP process equivalent model, so that the adversarial RGB data can be used to perform model training on the image classification model to be trained.
2. The image classification model training method according to claim 1, characterized in that: The step of obtaining the original RAW data input and the original RGB image output during the visual application ISP process to obtain corresponding original RAW-RGB data pairs includes: Original RAW data is read from a preset data set, and the original RAW data is input into a visual application OpenISP process to obtain an output original RGB image, and the original RAW data and the original RGB image are determined as an original RAW-RGB data pair.
3. The image classification model training method according to claim 1, characterized in that: The initial ISP process equivalent model is constructed based on the preset encoder-decoder architecture, including: A convolutional neural network model is constructed based on a preset encoder-decoder architecture, and a residual block is added to the convolutional neural network model to obtain an equivalent model of the initial ISP process.
4. The image classification model training method according to claim 1, characterized in that: The training of the initial ISP process equivalent model using the original RAW-RGB data to obtain a target ISP process equivalent model includes: The original RAW-RGB data pair is input into the initial ISP process equivalent model, and the initial ISP process equivalent model is trained using a preset supervised learning method combined with a target loss function to obtain a target ISP process equivalent model.
5. The image classification model training method according to claim 4, characterized in that: Before the initial ISP process equivalent model is trained by using a preset supervised learning method in combination with a target loss function to obtain a target ISP process equivalent model, the method further includes: Constructing a content loss function based on the image distance difference between the equivalent model of the initial ISP process and the image distance difference of the visual application ISP process; Constructing a structural similarity loss function based on the image color deviation of the initial ISP process equivalent model and the visual application ISP process; The target loss function is generated according to the content loss function and the structural similarity loss function.
6. The image classification model training method according to claim 1, characterized in that: The processing of the original RAW data using a preset optimization target to obtain adversarial RAW data, and generating adversarial RGB data corresponding to the adversarial RAW data using the target ISP process equivalent model, includes: Determining a preset optimization target corresponding to the original RAW data, and processing the preset optimization target using a gradient optimization method to obtain adversarial RAW data; The adversarial RAW data is input into the target ISP process equivalent model to generate adversarial RGB data corresponding to the adversarial RAW data.
7. The image classification model training method according to any one of claims 1 to 6, characterized in that: After the original RAW data is processed using the preset optimization target to obtain the countermeasure RAW data, the method further includes: Inputting the adversarial RAW data into the visual application ISP process to obtain a target attack image; Scaling the target attack image to obtain a corresponding image to be detected, and inputting the image to be detected into an image classification detection model to obtain an image classification result; If the image classification result is inconsistent with the actual image classification of the adversarial RAW data, the adversarial RGB data is determined as an interference image, so as to use the interference image to perform model training on the image classification model to be trained.
8. An image classification model training device, characterized in that: include: The data pair acquisition module is used to acquire the original RAW data input and the original RGB image output during the ISP process of the visual application to obtain the corresponding original RAW-RGB data pairs; An equivalent model training module is used to construct an initial ISP process equivalent model based on a preset encoder-decoder architecture, and train the initial ISP process equivalent model using the original RAW-RGB data to obtain a target ISP process equivalent model; A model training module is used to process the original RAW data using a preset optimization target to obtain adversarial RAW data, and to generate adversarial RGB data corresponding to the adversarial RAW data using the target ISP process equivalent model, so as to use the adversarial RGB data to perform model training on the image classification model to be trained.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the image classification model training method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program, which, when executed by a processor, implements the image classification model training method according to any one of claims 1 to 7.