Artificial Intelligence-Based Image Augmentation Processing Method, Device, Equipment, and Storage Medium
By encoding, noise modulation, decoding and superimposing processing of the target image, the noise perturbation image is generated as an augmented image, which solves the problem that image augmentation in the prior art is difficult to improve the generalization ability of image classification model, and achieves a more efficient and better quality image augmentation effect.
Patent Information
- Application Number
- CN202011074076.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-10-09
AI Technical Summary
The prior art uses conventional operations in image augmentation, making it difficult to effectively improve the generalization ability of image classification models.
By encoding, noise modulation, decoding and superimposing processing of the target image, noise perturbation images are generated as augmented images, and the image augmentation process is optimized using artificial intelligence technology.
It improves the performance and quality of image augmentation, expands the sample number, and enhances the generalization ability of image classification model.
Smart Images

Figure CN112132106B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the graphic image processing technology of artificial intelligence, and particularly relates to an image augmentation processing method, device, electronic device and computer-readable storage medium based on artificial intelligence. Background Art
[0002] Artificial Intelligence (AI) is a comprehensive technology in computer science. By studying the design principles and implementation methods of various intelligent machines, machines are enabled to have the functions of perception, reasoning and decision-making. The graphic processing technology based on artificial intelligence has been applied in many fields and plays an increasingly important role.
[0003] In the application of image classification models, it is usually necessary to obtain diverse image samples through image augmentation, and then train the image classification model to improve the generalization ability of the image classification model. Taking the face recognition model as an example, if the face recognition model can recognize faces from various face images including noise, it indicates that it has good generalization ability.
[0004] In related technologies, conventional image operations are often used in image augmentation, such as horizontal flipping, translation, rotation, etc. However, training the image classification model with the augmented images obtained by these methods has very limited effect on improving the generalization ability of the image classification model.
[0005] Therefore, there is no effective solution in related technologies for how to perform image augmentation to improve the generalization ability of the image classification model. Summary of the Invention
[0006] The embodiments of this application provide an image augmentation processing method, device, electronic device and computer-readable storage medium based on artificial intelligence, which can effectively improve the performance and quality of image augmentation.
[0007] The technical solution of the embodiments of this application is implemented as follows:
[0008] The embodiments of this application provide an image augmentation processing method based on artificial intelligence, including:
[0009] Performing encoding processing on a target image to obtain the image features of the target image;
[0010] Querying a feature library based on the first label type of the target image to obtain the first feature parameter of the normal distribution followed by the first label type;
[0011] Performing noise modulation processing on the image features based on the first feature parameter to obtain the first adversarial noise features;
[0012] Decode the first adversarial noise feature to obtain a first adversarial noise image;
[0013] Superimpose the target image and the first adversarial noise image to obtain a noise perturbation image as the augmented image of the target image.
[0014] An embodiment of the present application provides an image augmentation processing device based on artificial intelligence, including:
[0015] A first encoder for encoding a target image to obtain an image feature of the target image;
[0016] A modulator for querying a feature library based on a first label type of the target image to obtain a first feature parameter of a normal distribution followed by the first label type;
[0017] Perform noise modulation processing on the image feature based on the first feature parameter to obtain a first adversarial noise feature;
[0018] A decoder for decoding the first adversarial noise feature to obtain a first adversarial noise image;
[0019] A superimposing module for superimposing the target image and the first adversarial noise image to obtain a noise perturbation image as the augmented image of the target image.
[0020] In the above solution, the feature library stores a mapping relationship between different label types and different feature parameters;
[0021] The modulator is further configured to query the mapping relationship stored in the feature library based on the first label type of the target image to obtain a first feature parameter of a normal distribution corresponding to the first label tag.
[0022] In the above solution, the feature parameters of the normal distribution followed by the first label type include a first mean vector and a first variance vector;
[0023] Among them, the first mean vector is used to represent the mean of the image features of the first label type, and the first variance vector is used to represent the jitter degree of the image features of the first label type;
[0024] The modulator is further configured to determine a first difference between the first mean vector and the image feature of the target image;
[0025] Determine the first adversarial noise feature by a first ratio between the first difference and the first variance vector.
[0026] In the above solution, the encoding process is implemented by the first encoder in the codec model, the noise modulation process is implemented by the modulator in the codec model, and the decoding process is implemented by the decoder in the codec model;
[0027] In the above solution, an image augmentation processing device based on artificial intelligence provided by an embodiment of the present application further includes:
[0028] A first training module, configured to iteratively perform the following training operations:
[0029] Jointly train the codec model, the first classification model, and the feature library based on the classification probability distribution of the image sample by the first classification model and the classification probability distribution of the noise-perturbed image of the image sample;
[0030] Train the first classification model based on the classification probability distribution of the image sample by the first classification model and the classification probability distribution of the noise-perturbed image of the image sample.
[0031] In the above solution, the first training module is further configured to generate a noise-perturbed image of the image sample through the codec model;
[0032] Generate the classification probability distribution of the noise-perturbed image of the image sample and the classification probability distribution of the image sample through the first classification model;
[0033] Construct a first loss function based on the difference between the classification probability distribution of the noise-perturbed image and the classification probability distribution of the image sample, and update the model parameters of the codec model, the model parameters of the first classification model, and the feature parameters of the feature library by maximizing the first loss function.
[0034] In the above solution, the classification probability distribution of the noise-perturbed image of the image sample includes the probabilities that the noise-perturbed image of the image sample belongs to the original image type and the noise image type respectively;
[0035] The classification probability distribution of the image sample includes the probabilities that the image sample belongs to the original image type and the noise image type respectively;
[0036] The first training module is further configured to determine the gradient values of the codec model, the first classification model, and the feature library when maximizing the first loss function;
[0037] Update the model parameters of the codec model based on the gradient value of the codec model;
[0038] Update the model parameters of the first classification model based on the gradient value of the first classification model;
[0039] Update the feature parameters of the normal distribution followed by the second label type in the feature library based on the gradient value of the feature library; wherein, the second label type is the pre-label type of the image sample.
[0040] In the above solution, the model parameters of the modulator include a modulation bias parameter and a modulation rate parameter;
[0041] The first training module is further configured to determine the gradient value of the modulator when maximizing the first loss function, and update the modulation bias parameter and the modulation rate parameter based on the gradient value of the modulator.
[0042] In the above solution, the first training module is further configured to perform downsampling processing on the image sample through the first encoder to obtain the image features of the image sample;
[0043] Query the feature library based on the second label type of the image sample to obtain the second feature parameters of the normal distribution followed by the second label type of the image sample; wherein, the second label type is the pre-label type of the image sample;
[0044] Perform noise modulation processing on the image features of the image sample through the modulator and based on the second feature parameters to obtain second adversarial noise features;
[0045] Perform upsampling processing on the second adversarial noise features through the decoder to obtain the noise perturbed image of the image sample.
[0046] In the above solution, the first classification model includes a second encoder, a third encoder, and a classifier;
[0047] The first training module is further configured to perform feature extraction processing on the noise perturbed image of the image sample through the second encoder and the third encoder to obtain the image features of the noise perturbed image of the image sample, and map the image features of the noise perturbed image of the image sample to the classification probability distribution of the noise perturbed image of the image sample through the classifier;
[0048] Perform feature extraction processing on the image sample through the second encoder and the third encoder to obtain the image features of the image sample, and map the image features of the image sample to the classification probability distribution of the image sample through the classifier.
[0049] In the above solution, the first encoder and the second encoder have the same structure and share the same model parameters.
[0050] In the above solution, the first training module is further configured to classify the noise-perturbed image of the image sample and the image sample through the first classification model, so as to obtain the probability that the noise-perturbed image of the image sample belongs to the original image type and the probability that the image sample belongs to the original image type;
[0051] Construct a second loss function according to the difference between the probability that the image sample belongs to the original image type obtained by the first classification model and the probability that the noise-perturbed image of the image sample belongs to the original image type;
[0052] Update the model parameters of the first classification model by minimizing the second loss function.
[0053] In the above solution, an image augmentation processing device based on artificial intelligence provided by an embodiment of the present application further includes:
[0054] A second training module, configured to establish a training set based on a target image and a noise-perturbed image of the target image;
[0055] Train a second classification model based on the training set;
[0056] Wherein, the labeled data in the training set is labeled according to the classification task of the second classification model, and the classification task of the second classification model is different from that of the first classification model.
[0057] An embodiment of the present application provides an image augmentation processing device based on artificial intelligence, including:
[0058] A memory, configured to store executable instructions;
[0059] A processor, configured to implement the image augmentation processing method based on artificial intelligence provided by an embodiment of the present application when executing the executable instructions stored in the memory.
[0060] An embodiment of the present application provides a computer-readable storage medium, storing executable instructions, which are configured to implement the image augmentation processing method based on artificial intelligence provided by an embodiment of the present application when being executed by a processor.
[0061] The embodiments of the present application have the following beneficial effects:
[0062] Perform a series of processes of encoding, modulating, decoding, and superimposing on the target image to automatically generate augmented images, effectively improving the performance and quality of image augmentation; expand the number of samples through the generated augmented images to obtain diversified image samples, and thus train an image classification model to improve the generalization ability of the image classification model. Description of the Drawings
[0063] Figure 1 is an optional schematic structural diagram of an artificial intelligence-based image augmentation processing system 100 provided by an embodiment of the present application;
[0064] Figure 2 is a schematic structural diagram of a server 200 for artificial intelligence-based image augmentation processing provided by an embodiment of the present application;
[0065] Figure 3 is a schematic structural diagram of an artificial intelligence-based image augmentation processing device 255 provided by an embodiment of the present application;
[0066] Figure 4A is an optional schematic flowchart of an artificial intelligence-based image augmentation processing method provided by an embodiment of the present application;
[0067] Figure 4B is an optional schematic flowchart of an artificial intelligence-based image augmentation processing method provided by an embodiment of the present application;
[0068] Figure 4C is an optional schematic flowchart of an artificial intelligence-based image augmentation processing method provided by an embodiment of the present application;
[0069] Figure 5 is a schematic structural diagram of an encoding and decoding model provided by an embodiment of the present application;
[0070] Figure 6 is a schematic structural diagram of a first classification model provided by an embodiment of the present application;
[0071] Figure 7 is a schematic flowchart of a training method for an encoding and decoding model provided by an embodiment of the present application;
[0072] Figure 8 is a schematic structural diagram of an encoding and decoding model for training provided by an embodiment of the present application. Detailed implementation manners
[0073] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0074] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0075] In the following description, the terms "first", "second", and "third" only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first", "second", and "third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0077] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.
[0078] 1) Adversarial noise, which is noise that interferes with the image classification model from making a correct classification result for an image. For example, after adding adversarial noise to the image samples in the training set, the classification accuracy of the image classification model for the training set will decrease compared to before adding the adversarial noise. The goal of training an image classification model is to make the image classification model "immune" to adversarial noise.
[0079] 2) Generalization ability, which refers to the adaptability of machine learning algorithms to fresh samples. Briefly speaking, it is to add new sample data to the original sample data set and output a reasonable result through training. The purpose of learning is to learn the rules hidden behind the sample data. For data outside the sample data set with the same rules, the trained network can also give a suitable output, which is called generalization ability.
[0080] 3) Image augmentation, randomly changing training samples can reduce the model's dependence on certain attributes, thereby improving the model's generalization ability. For example, the image can be cropped in different ways so that the object of interest appears in different positions, thereby enabling the model to reduce its dependence on the position where the object appears.
[0081] 4) The first classification model is a classification model used to assist in training the codec model, and its classification result is a probability distribution formed by the probabilities that the image to be classified belongs to the image and the probability of belonging to the noise image respectively.
[0082] 5) The second classification model is used to complete a classification task different from the first classification model.
[0083] 6) The original image is a "pure" image without added noise, such as an image obtained through image acquisition.
[0084] 7) The target image, which is the original image that needs to have noise added for image augmentation, is used to combine with the original image to form a training set for training the second classification model.
[0085] 8) The noise image is an image containing noise formed by means of image augmentation on the original image.
[0086] Generally, an image classification model needs to obtain diverse image samples through image augmentation, and then train the image classification model to improve its generalization ability, so that the image classification model can achieve better accuracy and recall rates for the training set and online data. In related technologies, in terms of image augmentation, conventional image operations are often adopted, such as horizontal flipping, translation, rotation, etc.; or the method of using a generative adversarial network to learn features is used to augment images specifically, for example, obtaining adversarial noise in the image space through gradient feedback, and then superimposing the adversarial noise on the original image to weaken the features of the original image, so that the image classification model can learn other general features to improve the generalization ability of the image classification model.
[0087] In the embodiments of the present application, it is found that the above methods in related technologies will have the following technical problems in the actual application process: The method of manually defined picture operations has limited effect on improving the generalization ability of the image classification model; For a single picture, it is necessary to go through multiple gradient feedbacks to obtain relatively good adversarial noise, and the efficiency is low.
[0088] In view of the above technical problems, the embodiments of the present application provide an artificial intelligence-based image augmentation processing method, device, electronic device, and computer-readable storage medium, which can effectively improve the performance and quality of image augmentation. The following describes an exemplary application of the artificial intelligence-based image augmentation processing electronic device provided by the embodiments of the present application. The artificial intelligence-based image augmentation processing electronic device provided by the embodiments of the present application can be implemented as a server, which performs a series of processes such as encoding, modulating, decoding, and superimposing on the target image, and automatically generates an augmented image corresponding to the marked type of the target image; It can also be implemented as various types of user terminals, which automatically generate an augmented image corresponding to the marked type of the target image according to the target image input by the user. The following will describe the exemplary application when the electronic device is implemented as a server.
[0089] See Figure 1 , Figure 1 FIG. is an optional schematic architecture diagram of an artificial intelligence-based image augmentation processing system 100 provided by the embodiments of the present application. The artificial intelligence-based image augmentation processing system 100 includes: a server 200, a network 300, and terminals (exemplarily shown as terminals 400-1 and 400-2). The terminals are connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0090] A server 200, configured to perform a series of processes including encoding, modulating, decoding, and superimposing on a target image based on the target image sent by a terminal, automatically generate an augmented image, train an image classification model according to the target image and the augmented image, and send the trained image classification model to the terminal.
[0091] A terminal, configured to run the image classification model sent by the server 200 according to an identification task to identify a target object in the target image, and perform subsequent tasks based on the identified target object.
[0092] In some embodiments, taking the image classification model as a face recognition model as an example, the terminal sends a target image to the server 200; the server 200 performs a series of processes including encoding, modulating, decoding, and superimposing on the target image to generate an augmented image corresponding to the target image, trains the face recognition model according to the target image and the augmented image, and sends the trained face recognition model to the terminal; a client of the terminal (such as a picture editing program) runs the face recognition model sent by the server 200 according to the target image uploaded by the user to perform an identification task, so as to automatically identify the face area in the target image, and display an editing tool for the user to perform further editing tasks such as adding special effects or changing faces based on the face area.
[0093] Combined with the exemplary applications and implementations of the server provided in the embodiments of the present application, it can be understood from the above that the image augmentation processing method based on artificial intelligence provided in the embodiments of the present application can be widely applied to image classification scenarios. For example, it can be applied to the field of remote sensing image recognition to perform image augmentation on satellite remote sensing images, and then perform image recognition based on the image classification model trained using the augmented image to improve the accuracy of terrain and geological exploration results; in the field of smart home, perform image augmentation on the images captured by a camera, and then perform image recognition based on the image classification model trained using the augmented image to improve the recognition degree and accuracy of the image content; in the medical field, perform image augmentation on scanned images, and then perform image recognition based on the image classification model trained using the augmented image to more accurately and quickly distinguish scanned images such as magnetic resonance imaging (MRI) and computed tomography (CT). In addition, scenarios related to image augmentation processing all belong to the potential application scenarios of the embodiments of the present application.
[0094] In the above fields, since the number of image samples is limited, the augmented images are generated by the artificial intelligence-based image augmentation processing method according to the embodiments of the present application to expand the number of samples, and then the image classification model is trained based on the expanded samples to improve the generalization ability of the image classification model; it can accurately identify the target image and has good anti-interference ability.
[0095] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server 200 may be directly or indirectly connected through wired or wireless communication means, which are not limited in the embodiments of the present application.
[0096] The following will make a detailed description of the hardware structure of the electronic device for the artificial intelligence-based image augmentation processing method provided by the embodiments of the present application. Taking the electronic device as Figure 1 the illustrated server 200 as an example, refer to Figure 2 , Figure 2 which is a schematic structural diagram of the server 200 for the artificial intelligence-based image augmentation processing provided by the embodiments of the present application. Figure 2 The server 200 shown in Figure 2 includes: at least one processor 210, a memory 250, at least one network interface 220, and a user interface 230. Each component in the server 200 is coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear description, in
[0097] all kinds of buses are labeled as the bus system 240.
[0098] The user interface 230 includes one or more output devices 231 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 230 also includes one or more input devices 232, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.
[0099] The memory 250 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid state memory, hard disk drives, optical disk drives, etc. The memory 250 optionally includes one or more storage devices that are physically remote from the processor 210.
[0100] The memory 250 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.
[0101] In some embodiments, the memory 250 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustratively described below.
[0102] The operating system 251, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0103] The network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;
[0104] The presentation module 253 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 231 associated with the user interface 230 (such as a display screen, a speaker, etc.).
[0105] The input processing module 254 is used to detect and translate one or more user inputs or interactions from one of one or more input devices 232.
[0106] In some embodiments, the artificial intelligence-based image augmentation processing apparatus provided by the embodiments of the present application may be implemented in software. Figure 2 Shown is the artificial intelligence-based image augmentation processing apparatus 255 stored in the memory 250, which may be software in the form of a program and plug-ins, etc., including the following software modules: a neural network model 2551, an overlay module 2552, a first training module 2553, and a second training module 2554. Among them, the neural network model 2551 includes an encoding and decoding model, and the encoding and decoding model includes a first encoder, a modulator, and a decoder. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.
[0107] In other embodiments, the artificial intelligence-based image augmentation processing apparatus provided by the embodiments of the present application may be implemented in hardware. As an example, the artificial intelligence-based image augmentation processing apparatus provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the artificial intelligence-based image augmentation processing method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0108] It can be understood that the artificial intelligence-based image augmentation processing method provided by the embodiments of the present application may be executed by an electronic device, and the electronic device includes but is not limited to a server or a terminal. The artificial intelligence-based image augmentation processing method provided by the embodiments of the present application will be described below in combination with an exemplary application in which the artificial intelligence-based image augmentation processing method provided by the embodiments of the present invention is implemented as a server.
[0109] See Figure 3 and Figure 4A , Figure 3 is a schematic structural diagram of the artificial intelligence-based image augmentation processing apparatus 255 provided by the embodiments of the present application. Figure 4A is an optional flowchart of the artificial intelligence-based image augmentation processing method provided by the embodiments of the present application. The steps shown will be described below in combination with Figure 3 for Figure 4A illustration.
[0110] In step 101, the target image is encoded to obtain the image features of the target image.
[0111] In some embodiments, based on Figure 3 , refer to Figure 5 , Figure 5 is a schematic structural diagram of the encoding and decoding model provided by the embodiments of the present application. Figure 3 The encoding and decoding model in Figure 5 is shown as follows. The encoding and decoding model includes a first encoder, a modulator, and a decoder.
[0112] The encoding process of the target image is implemented by the first encoder in the encoding and decoding model. The first encoder is used to extract the image features of the target image (for example, in the face recognition scenario, the image features are face features), that is, to compress the target image into a feature map including the image features.
[0113] In some examples, the first encoder is implemented by a downsampling layer (for example, a convolutional layer), and may include multiple cascaded downsampling layers to extract deep face features. Since the same target at different pixel positions in a target image has basically the same features, each downsampling layer extracts the same features at different pixel positions of the target image through the convolution operation of a convolution kernel. By compressing the image through the downsampling layer, a thumbnail feature map of the target image is generated, and the target image can be downsampled according to the pixel positions reflecting the image feature regions to obtain a feature map including the face features of the image.
[0114] In step 102, the feature library is queried based on the first label type of the target image to obtain the first feature parameter of the normal distribution followed by the first label type.
[0115] In some embodiments, refer to Figure 5 , the feature library stores the mapping relationship between different label types and different feature parameters; querying the feature library based on the first label type of the target image to obtain the first feature parameter of the normal distribution followed by the first label type includes: querying the mapping relationship stored in the feature library based on the first label type of the target image to obtain the first feature parameter of the normal distribution corresponding to the first label tag.
[0116] In some examples, the characteristic parameters of the normal distribution followed by the first marker type are the characteristic parameters of the normal distribution followed by the image features of multiple images of the first marker type. For example, when the target image is a face image, the first marker type is a face; when the target image is a non-face image, the first marker type is non-face; a mapping relationship between the first marker type and a feature vector group is stored in the feature library; when the first marker type corresponding to the target image is a face, the feature vector group with the first marker type being a face is retrieved from the database, and when the first marker type corresponding to the target image is non-face, the feature vector group with the first marker type being non-face is retrieved from the database. Here, the first marker type is the pre-marked type of the target image.
[0117] In step 103, the image features are subjected to noise modulation processing based on the first characteristic parameter to obtain the first adversarial noise feature.
[0118] In some embodiments, referring to Figure 5 , the image features are subjected to noise modulation processing by the modulator of the codec model. The characteristic parameters of the normal distribution followed by the first marker type include a first mean vector and a first variance vector; wherein, the first mean vector is used to represent the mean of the image features of the first marker type, and the first variance vector is used to represent the degree of jitter of the image features of the first marker type.
[0119] Subjecting the image features to noise modulation processing based on the first characteristic parameter to obtain the first adversarial noise feature includes: determining a first difference between the first mean vector and the image features of the target image; and determining the first adversarial noise feature based on a first ratio between the first difference and the first variance vector.
[0120] For example, assuming that the image features of the obtained face image are f, the first variance vector v and the first mean vector m corresponding to the face image are retrieved from the database. The image features f are subjected to noise modulation processing based on the first characteristic parameter to obtain the first adversarial noise feature. The specific calculation method is the first adversarial noise feature wherein, the noise modulation processing is to incorporate noise into the image features.
[0121] In some other embodiments, the first marker type can also be the image type corresponding to the classification probability distribution of the first classification model. The first classification model and the codec model are integrated into an adversarial noise generation model. The input can be any target image. The classification type of the target image is obtained according to the first classification model, and this classification type is used as the first marker type of the target image. An augmented image of the target image is obtained according to the target image and the first marker type.
[0122] In the embodiments of the present application, by using the characteristic parameters of the normal distribution followed by the image features of multiple images of the first marker type to modulate the target image corresponding to the first marker type, it is possible to specifically learn the image features of the target image, incorporate noise into the image features of the target image, improve the training accuracy of the codec model, and moreover, the generated first adversarial noise related to the first marker type of the image features can accelerate the fitting speed of the codec model.
[0123] In step 104, the first adversarial noise feature is decoded to obtain a first adversarial noise image.
[0124] In some embodiments, based on Figure 3 , see Figure 5 , the decoding process of the first adversarial noise feature is implemented by a decoder in the codec model. The decoder is used to restore the first adversarial noise image according to the first adversarial noise feature (for example, in the application of the face recognition scenario, the image feature is the face feature), that is, to enlarge the first adversarial noise image to the size of the target image by means of image interpolation in the first adversarial noise feature map to obtain the first adversarial noise image.
[0125] In some examples, the decoder is implemented by an upsampling layer, which may include multiple cascaded upsampling layers. The upsampling operation performed by the upsampling layer includes interpolation processing and deconvolution processing. Among them, interpolation processing refers to inserting new elements between pixel points on the basis of the pixels of the first adversarial noise feature map using a suitable interpolation algorithm, and deconvolution processing refers to improving the vertical resolution of data by compressing basic wavelets.
[0126] In step 105, the target image and the first adversarial noise image are superimposed to obtain a noise perturbation image, which is used as an augmented image of the target image.
[0127] In some embodiments, see Figure 3 , the target image and the first adversarial noise image are superimposed by a superimposing module to obtain a noise perturbation image. Obtain the pixel values, valid values, and transparency values of the red components of the pixel points at the same position of each layer of the target image and the first adversarial noise image; as well as, the pixel values, valid values, and transparency values of the green components; and, the pixel values, valid values, and transparency values of the blue components; respectively calculate the product sum of the pixel values, valid values, and transparency values of the red components at the same position of each layer, the product sum of the pixel values, valid values, and transparency values of the green components, and the product sum of the pixel values, valid values, and transparency values of the blue components; output layer superimposed data according to the product sums of the red, blue, and green components to obtain a noise perturbation image.
[0128] In some other embodiments, the color of a pixel can also be represented in the luminance - blue chrominance - red chrominance YcbCr color space, where Y represents luminance, Cb represents blue chrominance, and Cr represents red chrominance; obtain the pixel values, valid values, and transparency values of the luminance of the pixels at the same position in each layer of the target image and the first adversarial noise image; and, the pixel values, valid values, and transparency values of the green chrominance; and, the pixel values, valid values, and transparency values of the blue chrominance; respectively calculate the sum of products of the pixel values, valid values, and transparency values of the luminance of the pixels at the same position in each layer, the sum of products of the pixel values, valid values, and transparency values of the green chrominance, and the sum of products of the pixel values, valid values, and transparency values of the blue chrominance; output layer superposition data according to the sum of products of luminance, the sum of products of blue chrominance, and the sum of products of green chrominance to obtain a noise perturbation image.
[0129] In some embodiments, referring to Figure 4B , Figure 4B is an optional flowchart of an image augmentation processing method based on artificial intelligence provided by an embodiment of the present application. Based on Figure 4A , after step 105, step 106 and step 107 can also be executed.
[0130] In step 106, a training set is established based on the target image and the noise perturbation image of the target image.
[0131] In step 107, a second classification model is trained based on the training set; wherein, the labeled data in the training set is labeled according to the classification task of the second classification model, and the classification task of the second classification model is different from that of the first classification model.
[0132] For example, the second classification model is Figure 1 the image classification model in
[0133] which is trained based on the training set to improve the recognition rate and accuracy of the image classification model. The classification task can be to identify whether a face image is included in the target image, whether the target image is a high - definition image, etc. Taking the classification task of whether a face image is included in the target image as an example, when a face image is included in the target image, the corresponding label is 1; when a face image is not included in the target image, the corresponding label is 0.
[0134] In some embodiments, the encoding process of the image sample is implemented by the first encoder in the codec model, the noise modulation process is implemented by the modulator in the codec model, and the decoding process is implemented by the decoder in the codec model; based on Figure 4A , referring toFigure 4C , Figure 4C is an optional flowchart of an image augmentation processing method based on artificial intelligence provided by an embodiment of the present application, Figure 4C showing that the following training operations can also be iteratively executed before step 101: step 108 and step 109, which will be described in conjunction with each step below.
[0135] In step 108, based on the classification probability distribution of the image sample by the first classification model and the classification probability distribution of the noise-perturbed image of the image sample, jointly train the codec model, the first classification model, and the feature library;
[0136] In step 109, based on the classification probability distribution of the image sample by the first classification model and the classification probability distribution of the noise-perturbed image of the image sample, train the first classification model.
[0137] In some embodiments, based on the classification probability distribution of the image sample by the first classification model and the classification probability distribution of the noise-perturbed image of the image sample, jointly training the codec model, the first classification model, and the feature library includes: generating a noise-perturbed image of the image sample through the codec model; generating the classification probability distribution of the noise-perturbed image of the image sample and the classification probability distribution of the image sample through the first classification model; constructing a first loss function based on the difference between the classification probability distribution of the noise-perturbed image and the classification probability distribution of the image sample, and updating the model parameters of the codec model, the model parameters of the first classification model, and the feature parameters of the feature library by maximizing the first loss function. Here, the first classification model is used to assist in training the codec model.
[0138] The classification probability distribution of the noise-perturbed image of the image sample includes the probabilities that the noise-perturbed image of the image sample belongs to the original image type and the noise image type respectively; the classification probability distribution of the image sample includes the probabilities that the image sample belongs to the original image type and the noise image type respectively.
[0139] Updating the model parameters of the codec model, the model parameters of the first classification model, and the feature parameters of the feature library by maximizing the first loss function includes: determining the gradient values of each layer of the codec model, the gradient values of the first classification model, and the gradient values of the feature parameters of each type of the feature library when maximizing the first loss function; updating the model parameters of each layer of the codec model based on the gradient values of each layer of the codec model; updating the model parameters of each layer of the first classification model based on the gradient values of each layer of the first classification model; updating the feature parameters of the normal distribution followed by the second label type in the feature library based on the gradient values of the feature parameters of each type of the feature library; where the second label type is the pre-label type of the image sample.
[0140] In some embodiments, the model parameters of the modulator include modulation bias parameters and modulation rate parameters; the training of the encoding and decoding model further includes: determining the gradient value of the modulator when maximizing the first loss function, and updating the modulation bias parameters and modulation rate parameters based on the gradient value of the modulator.
[0141] In other embodiments, when the modulator only includes two model parameters, namely the modulation bias parameter b and the modulation rate parameter a, at this time, the training process of the modulator can be omitted.
[0142] In some embodiments, generating a noise perturbation image of an image sample through an encoding and decoding model includes: performing downsampling on the image sample through a first encoder to obtain the image features of the image sample; querying a feature library based on the second label type of the image sample to obtain the second feature parameters of the normal distribution followed by the second label type of the image sample; wherein, the second label type is the pre-label type of the image sample.
[0143] Performing noise modulation processing on the image features of the image sample through the modulator and based on the second feature parameters to obtain second adversarial noise features;
[0144] Performing upsampling on the second adversarial noise features through a decoder to obtain a noise perturbation image of the image sample.
[0145] In some examples, performing downsampling on an image sample, that is, the process of shrinking the image sample to obtain the local image features of the image sample, can be implemented according to the relevant techniques of pooling. The purpose is to reduce the dimension of the features and only retain the local image features, to a certain extent avoiding overfitting. For example, downsampling includes maximum sampling, average sampling, sum region sampling, and random region sampling, etc. For example, for average sampling, for an image I with size M*N, performing s-fold downsampling on it will obtain an image with a resolution of, for an image in matrix form, changing the image within an s*s window of the original image sample into one pixel, and the value of this pixel is the average value of all pixels within the window.
[0146] Performing upsampling on the second adversarial noise features, that is, the process of enlarging the second adversarial noise features through image interpolation. For example, using the inner interpolation method, that is, on the basis of the pixels of the second adversarial noise feature map, inserting new pixels between the pixel points by using an interpolation algorithm to enlarge the second adversarial noise feature map to the size of the original image sample.
[0147] In some embodiments, based on Figure 3 , see Figure 6 , Figure 6 is the structural schematic diagram of the first classification model provided by the embodiments of the present application.Figure 3 The first classification model in Figure 6 is shown as follows. The first classification model includes a second encoder, a third encoder, and a classifier; generating the classification probability distribution of the noise-perturbed image of the image sample and the classification probability distribution of the image sample through the first classification model, including: performing feature extraction processing on the noise-perturbed image of the image sample through the second encoder and the third encoder to obtain the image features of the noise-perturbed image of the image sample, and mapping the image features of the noise-perturbed image of the image sample to the classification probability distribution of the noise-perturbed image of the image sample through the classifier; performing feature extraction processing on the image sample through the second encoder and the third encoder to obtain the image features of the image sample, and mapping the image features of the image sample to the classification probability distribution of the image sample through the classifier.
[0148] For example, the classifier may include a fully connected layer and a logistic regression softmax function. The fully connected layer integrates all the obtained features into a feature vector, and the logistic regression softmax function is used to classify this feature vector to output the classification probability distribution of the target image.
[0149] In some embodiments, the structures of the first encoder and the second encoder are the same and share the same model parameters, that is, the first encoder in the codec model and the second encoder in the first classification model update the model parameters in a weight-sharing manner.
[0150] In other embodiments, the first encoder in the codec model and the second encoder in the first classification model may also update the model parameters in a weight-independent manner. That is, a first loss function is constructed based on the difference between the classification probability distribution of the noise-perturbed image of the image sample and the classification probability distribution of the image sample, and the model parameters of the codec model and the feature parameters of the feature library are updated by maximizing the first loss function; a second loss function is constructed according to the difference between the classification probability distribution of the image sample by the first classification model and the classification probability distribution of the noise-perturbed image of the image sample; the model parameters of the first classification model are updated by minimizing the second loss function. Here, the model parameters of the first encoder in the codec model and the model parameters of the second encoder in the first classification model are independent of each other.
[0151] In some embodiments, training the first classification model based on the classification probability distribution of the image sample by the first classification model and the classification probability distribution of the noise-perturbed image of the image sample includes: classifying the noise-perturbed image and the image sample of the image sample through the first classification model to obtain the probability that the noise-perturbed image of the image sample belongs to the original image type and the probability that the image sample belongs to the original image type; constructing a second loss function according to the difference between the probability that the image sample belongs to the original image type by the first classification model and the probability that the noise-perturbed image of the image sample belongs to the original image type; and updating the model parameters of the first classification model by minimizing the second loss function.
[0152] In some examples, updating the model parameters of the first classification model by minimizing the second loss function includes: determining the gradient values of the fully connected layer in the classifier and the gradient values of each layer in the second encoder and the third encoder when the second loss function obtains the minimum value; updating the model parameters of the classifier according to the gradient values of the fully connected layer of the classifier, and updating the model parameters of the second encoder and the third encoder respectively according to the gradient values of each layer in the second encoder and the third encoder.
[0153] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described. Taking a binary classification face recognition (i.e., identifying whether an image is a face image) model as an example, the developer generates a noise-perturbed image through a trained encoding and decoding model. Training the binary classification face recognition model according to the generated noise-perturbed image can reduce the overfitting of the binary classification face recognition model on the training set, improve the accuracy and recall rate of the binary classification face recognition model for the training set and online data, integrate the binary classification face recognition model function into the face-swiping application program, and when the user uses the face-swiping application program, the user's face can be efficiently and accurately recognized, improving the user experience. Refer to Figure 7 , Figure 7 is a schematic flowchart of the training method of the encoding and decoding model provided by the embodiments of the present application. The training method of the encoding and decoding model provided by the embodiments of the present application includes:
[0154] Step 201: Input the original image I into the second encoder to obtain the image feature f.
[0155] Step 202: Continuously input the image feature f into the third encoder to continue extracting features and classifying to obtain the classification probability distribution P(y|original image) of the original image I.
[0156] Step 203: According to the label of the original image, select the corresponding feature vector group (m, v) from the feature library, then modulate f to obtain and input it into the decoder to obtain the adversarial noise N.
[0157] Step 204: Superimpose the original image I and the adversarial noise N to obtain a noise-perturbed image I’ = I + N. Input I’ into the second encoder, and then through the third encoder, to obtain the classification probability distribution P(y|noisy image) of the noise-perturbed image.
[0158] Step 205: Maximize the distance between the classification probability distributions corresponding to I and I’, and backpropagate the gradient accordingly to update the parameters of the second encoder, the decoder, and the feature library. This is used to train the generative model to generate a noise-perturbed image related to the original image.
[0159] Step 206: Re-input the original image I and the noise-perturbed image into the second encoder and the third encoder, and minimize the classification loss function of I and I’, to update the parameters of the second encoder and the third encoder. This is used to train the first classification model so that the recognition results of both I and I’ are I.
[0160] Step 207: Stop training after training for a preset number of times.
[0161] In some examples, refer to Figure 8 , Figure 8 which is a schematic structural diagram of the training encoder-decoder model provided by an embodiment of the present application. The first classification model includes a second encoder, a third encoder, and a classifier. The output of the second encoder is image features (not shown in the figure). The process from the second encoder to the third encoder is a process of continuous abstraction of the image features and continuous increase in dimensions. The output of the second encoder is input into the third encoder again. The output of the third encoder is deeper image features. Based on this image feature, the probability that the corresponding image belongs to a face image is output through the classifier, that is, the classification probability distribution P(y|original image) of the original face image I and the classification probability distribution P(y|noisy image) of the noise-perturbed image. In the training stage, this binary classification face recognition model has two types of inputs. One type is the original image, such as a face image or a non-face image, and the other type is the image after this original image is perturbed by noise, that is, the noise-perturbed image. Both types of images will be input into the first classification model for classification.
[0162] Refer to Figure 8, the noise-disturbed image is mainly generated by an encoding and decoding model, which includes a first encoder, a modulator, and a decoder. The first encoder can share model parameters with the second encoder in the first classification model, and its decoder is independent. The output of the decoder is a noise-disturbed image. The input of this first classification model is an original image and the label corresponding to the original image. For example, here a face image can be represented by label 1, and a non-face image can be represented by label 0. For the feature library, the feature library has several groups of feature vectors, and the specific quantity corresponds to the categories that the classifier needs to distinguish. For example, here there are a total of 2. Each group of feature vectors contains two vectors, which respectively represent the first mean vector m and the first variance vector v of the features of this category in the high-dimensional space. If the label corresponding to the image is 0, the feature vector group numbered 0 in the feature library is retrieved through query. If the label is 1, the feature vector group numbered 1 is retrieved from the feature library through query. Assume that the output of the first encoder is the image feature f, and the feature vector group is used to modulate f. The specific method is Then, f’ is input into the decoder to obtain an adversarial noise image corresponding to the original image. The original image and the adversarial noise image are superimposed to obtain a noise-disturbed image, which is used as the augmented image of the original image.
[0163] Training the binary classification face recognition model based on the augmented image can reduce the overfitting of the binary classification face recognition model on the training set, improve the accuracy and recall rate of the binary classification face recognition model for the training set and online data, integrate the function of the binary classification face recognition model into the face-swiping application program, and when the user uses the face-swiping application program, the user's face can be recognized efficiently and accurately, improving the user experience. Here, the binary classification face recognition model is Figure 3 the second classification model shown in
[0164] Here, the second encoder in the first classification model and the first encoder in the generation model can also be constructed in a weight-independent manner, that is, the first encoder and the second encoder use different model parameters; the encoding and decoding model can also directly generate the noise-disturbed image I’; the feature vectors used for feature modulation here can also be generated by a fully connected network to generate two vectors corresponding to the first mean vector m and the first variance vector v respectively.
[0165] Next, continue to describe the exemplary structure of the software module implementation of the artificial intelligence-based image augmentation processing device 255 provided in the embodiments of the present application. In some embodiments, as Figure 2 shown, the software module stored in the artificial intelligence-based image augmentation processing device 255 in the memory 240 may include:
[0166] The neural network module 2551 is used to encode the target image to obtain the image features of the target image; query the feature library based on the first label type of the target image to obtain the first feature parameters of the normal distribution followed by the first label type; perform noise modulation processing on the image features based on the first feature parameters to obtain the first adversarial noise features; perform decoding processing on the first adversarial noise features to obtain the first adversarial noise image; The superimposing module 2552 is used to superimpose the target image and the first adversarial noise image to obtain a noise perturbation image as the augmented image of the target image.
[0167] In some embodiments, the feature library stores the mapping relationship between different label types and different feature parameters; the neural network module 2551 is further used to query the mapping relationship stored in the feature library based on the first label type of the target image to obtain the first feature parameters of the normal distribution corresponding to the first label tag.
[0168] In some embodiments, the feature parameters of the normal distribution followed by the first label type include a first mean vector and a first variance vector; wherein, the first mean vector is used to characterize the mean of the image features of the first label type, and the first variance vector is used to characterize the jitter degree of the image features of the first label type; the neural network module 2551 is further used to determine the first difference between the first mean vector and the image features of the target image; determine the first adversarial noise feature by the first ratio between the first difference and the first variance vector.
[0169] In some embodiments, the encoding process is implemented by the first encoder in the codec model, the noise modulation process is implemented by the modulator in the codec model, and the decoding process is implemented by the decoder in the codec model; An image augmentation processing device based on artificial intelligence provided by an embodiment of the present application further includes: a first training module 2553, configured to iteratively perform the following training operations: jointly train the codec model, the first classification model, and the feature library based on the classification probability distribution of the image samples by the first classification model and the classification probability distribution of the noise perturbation images of the image samples; train the first classification model based on the classification probability distribution of the image samples by the first classification model and the classification probability distribution of the noise perturbation images of the image samples.
[0170] In some embodiments, the first training module 2553 is further used to generate a noise perturbation image of the image sample through the codec model; generate the classification probability distribution of the noise perturbation image of the image sample and the classification probability distribution of the image sample through the first classification model; construct a first loss function based on the difference between the classification probability distribution of the noise perturbation image and the classification probability distribution of the image sample, and update the model parameters of the codec model, the model parameters of the first classification model, and the feature parameters of the feature library by maximizing the first loss function.
[0171] In some embodiments, the classification probability distribution of the noise-perturbed image of the image sample includes the probabilities that the noise-perturbed image of the image sample belongs to the original image type and the noise image type respectively; the classification probability distribution of the image sample includes the probabilities that the image sample belongs to the original image type and the noise image type respectively; the first training module 2553 is further configured to determine the gradient values of the codec model, the first classification model, and the feature library when maximizing the first loss function; update the model parameters of the codec model based on the gradient value of the codec model; update the model parameters of the first classification model based on the gradient value of the first classification model; update the feature parameters of the normal distribution followed by the second label type in the feature library based on the gradient value of the feature library; wherein the second label type is the pre-label type of the image sample.
[0172] In some embodiments, the model parameters of the modulator include a modulation bias parameter and a modulation rate parameter; the first training module 2553 is further configured to determine the gradient value of the modulator when maximizing the first loss function, and update the modulation bias parameter and the modulation rate parameter based on the gradient value of the modulator.
[0173] In some embodiments, the first training module 2553 is further configured to perform downsampling processing on the image sample through the first encoder to obtain the image features of the image sample; query the feature library based on the second label type of the image sample to obtain the second feature parameters of the normal distribution followed by the second label type of the image sample; wherein the second label type is the pre-label type of the image sample; perform noise modulation processing on the image features of the image sample through the modulator and based on the second feature parameters to obtain second adversarial noise features; and perform upsampling processing on the second adversarial noise features through the decoder to obtain the noise-perturbed image of the image sample.
[0174] In some embodiments, the first classification model includes a second encoder, a third encoder, and a classifier; the first training module 2553 is further configured to perform feature extraction processing on the noise-perturbed image of the image sample through the second encoder and the third encoder to obtain the image features of the noise-perturbed image of the image sample, and map the image features of the noise-perturbed image of the image sample to the classification probability distribution of the noise-perturbed image of the image sample through the classifier; perform feature extraction processing on the image sample through the second encoder and the third encoder to obtain the image features of the image sample, and map the image features of the image sample to the classification probability distribution of the image sample through the classifier.
[0175] In some embodiments, the first encoder and the second encoder have the same structure and share the same model parameters.
[0176] In some embodiments, the first training module 2553 is further configured to classify the noise-perturbed image of the image sample and the image sample through the first classification model, to obtain the probability that the noise-perturbed image of the image sample belongs to the original image type and the probability that the image sample belongs to the original image type; construct a second loss function according to the difference between the probability that the image sample belongs to the original image type by the first classification model and the probability that the noise-perturbed image of the image sample belongs to the original image type; and update the model parameters of the first classification model by minimizing the second loss function.
[0177] In some embodiments, an image augmentation processing apparatus based on artificial intelligence provided by an embodiment of the present application further includes: a second training module 2554, configured to establish a training set based on a target image and a noise-perturbed image of the target image; train a second classification model based on the training set; wherein the annotation data in the training set is annotated according to the classification task of the second classification model, and the classification task of the second classification model is different from that of the first classification model.
[0178] An embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned image augmentation processing method based on artificial intelligence in the embodiment of the present application.
[0179] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions are stored, and when the executable instructions are executed by a processor, the processor will be caused to execute the image augmentation processing method based on artificial intelligence provided by the embodiment of the present application. For example, Figure 4A 、 4B 、the image augmentation processing method based on artificial intelligence shown in 4C.
[0180] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0181] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted language, or declarative or procedural language), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0182] As an example, the executable instructions may or may not correspond to files in a file system, and may be stored as part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program under discussion, or stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or code portions).
[0183] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed across multiple locations and interconnected by a communication network.
[0184] In summary, through the embodiments of the present application, it is possible to automatically generate noise perturbation images related to the marking type, effectively improving the performance and quality of image augmentation; by adding the target image and the noise perturbation image of the target image to the training set to further train the second classification model, the generalization ability of the second classification model on the training set and online data can be improved; training the classification model according to the generated noise perturbation image can reduce the overfitting of the classification model on the training set, and improve the accuracy and recall rate of the classification model for the training set and online data; integrating the classification model function into the application program enables users to classify efficiently and accurately when using the application program, enhancing the user experience.
[0185] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. An image augmentation processing method based on artificial intelligence, characterized in that, Including: Performing encoding processing on a target image to obtain image features of the target image; Querying a mapping relationship stored in a feature library based on a first marking type of the target image to obtain first feature parameters of a normal distribution followed by the first marking type; different mapping relationships between different marking types and different feature parameters are stored in the feature library; the feature parameters of the normal distribution followed by the first marking type include a first mean vector and a first variance vector; Performing noise modulation processing on the image features based on the first feature parameters to obtain first adversarial noise features; wherein, performing noise modulation processing on the image features based on the first feature parameters to obtain first adversarial noise features includes: determining a first difference between the first mean vector and the image features of the target image; determining a first ratio between the first difference and the first variance vector as the first adversarial noise features; Performing decoding processing on the first adversarial noise features to obtain a first adversarial noise image; Superimposing the target image and the first adversarial noise image to obtain a noise perturbation image as an augmented image of the target image.
2. The method according to claim 1, characterized in that, The first mean vector is used to represent the mean of the image features of the first marking type, and the first variance vector is used to represent the jitter degree of the image features of the first marking type.
3. The method according to claim 1, characterized in that, The encoding processing is implemented by a first encoder in an encoding and decoding model, the noise modulation processing is implemented by a modulator in the encoding and decoding model, and the decoding processing is implemented by a decoder in the encoding and decoding model; Before performing encoding processing on the target image, it further includes: Iteratively performing the following training operations: Jointly training the encoding and decoding model, the first classification model, and the feature library based on the classification probability distribution of an image sample by the first classification model and the classification probability distribution of a noise perturbation image of the image sample; Training the first classification model based on the classification probability distribution of the image sample by the first classification model and the classification probability distribution of a noise perturbation image of the image sample.
4. The method according to claim 3, characterized in that, The jointly training the encoding and decoding model, the first classification model, and the feature library based on the classification probability distribution of an image sample by the first classification model and the classification probability distribution of a noise perturbation image of the image sample includes: Generating a noise perturbation image of the image sample through the encoding and decoding model; Generating the classification probability distribution of the noise perturbation image of the image sample and the classification probability distribution of the image sample through the first classification model; Constructing a first loss function based on the difference between the classification probability distribution of the noise perturbation image and the classification probability distribution of the image sample, and updating the model parameters of the encoding and decoding model, the model parameters of the first classification model, and the feature parameters of the feature library by maximizing the first loss function.
5. The method according to claim 4, characterized in that, The classification probability distribution of the noise perturbation image of the image sample includes the probabilities that the noise perturbation image of the image sample belongs to the original image type and the noise image type respectively; The classification probability distribution of the image sample includes the probabilities that the image sample belongs to the original image type and the noise image type respectively; The method of updating the model parameters of the codec model, the model parameters of the first classification model, and the feature parameters of the feature library by maximizing the first loss function includes: Determining the gradient values of the codec model, the first classification model, and the feature library when maximizing the first loss function; Updating the model parameters of the codec model based on the gradient value of the codec model; Updating the model parameters of the first classification model based on the gradient value of the first classification model; Updating the feature parameters of the normal distribution followed by the second label type in the feature library based on the gradient value of the feature library; wherein, the second label type is the pre-label type of the image sample.
6. The method according to claim 5, characterized in that, The model parameters of the modulator include a modulation bias parameter and a modulation rate parameter; The method further includes: Determining the gradient value of the modulator when maximizing the first loss function, and updating the modulation bias parameter and the modulation rate parameter based on the gradient value of the modulator.
7. The method according to claim 4, characterized in that, Generating the noise perturbation image of the image sample through the codec model includes: Performing downsampling processing on the image sample through the first encoder to obtain the image features of the image sample; Querying the feature library based on the second label type of the image sample to obtain the second feature parameters of the normal distribution followed by the second label type of the image sample; wherein, the second label type is the pre-label type of the image sample; Performing noise modulation processing on the image features of the image sample through the modulator and based on the second feature parameters to obtain second adversarial noise features; Performing upsampling processing on the second adversarial noise features through the decoder to obtain the noise perturbation image of the image sample.
8. The method according to claim 4, characterized in that The first classification model includes a second encoder, a third encoder, and a classifier; Generating the classification probability distribution of the noise perturbation image of the image sample and the classification probability distribution of the image sample through the first classification model, includes: Performing feature extraction processing on the noise perturbation image of the image sample through the second encoder and the third encoder to obtain the image features of the noise perturbation image of the image sample, and mapping the image features of the noise perturbation image of the image sample to the classification probability distribution of the noise perturbation image of the image sample through the classifier; Performing feature extraction processing on the image sample through the second encoder and the third encoder to obtain the image features of the image sample, and mapping the image features of the image sample to the classification probability distribution of the image sample through the classifier.
9. The method according to claim 8, characterized in that The structures of the first encoder and the second encoder are the same and share the same model parameters.
10. The method according to claim 3, characterized in that Training the first classification model based on the classification probability distribution of the image sample and the classification probability distribution of the noise perturbation image of the image sample through the first classification model, includes: Classify the noise-perturbed image and the image sample of the image sample through the first classification model to obtain the probability that the noise-perturbed image of the image sample belongs to the original image type and the probability that the image sample belongs to the original image type; Construct a second loss function according to the difference between the probability that the image sample belongs to the original image type by the first classification model and the probability that the noise-perturbed image of the image sample belongs to the original image type; Update the model parameters of the first classification model by minimizing the second loss function.
11. The method according to any one of claims 3 to 10, characterized in that The method further includes: Establish a training set based on the target image and the noise-perturbed image of the target image; Train a second classification model based on the training set; Wherein, the labeled data in the training set is labeled according to the classification task of the second classification model, and the classification task of the second classification model is different from that of the first classification model.
12. An image augmentation processing device based on artificial intelligence, characterized in that Includes: A first encoder for encoding the target image to obtain the image features of the target image; A modulator for: Query the mapping relationship stored in the feature library based on the first label type of the target image to obtain the first feature parameters of the normal distribution followed by the first label type; different mapping relationships between different label types and different feature parameters are stored in the feature library; the feature parameters of the normal distribution followed by the first label type include a first mean vector and a first variance vector; Perform noise modulation processing on the image features based on the first feature parameters to obtain first adversarial noise features; wherein, performing noise modulation processing on the image features based on the first feature parameters to obtain first adversarial noise features includes: determining a first difference between the first mean vector and the image features of the target image; determining a first ratio between the first difference and the first variance vector as the first adversarial noise features; A decoder for decoding the first adversarial noise features to obtain a first adversarial noise image; An overlay module for overlaying the target image and the first adversarial noise image to obtain a noise-perturbed image as an augmented image of the target image.
13. An electronic device, characterized in that Includes: A memory for storing executable instructions; A processor for implementing the artificial intelligence-based image augmentation processing method according to any one of claims 1 to 11 when executing the executable instructions stored in the memory.
14. A computer-readable storage medium, characterized in that Stored with executable instructions for causing a processor to implement the artificial intelligence-based image augmentation processing method according to any one of claims 1 to 11 when executed.
15. A computer program product, the computer program product comprising executable instructions stored in a computer-readable storage medium; When a processor of an electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the image augmentation processing method based on artificial intelligence according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
A training sample data expansion method and device based on a variational auto-encoder
CN109886388A
Adversarial sample generation method and device, medium and computing device
CN110245598A