An image recognition method, device, system and storage medium
By establishing an image recognition model based on GAN fingerprints of each synthetic image in advance, the problems of low recognition accuracy and reduced detection performance in the prior art are solved, and high-precision image recognition suitable for a variety of GAN architectures are achieved.
Patent Information
- Application Number
- CN202210763806.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-06-30
AI Technical Summary
The prior art is difficult to apply to various types of GAN architectures, with low recognition accuracy and significantly reduced detection performance when the GAN network architecture changes.
By establishing an image recognition model based on the GAN fingerprints of each synthetic image in advance, obtaining the image to be identified and extracting its fingerprints, and using the image recognition model to classify and identify the fingerprints to determine whether the image to be identified is a synthetic image.
It realizes suitable for many types of GAN architectures, with high recognition accuracy, certain robustness, and improved detection performance.
Smart Images

Figure CN114913556B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an image recognition method, device, system and computer-readable storage medium. Background Art
[0002] Synthetic images refer to images generated by a Generative Adversarial Network (GANs). While the development of the GAN has advanced image generation technology by leaps and bounds, it has also raised concerns and panic in society about the abuse of face synthetic images. Therefore, technologies for detecting and attributing face synthetic images have emerged. However, most current detection methods can only identify images generated by GAN architectures with specific features, with a small scope of application, and the detection performance will be greatly reduced when the GAN network architecture changes.
[0003] In view of this, how to provide an image recognition method, device, system and computer-readable storage medium with stronger applicability and high recognition accuracy has become a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide an image recognition method, device, system and computer-readable storage medium, which can be applicable to various types of GAN architectures during use, have a wide scope of application, high recognition accuracy, certain robustness, and are beneficial to improving detection performance.
[0005] To solve the above technical problems, the embodiments of the present invention provide an image recognition method, including:
[0006] Obtain an image to be recognized;
[0007] Extract the fingerprint of the image to be recognized;
[0008] Use a pre-established image recognition model to classify and recognize the fingerprint to obtain a classification result corresponding to the fingerprint; the image recognition model is established based on the GAN fingerprints of each synthetic image;
[0009] Determine whether the image to be recognized is a synthetic image based on the classification result.
[0010] Optionally, the establishment of the image recognition model includes:
[0011] For each Generative Adversarial Network (GAN), use the GAN to generate multiple synthetic images in advance;
[0012] Extract the GAN fingerprint of each of the synthetic images, and add a GAN identification label to the fingerprint;
[0013] Input the GAN fingerprints with the GAN identification labels added into a classifier, and train the classifier to obtain an image recognition model.
[0014] Optionally, the extracting the GAN fingerprint of each of the synthetic images includes:
[0015] For each of the synthetic images, use an autoencoder to process the synthetic image to obtain a corresponding predicted image;
[0016] Calculate the reconstruction residual between the predicted image and the synthetic image;
[0017] Take the reconstruction residual as the GAN fingerprint corresponding to the synthetic image.
[0018] Optionally, the autoencoder is trained based on a preset loss function, and the preset loss function is:
[0019] L = ||Fx predicted - F(x input )|| 2 , where F(x input ) represents the input feature, Fx predicted represents the predicted result feature, and |||| 2 represents the distance L.
[0020] Optionally, the classification results include real images, images containing composite GAN fingerprints, and single GAN images.
[0021] Optionally, the determining whether the image to be recognized is a synthetic image based on the classification result includes:
[0022] When the classification result is a real image, the image to be recognized is a real image;
[0023] When the classification result is an image containing a composite GAN fingerprint or a single GAN image, the image to be recognized is a synthetic image.
[0024] Optionally, it further includes:
[0025] When the image to be recognized is a synthetic image, determine the GAN identification corresponding to the synthetic image based on the classification result.
[0026] An embodiment of the present invention further provides an image recognition device, including:
[0027] An acquisition module, configured to acquire an image to be recognized;
[0028] An extraction module, configured to extract the fingerprint of the image to be recognized;
[0029] A recognition module, configured to classify and recognize the fingerprint by using an image recognition model established in advance by a building module, and obtain a classification result corresponding to the fingerprint; the image recognition model is established based on GAN fingerprints of each synthetic image;
[0030] A judgment module, configured to determine whether the image to be recognized is a synthetic image based on the classification result.
[0031] Optionally, the building module includes:
[0032] A generation unit, configured to generate multiple synthetic images in advance for each generative adversarial network (GAN);
[0033] An extraction unit, configured to extract the GAN fingerprint of each synthetic image and add a GAN identification label to the fingerprint;
[0034] A training unit, configured to input each GAN fingerprint added with a GAN identification label into a classifier and train the classifier to obtain an image recognition model.
[0035] Optionally, the extraction unit includes:
[0036] A processing subunit, configured to process each synthetic image by using an autoencoder to obtain a corresponding predicted image;
[0037] A calculation subunit, configured to calculate the reconstruction residual between the predicted image and the synthetic image;
[0038] A determination subunit, configured to use the reconstruction residual as the GAN fingerprint corresponding to the synthetic image.
[0039] Optionally, the processing subunit is specifically configured to process each synthetic image by using an autoencoder trained based on a preset loss function to obtain a corresponding predicted image;
[0040] The preset loss function is:
[0041] L = ||Fx predicted - F(x input )|| 2 , where F(x input ) represents the input feature, Fx predicted represents the predicted result feature, and |||| 2 represents the distance L.
[0042] Optionally, the classification result includes a real image, an image containing a composite GAN fingerprint, and a single GAN image.
[0043] Optionally, the determination module includes:
[0044] A first determination unit, configured to determine that the image to be recognized is a real image when the classification result is a real image;
[0045] A second determination unit, configured to determine that the image to be recognized is a synthetic image when the classification result is an image containing a composite GAN fingerprint or a single GAN image.
[0046] Optionally, it further includes:
[0047] A determination module, configured to determine a GAN identifier corresponding to the synthetic image based on the classification result when the image to be recognized is a synthetic image.
[0048] An embodiment of the present invention further provides an image recognition system, including:
[0049] A memory, configured to store a computer program;
[0050] A processor, configured to implement the steps of the image recognition method as described above when executing the computer program.
[0051] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the image recognition method as described above are implemented.
[0052] An embodiment of the present invention provides an image recognition method, device, system, and computer-readable storage medium. The method includes: obtaining an image to be recognized; extracting a fingerprint of the image to be recognized; classifying and recognizing the fingerprint by using a pre-established image recognition model to obtain a classification result corresponding to the fingerprint, where the image recognition model is established based on GAN fingerprints of each synthetic image; and determining whether the image to be recognized is a synthetic image based on the classification result.
[0053] It can be seen that in the embodiment of the present invention, an image recognition model is pre-established based on GAN fingerprints of each synthetic image, then the image to be recognized is obtained and the fingerprint of the image to be recognized is extracted, the fingerprint is classified and recognized by using the image recognition model to obtain a corresponding classification result, and whether the image to be recognized is a synthetic image can be determined according to the classification result. The present invention can be applied to various types of GAN architectures, has a wide application range, high recognition accuracy, certain robustness, and is beneficial to improving the detection performance. Description of the Drawings
[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the prior art and the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0055] Figure 1 It is a schematic flowchart of an image recognition method provided by an embodiment of the present invention;
[0056] Figure 2 It is a schematic flowchart of another image recognition method provided by an embodiment of the present invention;
[0057] Figure 3 It is a schematic diagram of an image recognition model training process provided by an embodiment of the present invention;
[0058] Figure 4 It is a schematic diagram of an image recognition process provided by an embodiment of the present invention;
[0059] Figure 5 It is a schematic diagram of the structure of an identification terminal provided by an embodiment of the present application;
[0060] Figure 6 It is a schematic diagram of the structure of an image recognition device provided by an embodiment of the present invention;
[0061] Figure 7 It is a schematic diagram of the structure of an image recognition system provided by an embodiment of the present invention. Detailed implementation manners
[0062] The embodiments of the present invention provide an image recognition method, device, system and computer-readable storage medium, which can be applicable to various types of GAN architectures during use, have a wide application range, high recognition accuracy, certain robustness, and are beneficial to improving the detection performance.
[0063] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] It should be noted that with the society's concerns and panic about the abuse of face synthetic images, technologies for the detection and attribution of face synthetic images have emerged as the times require. However, the emergence of these technologies also brings some new problems. On the one hand, once a detection technology is proposed, it will be applied by the next-generation generation technology to adversarial learning, enabling the newly generated samples to escape the detection of this technology. Also, since most current detection methods only target specific GAN architectures, and there are various types of GAN architectures, corresponding detection methods need to be developed for different types of GAN architectures. Therefore, when inspecting the images to be detected, multiple detection methods may be needed to detect them to reduce misjudgments. In addition, for a detection method developed for a specific GAN architecture, when the GAN architecture changes during training or the architecture itself changes, the detection performance corresponding to this structure will drop significantly. Therefore, in the embodiments of the present invention, to solve this problem, an image recognition model is established in advance based on the GAN fingerprints of a large number of artificial synthetic images, and then the fingerprint of the image to be recognized is extracted, and the pre-established image recognition model is used to classify and recognize the fingerprint to obtain the classification result corresponding to the fingerprint. Then, based on this classification result, it is determined whether the image to be recognized is an artificial synthetic image, which can be applicable to various types of GAN network structures.
[0065] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a kind of image recognition method provided by the embodiments of the present invention.
[0066] The method includes:
[0067] S110: Obtain the image to be recognized;
[0068] It should be noted that the GAN fingerprint is a feature used to characterize the GAN network. By analogy with human fingerprints, since each person has a unique fingerprint, the identity of a person can be determined based on the fingerprint, that is, the identity of a person can be recognized based on the fingerprint. Similarly, each GAN network also has unique features (which can be called GAN fingerprints), and the images generated by each GAN network will carry this feature. Therefore, it is possible to determine which GAN network generated the image by recognizing the GAN fingerprint. Of course, an image generated by a GAN network is an artificial synthetic image. Therefore, in the embodiments of the present invention, the GAN fingerprints of a large number of artificial synthetic image samples can be used in advance to train the image recognition model to obtain the image recognition model. Specifically, when it is necessary to recognize whether the image to be recognized is an artificial synthetic image, first obtain the fingerprint of the image to be recognized.
[0069] S120: Extract the fingerprint of the image to be recognized;
[0070] Specifically, after obtaining the image to be recognized, extract the fingerprint of the image to be recognized.
[0071] S130: Use a pre-established image recognition model to classify and recognize the fingerprint, and obtain the classification result corresponding to the fingerprint; the image recognition model is established based on the GAN fingerprints of each synthetic image;
[0072] Specifically, after extracting the fingerprint of the image to be recognized, use a pre-established image recognition model to classify and recognize the fingerprint of the image to be recognized. Specifically, the fingerprint of the image to be recognized can be used as the input of the image recognition model, and the image recognition model will output the classification result after the recognition is completed.
[0073] S140: Determine whether the image to be recognized is a synthetic image based on the classification result.
[0074] It can be understood that after obtaining the recognition result output by the image recognition model, it can be further determined whether the image to be recognized is a synthetic image according to the classification result.
[0075] Further, please refer to Figure 2 , the establishment process of the image recognition model in the embodiments of the present invention may specifically include:
[0076] S210: For each generative adversarial network GAN, use the GAN to generate multiple synthetic images in advance;
[0077] It should be noted that each GAN network can be trained in advance to obtain each GAN, and then each trained GAN can be used to generate the required data set. Specifically, for each GAN, a large number of synthetic images can be generated through the GAN. These synthetic images can be used as a data set, that is, each data set includes multiple synthetic images, and each GAN corresponds to a data set, so as to obtain a large number of synthetic images for different GANs.
[0078] S220: Extract the GAN fingerprint of each synthetic image, and add a GAN identification label to the fingerprint;
[0079] Specifically, after obtaining a large number of synthetic images, for each synthetic image of a GAN, the GAN fingerprint of the synthetic image is extracted, and a corresponding label (specifically, it can be a GAN identifier) also needs to be added to the GAN fingerprint. That is, the corresponding GAN identifier is used as a label and added to the GAN fingerprint corresponding to the synthetic image. For each synthetic image generated by each GAN network, the corresponding GAN fingerprint is extracted and the corresponding GAN identifier label is added, so as to obtain the GAN fingerprint with the GAN identifier label corresponding to each synthetic image respectively. It should be noted that adding the GAN identifier label to the GAN fingerprint is to facilitate the targeted identification of each GAN fingerprint corresponding to each GAN during the subsequent training process. For example, for GAN1, all GAN fingerprints corresponding to GAN1 can be identified, etc.
[0080] S230: Input the GAN fingerprints with the GAN identifier labels into a classifier and train the classifier to obtain an image recognition model.
[0081] It can be understood that after obtaining the GAN fingerprints with the GAN identifier labels corresponding to each synthetic image respectively, these GAN fingerprints with the GAN identifier labels are used as a sample set for training the image recognition model. Specifically, these GAN fingerprints with the GAN identifier labels are input into a pre-established classifier for classification training, and after the training is completed, a trained image recognition model is obtained for subsequent use in image recognition of the image to be recognized. Among them, each GAN fingerprint has a corresponding label. Training the classifier with each GAN fingerprint with the GAN identifier label is conducive to better learning the characteristics of the GAN fingerprint of each GAN during the training process, making the obtained image recognition model more accurate, so that when the image recognition model is subsequently used to recognize the image to be recognized, accurate recognition and detection can be performed.
[0082] Furthermore, the process of extracting the GAN fingerprint of each synthetic image in S220 in the embodiment of the present invention specifically may include:
[0083] For each synthetic image, use an autoencoder to process the synthetic image to obtain a corresponding predicted image;
[0084] Calculate the reconstruction residual between the predicted image and the synthetic image;
[0085] Take the reconstruction residual as the GAN fingerprint corresponding to the synthetic image.
[0086] It should be noted that in the embodiments of the present invention, an autoencoder is established in advance, and during the process of extracting the GAN fingerprints of each synthetic image, each synthetic image is processed. Specifically, image processing can be performed on the synthetic image to obtain a predicted image corresponding to the synthetic image, that is, each synthetic image corresponds to a predicted image respectively, so as to obtain multiple groups of synthetic images and predicted images. For any group of synthetic images and predicted images, the reconstruction residual between the predicted image and the synthetic image is calculated, and this reconstruction residual is used as the GAN fingerprint of the corresponding synthetic image, thereby obtaining the GAN fingerprint corresponding to each synthetic image respectively.
[0087] Furthermore, the autoencoder in the embodiments of the present invention is trained based on a preset loss function, where the preset loss function can specifically be:
[0088] L = ||Fx predicted -F(x input )|| 2 , where F(x input ) represents the input feature, and Fx predicted represents the predicted result feature, and |||| 2 represents the distance L.
[0089] It should be noted that in the embodiments of the present invention, an autoencoder trained based on a preset loss function can be used to process the synthetic image to obtain a corresponding predicted image. Among them, the preset loss function can specifically be L = ||Fx predicted -F(x input )|| 2 . The purpose of using this preset loss function is to make the prediction result closer to the input result. For example, when using an autoencoder trained based on a preset loss function to process a synthetic image to obtain a corresponding predicted image, if the input synthetic image is image 0 and the output preset image is image 0', then the loss function L = ||Fx predicted -F(x input )|| 2 used in the embodiments of the present invention can make image 0' closer to image 0, so that the reconstruction residual obtained based on this predicted image and the corresponding synthetic image will be more accurate, and thus a more accurate GAN fingerprint can be obtained.
[0090] Next, the embodiments of the present invention will be described in detail taking two GAN networks as an example. Among them, the two GAN networks are GAN1 and GAN2 respectively. Specifically, please refer to Figure 3 :
[0091] Pre-train the GAN1 network and the GAN2 network to obtain GAN1 and GAN2, and then use the trained GAN1 and GAN2 to generate the required dataset. Specifically, for each GAN in GAN1 and GAN2, a large number of synthetic images can be generated respectively through GAN1 and GAN2. For example, a dataset 1 containing multiple synthetic images is generated through GAN1, and a dataset 2 containing multiple synthetic images is generated through GAN2, so as to obtain the dataset 1 corresponding to GAN1 and the dataset 2 corresponding to GAN2.
[0092] Specifically, after each synthetic image corresponding to GAN1 and GAN2 respectively, for each synthetic image of GAN1 and GAN2, extract the GAN fingerprint of each synthetic image, and add a corresponding label (specifically, it can be a GAN identifier) to the extracted GAN fingerprint. For example, for the image 1 generated by GAN1, the extracted fingerprint is fingerprint 1, and add the GAN1 label to this fingerprint 1, that is, indicate that this fingerprint 1 is the fingerprint of GAN1 through this label; for the image 2 generated by GAN2, the extracted fingerprint is fingerprint 2, and add the GAN2 label to this fingerprint 2, that is, indicate that this fingerprint 2 is the fingerprint of GAN2 through this label.
[0093] When performing fingerprint extraction, specifically, the synthetic image can be processed by an autoencoder trained based on a preset loss function to obtain a corresponding predicted image. Among them, the preset loss function can specifically be L = ||Fx predicted - F(x input )|| 2 . As Figure 3 shown, for the synthetic image 1 generated by GAN1, input the image 1 into the autoencoder (composed of an encoder and a decoder), and after the autoencoder processes the image 1, obtain the image 1'. Then, perform reconstruction residual calculation on the image 1 and the image 1' to obtain the reconstruction residual, and use this reconstruction residual as the fingerprint 1 of the image 1; for the synthetic image 2 generated by GAN2, input the image 2 into the autoencoder (composed of an encoder and a decoder), and after the autoencoder processes the image 2, obtain the image 2'. Then, perform reconstruction residual calculation on the image 2 and the image 2' to obtain the reconstruction residual, and use this reconstruction residual as the fingerprint 2 of the image 2. Of course, for other synthetic images generated by GAN1, the above method is used to extract fingerprints to obtain corresponding fingerprints. In this embodiment of the present invention, only one synthetic image of GAN1 and one synthetic image of GAN2 are used as examples for illustration.
[0094] Specifically, after obtaining the GAN fingerprints with the GAN1 identification tags added to each synthetic image of GAN1 and the GAN fingerprints with the GAN2 identification tags added to each synthetic image of GAN2, these GAN fingerprints with the GAN identification tags are used as the input of the classifier for classification training in the classifier, and after the training is completed, a trained image recognition model is obtained for subsequent use in image recognition of the image to be recognized.
[0095] When subsequently using the image recognition model trained in the embodiments of the present invention to recognize the image to be recognized, as Figure 4 shown, the image to be recognized is input into an autoencoder (encoder and decoder), and an image to be recognized' corresponding to the image to be recognized is obtained, and the reconstruction residual between the image to be recognized' and the image to be recognized is calculated. The obtained reconstruction residual is used as the fingerprint of the image to be recognized, and then this fingerprint is input into the trained image recognition model. The fingerprint is classified and recognized by the image recognition model, and the classification result is output, and then based on this classification result, it is determined whether the image to be recognized is a synthetic image.
[0096] Further, in practical applications, the classification result can specifically include three types, which can be a real image, an image containing a composite GAN fingerprint, and a single GAN image, respectively.
[0097] It should be noted that the artificial image detection methods in the prior art cannot recognize images with composite GAN fingerprints, resulting in low recognition accuracy. Among them, there are various reasons for the appearance of composite GAN fingerprints, such as artificial addition, etc. The method provided in the embodiments of the present invention can identify whether the image to be recognized is a real image, an image containing a composite GAN fingerprint, or a single GAN image through the trained image recognition model. Specifically, after the image recognition model is trained, the GAN fingerprint of each GAN can be determined. During the process of recognizing the image to be recognized, the GAN fingerprint corresponding to this fingerprint can be determined according to the fingerprint of the image to be recognized. When there is no GAN fingerprint corresponding to the image to be recognized, it means that the image to be recognized is a real image. When there are multiple GAN fingerprints corresponding to the image to be recognized, it means that the image to be recognized is an image containing a composite GAN fingerprint. When there is one GAN fingerprint corresponding to the image to be recognized, it means that the image to be recognized is a single GAN image.
[0098] Furthermore, the process of determining whether the image to be recognized is a synthetic image based on the classification result in S140 can specifically include:
[0099] When the classification result is a real image, the image to be recognized is a real image;
[0100] In the case where the classification result is an image containing a composite GAN fingerprint or a single GAN image, the image to be recognized is a synthetic image.
[0101] It can be understood that both the image containing the composite GAN fingerprint and the single GAN image are synthetic images. That is, as long as there is a GAN fingerprint corresponding to the image to be recognized, it indicates that the image to be recognized is generated by the GAN network. Therefore, the image to be recognized is a synthetic image. For the image to be recognized for which no corresponding GAN fingerprint is matched, it can be determined that the image to be recognized is a real image.
[0102] Furthermore, the method may further include:
[0103] In the case where the image to be recognized is a synthetic image, determine the GAN identifier corresponding to the synthetic image based on the classification result.
[0104] Specifically, in the case where it is determined that the image to be recognized is a synthetic image, it is also possible to further determine which GAN networks the image to be recognized is generated by according to the classification result corresponding to the image to be recognized, so as to determine the GAN identifier of the GAN network, which is beneficial to providing a basis for users to trace the generation of the image. For example, if the classification result of the image to be recognized is a single GAN image, and the label of the corresponding GAN fingerprint is GAN1, it indicates that the image to be recognized is generated by the GAN1 network. If the classification result of the image to be recognized is an image containing a composite GAN fingerprint, and the labels of the corresponding GAN fingerprints are GAN1 and GAN2 respectively, it indicates that the image to be recognized is associated with both GAN1 and GAN2.
[0105] It can be seen that in the embodiments of the present invention, an image recognition model is established in advance based on the GAN fingerprints of each synthetic image, then the image to be recognized is obtained and the fingerprint of the image to be recognized is extracted, and the image recognition model is used to classify and recognize the fingerprint to obtain the corresponding classification result. According to this classification result, it can be determined whether the image to be recognized is a synthetic image. The present invention can be applied to various types of GAN architectures, has a wide application range, high recognition accuracy, certain robustness, and is beneficial to improving the detection performance.
[0106] In addition, in practical applications, the above image recognition model and image recognition method can be applied to a recognition terminal. The recognition terminal may include a processor, an input component, and a display screen. The processor is respectively connected to the input component and the display screen. The processor can acquire the image to be recognized, extract the fingerprint of the image to be recognized, classify and recognize the fingerprint using a pre-established image recognition model, and obtain the classification result corresponding to the fingerprint. The image recognition model is established based on the GAN fingerprints of each synthetic image, determine whether the image to be recognized is a synthetic image based on the classification result, and display the recognition result through the display screen.
[0107] In practical applications, the input component may include an input interface and an input keyboard. The input interface can achieve connection with external devices, and the input keyboard can facilitate the user to input relevant instructions or data, etc., to the recognition terminal. To reduce the wiring difficulty and meet the data transmission requirements, a wireless transmission module may also be provided on the recognition terminal. Among them, the wireless transmission module may be a Bluetooth module or a wifi module, etc.
[0108] Figure 5 The figure is a schematic structural diagram of a recognition terminal provided by an embodiment of the present application. The recognition terminal may include a processor, a display screen 51, an input interface 52, an input keyboard 53, and a wireless transmission module 54. When the display screen 51 is a touch screen, the input keyboard 53 may be a soft keyboard presented on the display screen 51. The input interface 52 can be used to achieve connection with external devices. There may be multiple input interfaces. Figure 5 In the figure, one input interface is taken as an example. The processor is embedded inside the recognition terminal, so it is not shown in Figure 5 the figure.
[0109] The recognition terminal may be a smart phone, a tablet computer, a notebook computer, or a desktop computer, etc. In the embodiments of the present application, the form of the recognition terminal is not limited. When the question and answer terminal is a smart phone or a tablet computer, the input interface 52 can be connected to external devices through a data cable, and the input keyboard 33 may be a soft keyboard presented on the display interface. When the recognition terminal is a notebook computer or a desktop computer, the input interface 52 may be a USB interface for connecting external devices such as USB flash drives, and the input keyboard 53 may be a hard keyboard.
[0110] Taking a desktop computer as an example, in practical applications, the user can import the image to be recognized into a USB flash drive and insert the USB flash drive into the input interface 52 of the recognition terminal. After the recognition terminal acquires the image to be recognized, it can extract the fingerprint of the image to be recognized, classify and recognize the fingerprint using the image recognition model, obtain the classification result corresponding to the fingerprint, determine whether the image to be recognized is a synthetic image based on the classification result, and display the recognition result through the display screen 51. It should be noted that Figure 5The functional modules such as the display screen 51, input interface 52, input keyboard 53, and wireless transmission module 54 included in the recognition terminal are only examples. In actual applications, the Q&A terminal may also include more or fewer functional modules based on actual needs, and this is not limited.
[0111] The method for determining the text answer provided in the embodiments of the present application can be deployed in a software platform based on an FPGA (Field Programmable Gate Array) neural network acceleration application or an AI (Artificial Intelligence) acceleration chip. It should be noted that the method of compressing the neural network model according to the offset in the embodiments of the present application can be applied not only to the determination of text answers but also to the processing of time series data based on LSTM (Long Short-Term Memory), such as scenarios like multi-object tracking.
[0112] Based on the above embodiments, the embodiments of the present invention further provide an image recognition device. For example, please refer to Figure 6 , the device includes:
[0113] An acquisition module 61, configured to acquire an image to be recognized;
[0114] An extraction module 62, configured to extract the fingerprint of the image to be recognized;
[0115] A recognition module 63, configured to classify and recognize the fingerprint by using the image recognition model established in advance by the establishment module, and obtain a classification result corresponding to the fingerprint; the image recognition model is established based on the GAN fingerprints of each synthetic image;
[0116] A judgment module 64, configured to determine whether the image to be recognized is a synthetic image based on the classification result.
[0117] It should be noted that in the embodiments of the present invention, a large number of GAN fingerprints of artificial synthetic image samples can be used in advance by the establishment module to train an image recognition model, and the image recognition model is obtained, and the image to be recognized is recognized by the image recognition model. Specifically, when recognizing the image to be recognized, first, the acquisition module 61 acquires the fingerprint of the image to be recognized, and the acquisition module 61 sends the image to be recognized to the extraction module 62. The extraction module 62 extracts the fingerprint of the image to be recognized, obtains the fingerprint of the image to be recognized, and sends the fingerprint to the recognition module 63. After receiving the fingerprint of the image to be recognized, the recognition module 63 uses the fingerprint of the image to be recognized as the input of the image recognition model, classifies and recognizes the fingerprint through the image recognition model, and outputs the classification result corresponding to the fingerprint to the judgment module 24. The judgment module 64 further determines whether the image to be recognized is an artificial synthetic image according to the classification result.
[0118] As can be seen from the above, the image recognition device provided in the embodiments of the present invention can be applicable to various types of GAN architectures, has a wide application range, high recognition accuracy, certain robustness, and is beneficial to improving the detection performance.
[0119] Furthermore, the establishment module in the embodiments of the present invention may specifically include:
[0120] A generation unit, configured to generate multiple artificial synthetic images in advance for each generative adversarial network (GAN);
[0121] An extraction unit, configured to extract the GAN fingerprint of each artificial synthetic image and add a GAN identification label to the fingerprint;
[0122] A training unit, configured to input each GAN fingerprint added with the GAN identification label into a classifier and train the classifier to obtain an image recognition model.
[0123] It should be noted that in practical applications, when building an image recognition model, each GAN network can be pre-trained to obtain each GAN. Then, the generation unit uses each trained GAN to generate multiple synthetic images, that is, each GAN corresponds to a dataset (including multiple synthetic images), so as to obtain a large number of synthetic images for different GANs. Then, the extraction unit extracts the GAN fingerprint of each synthetic image for each GAN, and corresponding labels (specifically, GAN identifiers) need to be added to the GAN fingerprint. That is, the corresponding GAN identifier is used as a label and added to the GAN fingerprint corresponding to the synthetic image. For each synthetic image generated by each GAN network, the corresponding GAN fingerprint is extracted and the corresponding GAN identifier label is added, so as to obtain the GAN fingerprint with the GAN identifier label corresponding to each synthetic image. After obtaining the GAN fingerprint with the GAN identifier label corresponding to each synthetic image, these GAN fingerprints with the GAN identifier label are used as the sample set for training the image recognition model. Specifically, the training unit can input these GAN fingerprints into a pre-established classifier for classification training, and after the training is completed, a trained image recognition model is obtained for subsequent image recognition of the image to be recognized.
[0124] Further, the extraction unit in the embodiment of the present invention may specifically include:
[0125] A processing subunit, configured to process each synthetic image by using an autoencoder to obtain a corresponding predicted image;
[0126] A calculation subunit, configured to calculate the reconstruction residual between the predicted image and the synthetic image;
[0127] A determination subunit, configured to use the reconstruction residual as the GAN fingerprint corresponding to the synthetic image.
[0128] Further, the above-mentioned processing subunit may specifically be configured to process each synthetic image by using an autoencoder trained based on a preset loss function to obtain a corresponding predicted image;
[0129] Among them, the preset loss function in the embodiment of the present invention may specifically be:
[0130] L = ||Fx predicted - F(x input )|| 2 , where F(x input ) represents the input feature, Fx predicted represents the predicted result feature, and |||| 2 represents the distance L.
[0131] It should be noted that in the embodiment of the present invention, the processing subunit specifically uses an autoencoder trained based on a preset loss function to process the synthetic image to obtain a corresponding predicted image. Among them, the preset loss function can specifically be L = ||Fx predicted - F(x input )|| 2 . The purpose of using this preset loss function is to make the prediction result closer to the input result, so that the reconstruction residual obtained based on the predicted image and the corresponding synthetic image will be more accurate, and thus a more accurate GAN fingerprint can be obtained.
[0132] Furthermore, the classification results in the embodiment of the present invention include real images, images containing composite GAN fingerprints, and single GAN images.
[0133] Even further, the determination module 64 can specifically include:
[0134] The first determination unit is used to determine that the image to be recognized is a real image when the classification result is a real image;
[0135] The second determination unit is used to determine that the image to be recognized is a synthetic image when the classification result is an image containing a composite GAN fingerprint or a single GAN image.
[0136] That is, whether it is an image with a composite GAN fingerprint or a single GAN image, as long as the image to be recognized corresponds to a GAN fingerprint, it means that the image to be recognized is generated by the GAN network, and the image to be recognized is a synthetic image. For the image to be recognized that does not match the corresponding GAN fingerprint, it can be determined that the image to be recognized is a real image.
[0137] Even further, the device can also include:
[0138] The determination module is used to determine the GAN identifier corresponding to the synthetic image based on the classification result when the image to be recognized is a synthetic image.
[0139] Specifically, after determining that the image to be recognized is a synthetic image, the determination module can further extract the corresponding GAN identifier according to the classification result, so as to determine which GAN network generated the synthetic image, which is beneficial to providing a basis for users to trace the generation of the image.
[0140] Based on the above embodiments, the embodiment of the present invention also provides an image recognition system, specifically please refer to Figure 7 . The system includes:
[0141] The memory 70 is used to store computer programs;
[0142] A processor 71, which is configured to implement the steps of the image recognition method as described above when executing a computer program.
[0143] The image recognition system provided in this embodiment may specifically be an electronic device, including but not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer, etc.
[0144] Among them, the processor 71 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 71 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 71 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 71 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 71 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0145] The memory 70 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 70 may further include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 70 is at least used to store the following computer program 701. After the computer program is loaded and executed by the processor 71, it can implement the relevant steps of the image recognition method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 70 may further include an operating system 702 and data 703, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system 702 may include Windows, Unix, Linux, etc. The data 703 may include but is not limited to set offsets, etc.
[0146] In some embodiments, the electronic device may further include a display screen 72, an input / output interface 73, a communication interface 74, a power supply 75, and a communication bus 76.
[0147] Those skilled in the art can understand that Figure 7 the structure shown in Figure 7 does not constitute a limitation on the electronic device, and it may include more or fewer components than those shown in the figure.
[0148] It can be understood that if the method for determining the text answer in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in various embodiments of this application. The aforementioned storage media include: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), electrically erasable programmable ROMs, registers, hard disks, removable disks, CD-ROMs, magnetic disks, or optical disks, etc., which can store program codes.
[0149] Based on this, the embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the image recognition method as described above.
[0150] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0151] It should also be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0152] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0153] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0154] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image recognition method, characterized in that, it includes: Obtain the image to be recognized; Extract the fingerprint of the image to be recognized; Use a pre-established image recognition model to classify and recognize the fingerprint to obtain the classification result corresponding to the fingerprint; The image recognition model is established based on the GAN fingerprints of each synthetic image; Determine whether the image to be recognized is a synthetic image based on the classification result; where: The classification result includes a real image, an image containing a composite GAN fingerprint, and a single GAN image; The step of using a pre-established image recognition model to classify and recognize the fingerprint to obtain the classification result corresponding to the fingerprint includes: Use a pre-established image recognition model to classify and recognize the fingerprint. When there is no GAN fingerprint corresponding to the image to be recognized, determine that the image to be recognized is a real image; when there are multiple GAN fingerprints corresponding to the image to be recognized, determine that the image to be recognized is an image containing a composite GAN fingerprint; when there is one GAN fingerprint corresponding to the image to be recognized, determine that the image to be recognized is a single GAN image; The step of determining whether the image to be recognized is a synthetic image based on the classification result includes: If the classification result is an image containing a composite GAN fingerprint or a single GAN image, the image to be recognized is a synthetic image; If the image to be recognized is a synthetic image, determine the GAN identifier corresponding to the synthetic image based on the classification result; If the classification result of the image to be recognized is a single GAN image and the GAN identifier label of the corresponding GAN fingerprint is GAN1, then the image to be recognized is generated by the GAN1 network; if the classification result of the image to be recognized is an image containing a composite GAN fingerprint and the GAN identifier labels of the corresponding GAN fingerprints are GAN1 and GAN2 respectively, then the image to be recognized is associated with both the GAN1 network and the GAN2 network.
2. The image recognition method according to claim 1, characterized in that, The establishment of the image recognition model includes: For each generative adversarial network GAN, use the GAN to generate multiple synthetic images in advance; Extract the GAN fingerprint of each synthetic image and add a GAN identifier label to the fingerprint; Input the GAN fingerprints with added GAN identifier labels into a classifier and train the classifier to obtain an image recognition model.
3. The image recognition method according to claim 1, characterized in that, The step of extracting the GAN fingerprint of each synthetic image includes: For each synthetic image, use an autoencoder to process the synthetic image to obtain a corresponding predicted image; Calculate the reconstruction residual between the predicted image and the synthetic image; Use the reconstruction residual as the GAN fingerprint corresponding to the synthetic image.
4. The image recognition method according to claim 3, characterized in that, The autoencoder is trained based on a preset loss function, and the preset loss function is: L = ||Fx predicted -F(x input )|| 2 , where F(x input ) represents the input feature, Fx predicted represents the predicted result feature, and || || 2 represents the distance L.
5. The image recognition method according to claim 1, wherein, determining whether the image to be recognized is a synthetic image based on the classification result further includes: when the classification result is a real image, the image to be recognized is a real image.
6. An image recognition device, wherein, it includes: an acquisition module for acquiring an image to be recognized; an extraction module for extracting the fingerprint of the image to be recognized; a recognition module for classifying and recognizing the fingerprint by using an image recognition model established in advance by a building module to obtain a classification result corresponding to the fingerprint; the image recognition model is established based on the GAN fingerprints of each synthetic image; a judgment module for determining whether the image to be recognized is a synthetic image based on the classification result; wherein: the classification result includes a real image, an image containing a composite GAN fingerprint, and a single GAN image; the recognition module is specifically configured to classify and recognize the fingerprint by using a pre-established image recognition model, and when there is no GAN fingerprint corresponding to the image to be recognized, determine that the image to be recognized is a real image; when there are multiple GAN fingerprints corresponding to the image to be recognized, determine that the image to be recognized is an image containing a composite GAN fingerprint; when there is one GAN fingerprint corresponding to the image to be recognized, determine that the image to be recognized is a single GAN image; the judgment module includes: a second judgment unit for, when the classification result is an image containing a composite GAN fingerprint or a single GAN image, determining that the image to be recognized is a synthetic image; a determination module for, when the image to be recognized is a synthetic image, determining a GAN identifier corresponding to the synthetic image based on the classification result; if the classification result of the image to be recognized is a single GAN image and the GAN identifier label of the corresponding GAN fingerprint is GAN1, then the image to be recognized is generated by the GAN1 network; if the classification result of the image to be recognized is an image containing a composite GAN fingerprint and the GAN identifier labels of the corresponding GAN fingerprints are GAN1 and GAN2 respectively, then the image to be recognized is associated with both the GAN1 network and the GAN2 network.
7. An image recognition system, wherein, it includes: a memory for storing a computer program; a processor for implementing the steps of the image recognition method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the image recognition method according to any one of claims 1 to 5 are implemented.