Image detection method, electronic equipment and storage medium

By combining visual and textual features in a multimodal model, the problem of existing technologies being unable to identify diverse image forgery methods is solved, achieving more efficient image authenticity detection and improving the accuracy and comprehensiveness of detection.

CN120808123APending Publication Date: 2025-10-17BEIJING GESHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510646232.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing image authenticity detection methods are unable to effectively identify diverse and unknown image forgery methods, especially for forged images based on methods other than generative adversarial networks and diffusion models, where the accuracy is insufficient.

Method used

By obtaining the visual features and text features of the image, using the multimodal visual language model to extract the visual features, and combining them with the text features representing the real image and the forged image, the authenticity probability and forgery probability of the image are calculated to comprehensively judge the authenticity of the image.

Benefits of technology

The accuracy and comprehensiveness of image authenticity detection are improved, and forged images produced by various forgery methods can be more accurately identified, thereby enhancing information security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808123A_ABST
    Figure CN120808123A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image detection method, electronic equipment and a storage medium. The method comprises the following steps: acquiring a first visual feature of a to-be-detected image; based on the first visual feature and the first text feature, determining a first probability that the to-be-detected image is identified as a real image, and based on the first visual feature and the second text feature, determining a second probability that the to-be-detected image is identified as a forged image; the first text feature is a text feature used for representing that the image is a description text of a real image, and the second text feature is a text feature used for representing that the image is a description text of a forged image; and determining an authenticity detection result of the to-be-detected image according to the first probability and the second probability. By comprehensively considering the visual features of the image and the general first text feature and second text feature, the forged image generated by various forging methods can be identified more accurately, and the accuracy, practicability and comprehensiveness of image authenticity detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing. More particularly, it relates to an image detection method, an electronic device and a storage medium. BACKGROUND

[0002] With the development of science and technology, image generation and processing technology has brought much convenience to people's production and life; but at the same time, it has also brought new challenges to information security. Criminals can use various image generation techniques and physical presentation means (such as using masks or printed images) to forge images required for identity verification, thereby threatening information security. For example, using image generation techniques can create fake face images, thereby deceiving face recognition systems. Therefore, detecting the authenticity of images is a key measure to ensure information security.

[0003] However, there are many types of image forgery methods. In related technologies, there are some methods for detecting the authenticity of images. However, these methods can generally only identify specific types of forged images, such as being able to detect images generated by image generation methods based on Generative Adversarial Network (GAN) or Diffusion Model, but cannot accurately determine the authenticity of images generated by other methods. In summary, existing image authenticity detection methods cannot cope with diversified or even unknown image forgery methods. SUMMARY

[0004] The present application is proposed in consideration of the above problems.

[0005] According to an aspect of the present application, an image detection method is provided, comprising:

[0006] obtaining a to-be-detected image;

[0007] obtaining a first visual feature of the to-be-detected image;

[0008] determining a first probability that the to-be-detected image is identified as a real image based on the first visual feature and a first text feature, and determining a second probability that the to-be-detected image is identified as a forged image based on the first visual feature and a second text feature; wherein the first text feature refers to a text feature of a first text, the first text refers to a description text for characterizing that the image is a real image, and the second text feature refers to a text feature of a second text, the second text refers to a description text for characterizing that the image is a forged image;

[0009] According to still another aspect of the present application, an electronic device is provided, comprising a processor and a memory, the memory having stored therein computer program instructions which, when executed by the processor, are configured to perform the image detection method as described above.

[0010] According to a further aspect of the present application, there is provided a storage medium having stored thereon program instructions which, when executed by a processor, are configured to perform the image detection method as described above.

[0011] In the technical solution described above, the first visual feature of the to-be-detected image is obtained, the first probability of the to-be-detected image being identified as a real image is determined based on the first visual feature and the first text feature, and the second probability of the to-be-detected image being identified as a fake image is determined based on the first visual feature and the second text feature, so as to determine the authenticity detection result of the to-be-detected image. In this way, the first visual feature of the to-be-detected image is compared with the first text feature and the second text feature respectively, so that the visual feature of a real image is closer to the first text feature in the feature space, thereby enabling the features of real images to be clustered into one category based on the technical solution, and thus separated from the features of other fake images. By comprehensively considering the visual feature of the image and the first text feature and the second text feature, rather than focusing on the semantic feature of the image itself, fake images generated by various fake methods can be more accurately identified, and the accuracy, practicability and comprehensiveness of image authenticity detection are improved.

[0012] The above description is only a summary of the technical solutions of the present application. In order to enable the technical solutions of the present application to be more clearly understood, and to be implemented according to the contents of the description, and in order to enable the above and other purposes, features and advantages of the present application to be more apparent, the specific embodiments of the present application are described below. BRIEF DESCRIPTION OF DRAWINGS

[0013] The above and other purposes, features and advantages of the present application will become more apparent from the following detailed description of the embodiments of the present application, taken in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, but do not constitute a limitation of the present application.

[0014] Figure 1 A schematic flowchart of an image detection method according to an embodiment of the present application is shown;

[0015] Figure 2 A schematic block diagram of an image detection method according to an embodiment of the present application is shown;

[0016] Figure 3 A schematic flowchart of a training process of a visual feature adaptive module according to an embodiment of the present application is shown;

[0017] Figure 4 A schematic flowchart of a training process of a text feature adaptive module according to an embodiment of the present application is shown;

[0018] Figure 5 A schematic flow chart showing a co-training process of the visual adaptation module and the text adaptation module according to an embodiment of the present application is shown;

[0019] Figure 6 A schematic block diagram showing a training process of the visual feature adaptation module and the text feature adaptation module according to an embodiment of the present application is shown;

[0020] Figure 7 A schematic block diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0021] It should be noted that the image data obtained by the scheme of the present application is accessed, collected, stored and applied to subsequent analysis and processing under the condition that the user or the relevant data owner agrees and authorizes after being informed of the data collection content, data use, processing method and other information of the data, and the user or the relevant data owner can be provided with a way to access, correct and delete the data, as well as a method to revoke the consent and authorization.

[0022] In order to make the purpose, technical scheme and advantages of the present application more obvious, the example embodiments according to the present application will be described in detail below with reference to the drawings. Obviously, the described embodiments are only part of the embodiments of the present application.

[0023] Image authenticity is used to represent whether an image is collected for a target object or is fake. If an image is collected for a target object, the image is real; if an image is generated by AI or is tampered, etc., the image is fake. Detecting image authenticity has important significance for preventing the spread of false information, protecting information security and ensuring the accuracy of identity verification. The existing image detection method usually trains a model for detection using real images and the same kind of fake images, and the obtained model is only more sensitive to the visual features of this kind of fake image. Thus, the model has higher accuracy only when detecting this kind of fake image, and it is difficult to obtain ideal detection results for other fake images.

[0024] In order to at least solve the above technical problems, the present application provides an image detection method. In the detection method, in addition to using the visual features of the image itself, two different text features respectively representing whether the image is real or not are also used to detect the authenticity of the image. Thus, it is avoided to learn too many semantic features and ignore the learning of the differences between real images and fake images. Figure 1 A schematic flow chart showing an image detection method according to an embodiment of the present application is shown. As shown in Figure 1 The image detection method includes steps S1100, S1200, S1300 and S1400.

[0025] At step S1100, an image to be tested is acquired.

[0026] At step S1200, a first visual feature of the image to be tested is acquired.

[0027] At step S1300, a first probability that the image to be tested is identified as a real image is determined based on the first visual feature and a first text feature, and a second probability that the image to be tested is identified as a fake image is determined based on the first visual feature and a second text feature; the first text feature refers to a text feature of a first text, the first text refers to a description text for characterizing that the image is a real image, and the second text feature refers to a text feature of a second text, the second text refers to a description text for characterizing that the image is a fake image.

[0028] At step S1400, a result of the authenticity detection of the image to be tested is determined according to the first probability and the second probability.

[0029] Exemplarily, at step S1100, an image to be tested is acquired. The image to be tested can be any image that needs to be tested for authenticity. For example, an image that needs to be subjected to face recognition can be taken as the image to be tested for image authenticity detection. If the detection result indicates that the image to be tested is a fake image, it can be directly determined that it fails to pass face recognition. This can improve information security and avoid fake images passing face recognition. The image to be tested can also be an image provided by a user according to user demand. Thus, it can be determined whether the image to be tested is a fake image, helping the user to identify false information. The image to be tested can include an image in any format, for example, the image to be tested can be in a Joint Photographic Experts Group (JPEG), Portable Network Graphics (PNG), Raw Image Format (RAW), Scalable Vector Graphics (SVG), or the like. In some embodiments, the image detection method can be performed by a terminal device. The image to be tested can include an image photographed or stored by the terminal device, and the terminal device can perform subsequent steps to detect the image to be tested. In other embodiments, the image detection method can be performed by a background server. The terminal device can send the photographed or stored image to be tested to the background server. After the background server acquires the image to be tested, it performs subsequent steps to detect the image to be tested.

[0030] At step S1200, a first visual feature of the image to be tested is acquired.

[0031] The first visual feature of the to-be-tested image can be extracted by using any existing or future developed method for extracting features of an image. In some embodiments, the first visual feature of the to-be-tested image can be extracted by using a deep learning model. The deep learning model can be, for example, a multi-modal visual language model. The multi-modal visual language model includes a visual encoding part and a text encoding part. The first visual feature of the to-be-tested image can be extracted by using the visual encoding part in the multi-modal visual language model. The multi-modal visual language model can be an existing or future developed multi-modal visual language model such as a Contrastive Language-Image Pre-Training (CLIP), a Fine-grained Interactive Language-Image Pre-train (FILIP), a Foundational Language And Vision Alignment Model (FLAVA), etc. The multi-modal visual language model can be trained, i.e., trained by using a large amount of different types of images and corresponding labeled data. In other embodiments, the first visual feature of the to-be-tested image can be extracted by using technologies such as Convolutional Neural Networks (CNN), Vision Transformer (ViT), etc. The above-mentioned models for extracting the first visual feature can all be models trained by using a large amount of data. In other words, the first visual feature of the to-be-tested image can be extracted by using a trained model. Thus, the trained model can not be limited to the visual features of a specific type of fake image, and can extract more generalized features in the to-be-tested image as the first visual feature. This can avoid the problem in the related art that the model is over-fitted to the inherent features of the data set due to training by using a specific fake technology data set, has low generalization, has low accuracy in real-time detection of images outside the data set, and cannot cope with various fake technologies.

[0032] In step S1300, a first probability that the to-be-tested image is identified as a real image is determined based on the first visual feature and a first text feature, and a second probability that the to-be-tested image is identified as a fake image is determined based on the first visual feature and a second text feature; wherein the first text feature refers to the text feature of the first text, the first text refers to a description text for characterizing that the image is a real image, and the second text feature refers to the text feature of the second text, the second text refers to a description text for characterizing that the image is a fake image.

[0033] The first text feature can be a feature obtained by feature extraction on the description text representing that the image is a real image. Thus, the first text feature also represents that the image is real. Similarly, the second text feature can be a feature obtained by feature extraction on the description text representing that the image is a fake image. Thus, the second text feature also represents that the image is fake.

[0034] In some embodiments, the text encoding part in the multi-modal visual language module can be employed to extract a feature vector of the first text representing that the image is a real image as the first text feature, and extract a feature vector of the second text representing that the image is a fake image as the second text feature. For example, the first text representing that the image is a real image can include “real”, “true”, “authentic”, “real”, etc., and the second text representing that the image is a fake image can include “fake”, “forgery”, “fake”, etc. Any one of “real-fake”, “true-fake”, “authentic-forgery”, “real-fake”, etc. can be input into the text encoding part of the multi-modal visual language model as different texts representing real-fake, respectively. The first text feature representing that the image is a real image and the second text feature representing that the image is a fake image can be obtained. Preferably, the corresponding first text feature and second text feature can be invariant when performing image authenticity detection. Thus, the first text feature and the second text feature can be pre-set. The first text feature and the second text feature can be used as fixed parameters to avoid multiple extractions of text features, reduce the amount of calculation, and improve the detection speed. In other embodiments, the first text feature and the second text feature of the above-mentioned set of texts representing real-fake can be extracted by using a bag-of-words model, word embedding, topic model, etc.

[0035] Exemplarily, the first text feature and the second text feature can be pre-stored to an electronic device for performing the above-mentioned detection method before performing the image detection method. Alternatively, the above-mentioned first text feature and second text feature can be obtained via a network before or simultaneously with performing the image detection method. Alternatively, the first text feature and the second text feature can be extracted from the above-mentioned first text and second text, respectively, as described above when performing the image detection method.

[0036] The first visual feature, the first text feature and the second text feature can be feature vectors with the same dimension, for example, dimension [n, 1]. In some embodiments, a similarity between the first visual feature of the to-be-tested image and the first text feature can be calculated as a first probability that the to-be-tested image is identified as a real image. And a similarity between the first visual feature and the second text feature can be calculated as a second probability that the to-be-tested image is identified as a fake image. The similarity can be calculated by cosine similarity or other suitable metric methods, such as Euclidean distance. It can be understood that the higher the similarity between the first visual feature and the second text feature, the greater the first probability value that the to-be-tested image is identified as a real image, and similarly, the higher the similarity between the first visual feature and the second text feature, the greater the second probability value that the to-be-tested image is identified as a fake image. In other embodiments, the similarity between the first visual feature of the to-be-tested image and the first text feature and the second text feature can be calculated respectively to obtain two similarity values, and the two similarity values can be input into a classification layer to obtain the first probability and the second probability. The classification layer can be implemented by using a Softmax layer, a logistic layer, a fully connected layer, etc.

[0037] In step S1200, the visual feature of the image is extracted by using the model. In such embodiments, the model for extracting the first visual feature of the image can be obtained by training using the training image, the authenticity label data of the training image, the first text feature and the second text feature.

[0038] In step S1400, the authenticity detection result of the to-be-tested image is determined according to the first probability and the second probability.

[0039] In some embodiments, the first probability and the second probability are compared, and if the first probability is greater than the second probability, the to-be-tested image is determined to be a real image; if the second probability is greater than the first probability, the to-be-tested image is determined to be a fake image. For the case where the first probability is equal to the second probability value, the detection result of the to-be-tested image can be invalid detection. The to-be-tested image with invalid detection can be processed, and the first visual feature is extracted again, and the first probability and the second probability are determined again to determine the authenticity detection result. In other embodiments, the authenticity detection result of the to-be-tested image can be determined by comparing the first probability with the first probability threshold value and comparing the second probability with the second probability threshold value. For example, the first probability threshold value can be 0.7, and the second probability threshold value can be 0.3. If the first probability is greater than the first probability threshold value and the second probability is less than the second probability threshold value, the to-be-tested image is a real image, otherwise the to-be-tested image is a fake image. The first probability threshold value and the second probability threshold value can be set according to user requirements.

[0040] In the technical solution, the first visual feature of the to-be-tested image is obtained, the first probability that the to-be-tested image is identified as a real image is determined based on the first visual feature and the first text feature, and the second probability that the to-be-tested image is identified as a fake image is determined based on the first visual feature and the second text feature, so as to determine the authenticity detection result of the to-be-tested image. Thus, the first visual feature of the to-be-tested image is compared with the first text feature and the second text feature respectively, so that the visual features of real images are closer to the first text feature in the feature space, thereby enabling the features of real images to be classified into one category and separated from the features of other fake images based on the technical solution. By comprehensively considering the visual features of the image and the first text feature and the second text feature commonly used, rather than focusing on the semantic features of the image itself, fake images generated by various fake methods can be more accurately identified, and the accuracy, practicability and comprehensiveness of image authenticity detection are improved.

[0041] For example, the to-be-tested image is a face image containing a face region. In step S1400, the authenticity detection result of the to-be-tested image is determined according to the first probability and the second probability, including determining the authenticity detection result of the face image according to the first probability and the second probability, wherein the authenticity detection result is used to indicate whether the face region contained in the face image is a real face. The face image at least contains a face region of a person. For example, the authenticity detection result of the face image indicates whether the face region in the face image is a real face; if the authenticity detection result of the face image indicates that the face region contained in the face image is a real face, it can be determined that the face image is a real one taken and can be subjected to subsequent operations, such as face recognition, informing the user of the detection result, etc. If the authenticity detection result of the face image indicates that the face region contained in the face image is not a real face, i.e., a fake face, the user can be informed of the detection result or determined as an attack behavior, etc. The face image can include an image obtained in a face recognition scenario and needs to be subjected to authenticity detection. If the authenticity detection result of the face image indicates that the face is a real face, face recognition can be performed, and if the authenticity detection result of the face image indicates that the face is a fake face, it can be determined as an attack behavior. The face image can also include a face image provided by a user, and the image authenticity detection can determine whether the face is real or fake.

[0042] In the technical solution, the authenticity detection result of the face image is determined according to the first probability and the second probability, and the authenticity detection result is used to indicate whether the face region contained in the face image is a real face. Thus, the authenticity of the face can be determined, which provides a basis for subsequent tasks and improves user experience.

[0043] For example, step S1300 includes step S1310, step S1320 and step S1330.

[0044] At step S1310, a first similarity between the first visual feature and the first text feature is calculated. The first visual feature, the first text feature and the second text feature can be feature vectors with the same dimension, for example, dimension [n, 1]. The similarity between the first visual feature and the first text feature of the image to be tested can be calculated as the first similarity. The similarity can be calculated by cosine similarity or other suitable metric method (such as Euclidean distance).

[0045] At step S1320, a second similarity between the first visual feature and the second text feature is calculated. The same similarity calculation method as in step S1310 can be used to calculate the second similarity between the first visual feature and the second text feature.

[0046] At step S1330, the first similarity is normalized by using a normalized exponential function to obtain a first probability, and the second similarity is normalized by using a normalized exponential function to obtain a second probability. For example, the first probability and the second probability can be calculated according to the following formula:

[0047]

[0048] wherein x can be a feature vector including the first similarity and the second similarity; W k may be a weight matrix; W j may represent the jthcolumn of the weight matrix W k ; j can include 1 or 2; P(1) can represent the first probability, i.e., the probability that the image to be tested is a real image; P(2) can represent the second probability, i.e., the probability that the image to be tested is a fake image.

[0049] The first probability that the image to be tested is a real image and the second probability that the image to be tested is a fake image can be calculated by using a normalized exponential function according to the first similarity and the second similarity. The sum of the first probability and the second probability is 1. For example, the second probability that the image to be tested is a fake image can be S, and the first probability that the image to be tested is a real image can be 1-S. The closer S is to 0, the more real the image to be tested is; on the contrary, the closer S is to 1, the more likely the image to be tested is fake.

[0050] Step S1400 includes step S1410. In step S1410, if the first probability is greater than or equal to the second probability, the authenticity detection result of the to-be-tested image indicates that the to-be-tested image is a real image, and if the first probability is less than the second probability, the authenticity detection result of the to-be-tested image indicates that the to-be-tested image is a fake image. Taking the second probability of the to-be-tested image being a fake image as S and the first probability of the to-be-tested image being a real image as 1-S as an example, the authenticity detection result of the to-be-tested image can be determined by comparing the first probability with the second probability. If the first probability is greater than or equal to the second probability, the authenticity detection result of the to-be-tested image indicates that the to-be-tested image is a real image, and if the first probability is less than the second probability, the authenticity detection result of the to-be-tested image indicates that the to-be-tested image is a fake image. In other words, if the first probability of the to-be-tested image is greater than or equal to 0.5, the to-be-tested image is a real image, and if the first probability is less than 0.5, the to-be-tested image is a fake image.

[0051] In the technical solution, the first visual feature is calculated to have the first similarity and the second similarity with the first text feature and the second text feature respectively, the first similarity and the second similarity are normalized by using the normalization exponential function to determine the first probability and the second probability, and the authenticity detection result of the to-be-tested image is determined according to the size of the first probability and the second probability. Thus, whether the to-be-tested image is a real image or a fake image can be more accurately measured, and the normalized data can also be used for subsequent analysis and processing.

[0052] In an embodiment, the first similarity between the first visual feature and the first text feature, and the second similarity between the first visual feature and the second text feature can be calculated based on a cosine distance algorithm. The first visual feature, the first text feature and the second text feature can all be represented as feature vectors. The first similarity and the second similarity can be represented by calculating the cosine value of the included angle between the first visual feature and the first text feature and the second text feature respectively. The numerical range of the first similarity and the second similarity is between [-1, 1]. The closer the numerical value is to 1, the higher the similarity. For example, matrix multiplication can be used to simplify the process of calculating the cosine distance. The dimensions of the first visual feature, the first text feature and the second text feature can all be [n, 1]. When calculating the first similarity and the second similarity, the first text feature and the second text feature can be spliced to obtain a [n, 2] dimensional matrix. In order to calculate the similarity, the first text feature can be transposed to [1, n] dimension and multiplied with the spliced [n, 2] matrix to obtain a [1, 2] result matrix to obtain the first similarity and the second similarity at the same time. The first similarity can be used as the first probability, and the second similarity can be used as the second probability. The first similarity and the second similarity can also be input into a classification layer to obtain the first probability and the second probability, and then determine the image authenticity detection result. The classification layer can use a Softmax layer, a logistic layer, a fully connected layer, etc.

[0053] The above technical solution calculates the first similarity between the first visual feature and the first text feature and the second similarity between the first visual feature and the second text feature based on the cosine distance algorithm. The technical method of cosine similarity is simple, saves computing resources, and is convenient for subsequent determination of the image authenticity detection result of the to-be-detected image.

[0054] Exemplarily, step S1200 obtains the first visual feature of the to-be-detected image, including step S1210 and step S1220.

[0055] In step S1210, the initial visual feature of the to-be-detected image is extracted by using a visual encoder. The initial visual feature of the to-be-detected image can be extracted by using the visual encoder of the multi-modal visual language model described in step S1200. For brevity, it will not be repeated here.

[0056] In step S1220, the initial visual feature is target adjusted by using the trained visual feature adaptive module to obtain the first visual feature. The target adjustment includes visual feature level adjustment and / or visual feature region of interest adjustment.

[0057] The trained visual feature adaptive module can be trained by using training data sets. Each training data set includes a training image, authenticity label data of the training image, a first training text feature representing authenticity of the training image, and a second training text feature representing forgery of the training image.

[0058] It can be understood that the visual encoder in the multi-modal visual language model can have been pre-trained by a large amount of data, and thus can be directly used to extract visual features of the to-be-tested image. However, the visual encoder described above can be a general feature extraction model, which is not specifically designed for the image authenticity detection task. It can be understood that the initial visual features are the final output features of the to-be-tested image processed by the visual encoder, which pay more attention to the global of the to-be-tested image and pay more attention to the semantic content of the image. Therefore, in order to improve the accuracy of image authenticity detection, the trained visual feature adaptive module can be used to target adjust the initial visual features extracted by the visual encoder. The target adjustment can be performed by visual feature level adjustment, reducing the semantic content in the initial visual features, increasing the image low-level features, and increasing the features related to the detection of image authenticity. The semantic content can interfere with the identification of forgery traces, and reducing the semantic content in the initial visual features can further improve the accuracy of image authenticity detection. The visual feature level adjustment can include downgrading the semantic content related features to obtain more image low-level features, such as more texture and color features related to the to-be-tested image. The target adjustment can be performed by visual feature region of interest adjustment, changing the region of interest of the initial visual features to obtain more local detail features. For example, for a face image, the region of interest of the initial visual features can be the entire image, and by visual feature region of interest adjustment, the region of interest can be adjusted to the region where the face is located in the image to obtain more detail features related to the face. The visual feature region of interest adjustment can further improve the accuracy of image authenticity detection. The trained visual feature adaptive module can perform visual feature level adjustment and visual feature region of interest adjustment on the initial visual features to improve the accuracy of image authenticity detection. The processed initial visual features are used as the first visual features in step S1300 described above for determining the first probability and the second probability. The trained visual feature adaptive module can be implemented by using techniques such as fully connected layers and one-dimensional convolution to target adjust the initial visual features. The trained visual feature adaptive module can not change its mathematical structure in the process of processing the initial visual features. The first visual features and the initial visual features can have the same dimension. The target adjustment of the initial visual features by the trained visual feature adaptive module can make the obtained first visual features more suitable for performing the image authenticity detection task

[0059] The visual feature self-adaption module can be trained to make the first visual feature more suitable for performing the authenticity detection task of the image. The visual feature self-adaption module can be trained by using a training data set to obtain a trained visual feature self-adaption module. Each training data set can include a training image, authenticity label data of the training image, a first training text feature representing authenticity of the training image, and a second training text feature representing forgery of the training image. For example, the first training text feature and the second training text feature can be fixed, i.e., the first training text feature and the second training text feature remain unchanged in all training data sets. In some embodiments, the authenticity of the training image can be detected by using the visual encoder and the visual feature self-adaption module to obtain an authenticity detection result of the training image. According to the authenticity detection result of the training image and the authenticity label data of the training image, only the parameters of the visual feature self-adaption module are adjusted by using a loss function to obtain the trained visual feature self-adaption module. In other embodiments, according to the authenticity detection result of the training image and the authenticity label data of the training image, the parameters of the visual feature self-adaption module and the visual encoder are adjusted together by using a loss function. The training image can include an image generated by any one or more forgery techniques. The training image is used to fine-tune the model for extracting the first visual feature. Since the model is trained by using the first training text feature and the second training text feature, even if the training image contains a small number of image forgery techniques, the overall generalization ability of the model will not be affected.

[0060] In the above technical solution, the initial visual feature of the image to be detected is extracted by using the visual encoder, and the initial visual feature is target-adjusted by using the trained visual feature self-adaption module to obtain the first visual feature. The target adjustment includes visual feature level adjustment and / or visual feature region of interest adjustment. Thus, the first visual feature can be more suitable for performing the authenticity detection task of the image, thereby improving the detection accuracy.

[0061] The first text feature and the second text feature are obtained by the following process: the first text feature is obtained by adjusting the first text encoding corresponding to the first text by using the trained text feature self-adaption module, and then extracting features from the adjusted first text encoding by using the text encoder; and the second text feature is obtained by adjusting the second text encoding corresponding to the second text by using the trained text feature self-adaption module, and then extracting features from the adjusted second text encoding by using the text encoder, wherein the encoding adjustment includes adjusting the semantic content corresponding to the text encoding.

[0062] For example, the first text can be "real". First, "real" can be converted into a text encoding vector through text encoding to obtain a first text encoding corresponding to the first text, so as to facilitate the text encoder to recognize and extract features. The semantic content of the first text encoding corresponding to the first text may not be particularly accurate in describing that the image is a real image. The semantic content corresponding to the first text encoding can represent the meaning of the first text, that is, the meaning of the first text may not be particularly accurate in describing the meaning of the image being a real image. Therefore, the first text encoding of the first text can be first adjusted by using the trained text feature adaptive module. The adjusted first text encoding can more accurately describe that the image is real. Finally, the text encoder is used to extract features from the text encoding vector of the processed first text to obtain the first text feature. Further, the first text feature can more accurately describe that the image is real. For example, adjusting the semantic content corresponding to the text encoding can include inserting a trained code in the first text encoding, so that the semantic content corresponding to the text encoding more accurately describes that the image is real. Adjusting the semantic content corresponding to the text encoding can also include adjusting the numerical value of the first text encoding, so that the first text encoding can more accurately describe that the image is real.

[0063] Similarly, the second text can be "fake", and the second text encoding corresponding to the second text can be adjusted by using the trained text feature adaptive module and the text encoder in turn, and then the adjusted second text encoding is extracted to obtain the second text feature, so that it can more accurately describe that the image is fake.

[0064] The trained text feature self-adaption module is trained by using training data sets. Each training data set includes a training image, authenticity label data of the training image, a first text, and a second text. It can be understood that, by using the text feature self-adaption module, the first text feature of the obtained first text can be more representative of the common features of real images, and the second text feature can be more representative of the common features of fake images. The text feature self-adaption module can be implemented by using a fully connected layer or a one-dimensional convolutional layer, etc. The parameters of the text feature self-adaption module can be changed through the training process of the text feature self-adaption module. The training data sets used in the training process can include training images, authenticity label data of the training images, first texts, and second texts. The first texts and the second texts in different training data sets can be the same. In some embodiments, the training data sets can be used to train only the text feature self-adaption module. Specifically, based on the authenticity detection results of the training images and the authenticity label data, only the parameters of the text feature self-adaption module are adjusted until a preset condition for ending the training process is reached. In other embodiments, the training data sets can be used to train, based on the authenticity detection results of the training images and the authenticity label data, the parameters of the text feature self-adaption module and the text encoder together until a preset condition for ending the training process is reached.

[0065] In the above technical solution, the first text feature is obtained by sequentially using the trained text feature self-adaption module and the text encoder to encode and adjust the first text encoding corresponding to the first text and then performing feature extraction, and the second text feature is obtained by sequentially using the trained text feature self-adaption module and the text encoder to sequentially encode and adjust the second text encoding corresponding to the second text and then performing feature extraction. The application of the text feature self-adaption module can make the first text feature and the second text feature more representative of the common features of real images and fake images, thereby improving the accuracy of image authenticity detection.

[0066] In the above technical solution, the joint application of the trained visual feature self-adaption module and the text feature self-adaption module makes the obtained first visual feature more suitable for performing the image authenticity detection task, and the obtained first text feature and second text feature are more representative of the common features of real images and fake images, thereby improving the accuracy of detection.

[0067] Exemplarily, Figure 2A schematic block diagram of an image detection method according to an embodiment of the present application is shown. In the image authenticity detection of a to-be-detected image, initial visual features of the to-be-detected image can be extracted by a visual encoder, and the initial visual features are processed by a visual feature adaptive module to obtain first visual features. The visual feature adaptive module can be trained. Similarities of the first visual features with first text features and second text features are calculated respectively, and first and second probabilities are calculated according to the similarities to determine an authenticity detection result of the to-be-detected image.

[0068] Exemplarily, step S1300 includes steps S1340a and S1350 of determining the first probability that the to-be-detected image is identified as a real image based on the first visual features and the first text features.

[0069] In step S1340a, the first visual features are weighted and summed with the initial visual features to obtain third visual features. The initial visual features and the first visual features can have the same dimension, and thus can be weighted and summed. In an embodiment, the weights of the two can be set according to the user's requirements. In another embodiment, the weights of the two can be dynamically adjusted according to the fitting effect when the visual feature adaptive module is trained. For example, if the visual feature adaptive module cannot converge well when it is trained, the weight of the second visual features can be increased so that the initial visual features have more influence on the third visual features. Conversely, the weight of the first visual features can be increased.

[0070] In step S1350, the similarity of the third visual features with the first text features is calculated as the first probability. The third visual features and the first text features can be feature vectors with the same dimension. The similarity of the third visual features and the first text features can be calculated as the first probability by using methods such as cosine similarity, Euclidean distance, etc.

[0071] Step S1300 determines a second probability that the to-be-tested image is identified as a fake image based on the first visual feature and the second text feature, including step S1340b and step S1360. In step S1340b, the first visual feature is weighted and summed with the initial visual feature to obtain a fourth visual feature. In step S1360, the similarity between the fourth visual feature and the second text feature is calculated as the second probability. It can be understood that when calculating the second probability, the fourth visual feature obtained by weighting and summing the first visual feature and the initial visual feature can be used for calculation. In some embodiments, the fourth visual feature obtained in step S1340b can be the same as the third visual feature obtained in step S1340a, that is, the same weight is used in step S1340a and step S1340b to weight and sum the first visual feature and the initial visual feature. In other embodiments, step S1340b can use a different weight than step S1340a to weight and sum the first visual feature and the initial visual feature. Step S1360 and step S1350 can use the same similarity calculation method. The similarity between the fourth visual feature and the second text feature is calculated as the second probability. The first similarity and the second similarity are ensured to be comparable. Alternatively, the similarity between the third visual feature and the first text feature, and the similarity between the fourth visual feature and the second text feature, can be input into a classification layer to process the similarity values to obtain the first probability and the second probability.

[0072] In the above technical solution, the first visual feature is weighted and summed with the initial visual feature to obtain a third visual feature, the similarity between the third visual feature and the first text feature and the second text feature is calculated as the first probability, the first visual feature is weighted and summed with the initial visual feature to obtain a fourth visual feature, and the similarity between the fourth visual feature and the second text feature is calculated as the second probability. In this way, overfitting of the visual feature adaptive module can be avoided, and the detection accuracy can be improved.

[0073] Exemplarily, Figure 3 A schematic flowchart of a training process of a visual feature adaptive module according to an embodiment of the present application is shown. As Figure 3 As shown, the visual feature adaptive module is obtained through the training process of steps S3100 to S3500.

[0074] In step S3100, an initial sample visual feature of a training image is extracted using a visual encoder. The visual encoder can be the visual encoder in the multi-modal visual language model described in step S1200. The initial sample visual feature of each training image can be extracted using the visual encoder.

[0075] At step S3200, the initial sample visual features are target adjusted by an initial visual feature adaptive module to obtain second visual features, wherein the target adjustment includes visual feature level adjustment and / or visual feature region of interest adjustment. The initial visual feature adaptive module can be untrained and can have initial parameters. The initial sample visual features of each training image can be target adjusted by the initial visual feature adaptive module, and the adjusted initial sample visual features are taken as the second visual features. The specific processing manner is similar to the method described in step S1220, and details are not described herein for brevity.

[0076] At step S3300, a third probability that the training image is identified as a real image is determined based on the second visual features and the first training text features, and a fourth probability that the training image is identified as a fake image is determined based on the second visual features and the second training text features. The method for determining the third probability and the fourth probability is similar to the determination method in step S1300 described above, and the similarity between the second visual features and the first training text features and the similarity between the second visual features and the second training text features can be calculated based on cosine similarity, Euclidean distance, etc., and then the first probability and the second probability are determined. The similarity calculation method and the probability determination method can be selected according to user demand or training effect.

[0077] At step S3400, the authenticity detection result of the training image is determined based on the third probability and the fourth probability. Similar to the determination of the authenticity detection result of the image to be detected in step S1400 described above, the authenticity detection result of the training image can be determined by comparing the sizes of the third probability and the fourth probability. If the third probability is large, the authenticity detection result indicates that the training image is a real image, and if the fourth probability is large, the authenticity detection result indicates that the training image is a fake image. The authenticity detection result of the image to be detected can also be determined by comparing the sizes of the third probability and a third probability threshold value and comparing the sizes of the fourth probability and a fourth probability threshold value. For example, the third probability threshold value can be 0.7, and the fourth probability threshold value can be 0.3. If the third probability is greater than the third probability threshold value and the fourth probability is less than the fourth probability threshold value, the image to be detected is a real image, otherwise the image to be detected is a fake image.

[0078] At step S3500, the parameters of the initial visual feature adaptive module are adjusted based on the authenticity detection result of the training image and the authenticity annotation data of the training image to obtain a trained visual feature adaptive module.

[0079] It can be understood that the authenticity detection result includes two cases, i.e. the image is a real image or a fake image. The detection of the authenticity of the image can be regarded as a binary classification task. The authenticity label data of the training image is used to label the actual authenticity of the training image, for example, whether it is real or fake. The cross-entropy loss function can be used to adjust the parameters of the initial visual feature adaptive module based on the difference between the authenticity detection result of the training image and the authenticity label data to obtain the trained visual feature adaptive module.

[0080] Exemplarily, the above step S3500 can be repeated multiple times. Each time, the parameters of the visual feature adaptive module are adjusted according to the authenticity detection result of the training image and the authenticity label data of the training image until a preset condition is reached to obtain the trained visual feature adaptive module. The preset condition can be that the number of repetitions of step S3500 reaches a number threshold, the loss function value reaches a numerical threshold, and the like.

[0081] In the above training process, only the parameters of the visual feature adaptive module are changed, and the parameters of the visual encoder remain unchanged. In other words, in this embodiment, the extraction of the visual feature is only affected by changing the parameters of the visual feature adaptive module. The visual encoder can be a model that has been pre-trained on a large amount of image data. In the above training process, the parameters of the visual encoder are not further adjusted, and its good generalization performance can be maintained. Through the adjustment of the parameters of the visual feature adaptive module, the second visual feature is more suitable for the image authenticity detection task.

[0082] The above technical solution adjusts the parameters of the initial visual feature adaptive module based on the authenticity detection result of the training image and the authenticity label data of the training image, and combines the first training text feature and the second training text feature to obtain the trained visual feature adaptive module. In the training process, the parameters of the visual encoder remain unchanged. Thus, while maintaining the generalization performance of the visual encoder, the accuracy of image authenticity detection is improved.

[0083] Exemplarily, Figure 4 A schematic flowchart of a training process of a text feature adaptive module according to an embodiment of the present application is shown. As shown in Figure 4 The text feature adaptive module is obtained through the training process of steps S4100 to S4600.

[0084] In step S4100, the initial sample visual feature of the training image is obtained. Similar to step S1200, the initial sample visual feature of the training image can be extracted by any method. For example, the initial sample visual feature of the training image can be extracted by the visual encoding part of the multi-modal visual language model, and can also be extracted by a convolutional neural network and the like.

[0085] At step S4200, the first text encoding corresponding to the first text is adjusted by the initial text feature adaptive module to obtain a third training text encoding, and the second text encoding corresponding to the second text is adjusted by the initial text feature adaptive module to obtain a fourth training text encoding. The adjustment of the encoding includes adjusting the semantic content corresponding to the text encoding.

[0086] The initial text feature adaptive module can be implemented by a fully connected layer or a convolutional layer, etc. The initial text feature adaptive module can have preset parameters. The parameters of the initial text feature adaptive module can be adjusted during the training process. For example, the first text can be "real" to describe that the image is real. The first text encoding after text encoding of "real" can be adjusted by the initial text feature adaptive module to obtain the third training text encoding. Similarly, the second text can be "fake" to describe that the image is fake. The second text encoding after text encoding of "fake" can be adjusted by the initial text feature adaptive module to obtain the fourth training text encoding.

[0087] At step S4300, the third training text encoding is feature extracted by the text encoder to obtain a first training text feature, and the fourth training text encoding is feature extracted by the text encoder to obtain a second training text feature.

[0088] The third training text encoding is obtained by adjusting the first text encoding corresponding to the first text by the text feature adaptive module. The first training text feature is obtained by feature extracting the third training text encoding by the text encoder, which can make the first training text feature better represent the common features of real images. Similarly, the fourth training text encoding is obtained by adjusting the second text encoding corresponding to the second text by the text feature adaptive module. The second training text feature is obtained by feature extracting the fourth training text encoding by the text encoder, which can make the second training text feature better represent the common features of fake images.

[0089] At step S4400, a fifth probability that the training image is identified as a real image is determined based on the initial sample visual feature and the first training text feature, and a sixth probability that the training image is identified as a fake image is determined based on the initial sample visual feature and the second training text feature.

[0090] At step S4500, the authenticity detection result of the training image is determined based on the fifth probability and the sixth probability.

[0091] The step S4400 and the step S4500 are similar to the step S1300 and the step S1400 respectively, and thus are not described herein again for simplicity.

[0092] In the step S4600, the parameters of the initial text feature adaptive module are adjusted based on the authenticity detection result of the training image and the authenticity label data of the training image, to obtain a trained text feature adaptive module.

[0093] In the training process, the parameters of the text encoder remain unchanged, and only the parameters of the text feature adaptive module are changed. The authenticity label data of the training image can indicate that the training image should be determined as a real image or a fake image. Based on the authenticity detection result of the training image and the authenticity label data, the parameters of the text feature adaptive module can be adjusted multiple times by using an arbitrary loss function, such as a cross-entropy loss function, until a preset condition is reached, to obtain a trained text feature adaptive module. The preset condition can be that the number of repetitions in the step S4600 reaches a number threshold, the value of the loss function reaches a value threshold, and the like. It can be understood that, in the training process, only the parameters of the initial text feature adaptive module are adjusted, and the parameters of the text encoder are not adjusted, i.e., the parameters of the text encoder remain unchanged. The text encoder can be a general text encoder, which can be a model pre-trained by using a large amount of text data.

[0094] In the above technical solution, based on the authenticity detection result of the training image and the authenticity label data, the parameters of the initial text feature adaptive module are adjusted to obtain a trained text feature adaptive module, and in the training process, the parameters of the text encoder remain unchanged. Thus, the generalization ability of the text encoder can be ensured, and the text features can be made to be more representative of the common features of real or fake images, thereby improving the accuracy of image authenticity detection.

[0095] Exemplarily, Figure 5 A schematic flowchart of a common training process of the visual adaptive module and the text adaptive module according to an embodiment of the present application is shown. As shown in Figure 5 The visual adaptive module and the text adaptive module can be trained together through the steps S3100, S3200, S4200, S4300, S3300, S3400 and S5100.

[0096] At step S5100: based on the authenticity detection result of the training image and the authenticity annotation data of the training image, the parameters of the initial visual feature adaptive module and the initial text feature adaptive module are adjusted to obtain a trained visual feature adaptive module and a trained text feature adaptive module. Wherein, in the training process, the parameters of the visual encoder and the parameters of the text encoder remain unchanged. While adjusting the parameters of the initial text feature adaptive module, the parameters of the initial visual feature adaptive module are also adjusted to obtain the trained text feature adaptive module and the trained visual feature adaptive module after the training is completed.

[0097] In the above technical solution, in the training process of the model, the parameters of the text feature adaptive module and the visual feature adaptive module are adjusted at the same time, and the parameters of the text encoder and the visual encoder remain unchanged. Therefore, the generalization ability of both the text encoder and the visual encoder can be guaranteed, the detection ability of images generated by different forgery methods can be improved, and the extracted visual features and text features are more suitable for image authenticity detection tasks, and the accuracy of image authenticity detection is improved.

[0098] Exemplarily, Figure 6 A schematic block diagram of the training process of the visual feature adaptive module and the text feature adaptive module according to an embodiment of the present application is shown. As Figure 6As shown, the training image is input into the visual encoder. With the visual encoder, initial sample visual features of the training image are extracted. The initial sample visual features are input into the visual feature adaptation module. With the visual feature adaptation module, the initial sample visual features are processed to obtain second visual features. With the text feature adaptation module, the first text encoding corresponding to the first text and the second text encoding corresponding to the second text are processed respectively to obtain corresponding third training text encoding and fourth training text encoding. The first text represents a real image, and the second text represents a fake image. With the text encoder, the third training text encoding and the fourth training text encoding are respectively subjected to feature extraction, and finally the first training text features representing the first text of the real image and the second training text features representing the second text of the fake image are obtained. The similarity of the second visual features with the first training text features and the second training text features is calculated respectively, and the probability that the training image is a real image and the probability that the image is a fake image are determined to determine the authenticity detection result. According to the authenticity detection result and the authenticity annotation data of the training image, the parameters of the text feature adaptation module and the visual feature adaptation module can be adjusted to realize the simultaneous training of the two. It can be understood that after the training is completed, the parameters of the text feature adaptation module and the text encoder can be used for image authenticity detection of the to-be-tested image. The first training text features and the second training text features obtained after the training can be used as the first text features and the second text features for image authenticity detection of the to-be-tested image.

[0099] Exemplarily, according to another aspect of the present application, an image authenticity detection device is also provided. The device comprises a processor. The processor is configured to execute the image detection method of any of the above embodiments.

[0100] Exemplarily, according to still another aspect of the present application, an electronic device is also provided. Figure 7 A schematic block diagram of an electronic device 700 according to an embodiment of the present application is shown. The electronic device 700 comprises a processor 710 and a memory 720. The memory 720 stores computer program instructions, which, when executed by the processor 710, are configured to execute the image detection method as described above.

[0101] Exemplarily, according to still another aspect of the present application, a storage medium is also provided, which stores program instructions, which, when executed, are configured to execute the image detection method as described above. The storage medium may, for example, include an erasable programmable read-only memory (EPROM), a compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The storage medium can be any combination of one or more computer readable storage media.

[0102] Exemplarily, according to yet another aspect of the present application, there is also provided a computer program product comprising computer program instructions for performing the image detection method as described above when executed.

[0103] Those skilled in the art can understand the specific implementation schemes and beneficial effects of the image detection apparatus, the electronic device, the storage medium and the computer program product by reading the above description about the image detection method, and for brevity, will not be described here again.

[0104] Although the example embodiments have been described with reference to the accompanying drawings, it is to be understood that the example embodiments are only exemplary and are not intended to limit the scope of the present application. Those skilled in the art can make various changes and modifications without departing from the scope and spirit of the present application. All these changes and modifications are intended to be included within the scope of the present application as claimed in the appended claims.

[0105] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or in a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0106] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the above-described device embodiments are merely illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be omitted or not executed.

[0107] In the specification provided herein, a large number of specific details are described. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some examples, well-known methods, structures and techniques are not shown in detail in order not to obscure the understanding of the present specification.

[0108] Similarly, it is to be understood that the application can be positioned and claimed in a manner wider than that which is described under a single embodiment, figure or description thereof and that the claims can represent all the different combinations and permutations of the various features described herein. Accordingly, the claims are not to be seen as being limited to the specific embodiments described under the general description of the application or under a section entitled "Detailed description of the embodiments".

[0109] Those skilled in the art will appreciate that all features described in this specification (including the summary of the application, abstract and claims) together with those implied by the description and drawings can be provided in any combination thereof, to provide any of the features described herein or any other technical solution described herein. Each feature disclosed in this specification (including the summary of the application, abstract and claims) can be replaced by alternative features serving the same, equivalent or similar purpose, unless expressly stated otherwise.

[0110] Furthermore, those skilled in the art will recognize that references in the specification to "one embodiment", "an embodiment", "an example embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily referring to a single embodiment. Furthermore, the terms "comprises", "comprising", "includes", "including", "has", "having" or variants thereof are used synonymously with each other in this disclosure and are intended to mean that the feature or features in question can be present or included, but do not exclude the potential presence or inclusion of one or more additional features.

[0111] Various component embodiments of the application can be implemented in hardware, or as software modules running in one or more processors, or in combinations thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some of the modules in the image detection method apparatus according to embodiments of the application. The application can also be implemented as a program (for example, a computer program and computer program product) for executing any or all of the methods described herein on a computer system. Such program(s) of the present application can be stored on a computer readable medium which can be any medium, tangible or intangible, in which the program can be stored and / or executed including a storage medium. Such storage medium can have patterns of holes, pits or other patterning on it which follow a patterned data layout, or it can be in the form of a signal which can have patterns of ones and zeros which follow a patterned data layout. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0112] It should be noted that the above-mentioned embodiments illustrate rather than limit the application, and that one skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps other than those listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of both hardware and software, and any combination thereof. In a unitary claim, several devices or sub-claims can be joined by means of the expression "and / or". The use of the term "at least" followed by a list of one or more items should be interpreted as including at least one of the items but it does not exclude the presence of others not listed. The use of the term "one" followed by a list of one or more items should be interpreted as including at least one of the items but it does not exclude the presence of others not listed. It is emphasized that the terms "comprises / comprising" when used in this specification are taken to specify the presence of stated features, integers, steps or components but do not preclude the presence or addition of one or more other features, integers, steps, components or groups thereof.

[0113] The above description is only specific embodiments of the present application or specific explanations of specific embodiments, and the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, and all of them should be covered in the protection scope of the present application. The protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image detection method, characterized in that: include: Acquire the image to be tested; Acquiring a first visual feature of the image to be measured; Determining a first probability that the image to be tested is identified as a genuine image based on the first visual feature and the first text feature, and determining a second probability that the image to be tested is identified as a forged image based on the first visual feature and the second text feature; wherein the first text feature refers to a text feature of a first text, the first text refers to a descriptive text used to characterize the image as a genuine image, and the second text feature refers to a text feature of a second text, the second text refers to a descriptive text used to characterize the image as a forged image; An authenticity detection result of the image to be tested is determined according to the first probability and the second probability.

2. The image detection method according to claim 1, wherein: The obtaining of the first visual feature of the image to be measured includes: Extracting initial visual features of the image to be tested using a visual encoder; Using the trained visual feature adaptation module, performing target adjustment on the initial visual feature to obtain the first visual feature, wherein the target adjustment includes visual feature level adjustment and / or visual feature region of interest adjustment; and / or The first text feature and the second text feature are obtained by the following process: The first text feature is obtained by using a trained text feature adaptive module to perform encoding adjustment on a first text encoding corresponding to the first text, and then using a text encoder to perform feature extraction on the adjusted first text encoding; the second text feature is obtained by using the trained text feature adaptive module to perform encoding adjustment on a second text encoding corresponding to the second text, and then using the text encoder to perform feature extraction on the adjusted second text encoding, wherein the encoding adjustment includes adjusting the semantic content corresponding to the text encoding.

3. The image detection method according to claim 2, wherein: The visual feature adaptive module is trained through the following training process: extracting initial sample visual features of a training image using the visual encoder; Using the initial visual feature adaptation module, target adjustment is performed on the initial sample visual feature to obtain a second visual feature, wherein the target adjustment includes visual feature level adjustment and / or visual feature region of interest adjustment; determining a third probability that the training image is identified as a real image based on the second visual feature and the first training text feature, and determining a fourth probability that the training image is identified as a forged image based on the second visual feature and the second training text feature; determining a authenticity detection result of the training image based on the third probability and the fourth probability; Adjusting parameters of the initial visual feature adaptive module based on the authenticity detection result of the training image and the authenticity annotation data of the training image to obtain a trained visual feature adaptive module; During the training process, the parameters of the visual encoder remain unchanged.

4. The image detection method according to claim 2, wherein: Determining a first probability that the image to be tested is recognized as a real image based on the first visual feature and the first text feature includes: Performing a weighted summation of the first visual feature and the initial visual feature to obtain a third visual feature; Calculating a similarity between the third visual feature and the first text feature as the first probability; Determining a second probability that the image to be tested is identified as a forged image based on the first visual feature and the second text feature includes: Perform a weighted summation of the first visual feature and the initial visual feature to obtain a fourth visual feature: The similarity between the fourth visual feature and the second text feature is calculated as the second probability.

5. The image detection method according to claim 2, wherein: The text feature adaptive module is trained through the following training process: Obtaining initial sample visual features of training images; Using an initial text feature adaptive module to perform encoding adjustment on a first text encoding corresponding to the first text to obtain a third training text encoding, and using the initial text feature adaptive module to perform encoding adjustment on a second text encoding corresponding to the second text to obtain a fourth training text encoding, wherein the encoding adjustment includes adjusting semantic content corresponding to the text encoding; Performing feature extraction on the third training text encoding using the text encoder to obtain first training text features, and performing feature extraction on the fourth training text encoding using the text encoder to obtain second training text features; Determining a fifth probability that the training image is identified as a real image based on the initial sample visual feature and the first training text feature, and determining a sixth probability that the training image is identified as a forged image based on the initial sample visual feature and the second training text feature; determining a authenticity detection result of the training image based on the fifth probability and the sixth probability; Adjusting parameters of the initial text feature adaptive module based on the authenticity detection result of the training image and the authenticity annotation data of the training image to obtain a trained text feature adaptive module; During the training process, the parameters of the text encoder remain unchanged.

6. The image detection method according to claim 2, wherein: The visual feature adaptive module and the text feature adaptive module are trained through the following training process: extracting initial sample visual features of a training image using the visual encoder; Using the initial visual feature adaptation module, target adjustment is performed on the initial sample visual feature to obtain a second visual feature, wherein the target adjustment includes visual feature level adjustment and / or visual feature region of interest adjustment: Using an initial text feature adaptive module to perform encoding adjustment on a first text encoding corresponding to the first text to obtain a third training text encoding, and using the initial text feature adaptive module to perform encoding adjustment on a second text encoding corresponding to the second text to obtain a fourth training text encoding, wherein the encoding adjustment includes adjusting semantic content corresponding to the text encoding; Performing feature extraction on the third training text encoding using the text encoder to obtain first training text features, and performing feature extraction on the fourth training text encoding using the text encoder to obtain second training text features; determining a third probability that the training image is identified as a real image based on the second visual feature and the first training text feature, and determining a fourth probability that the training image is identified as a forged image based on the second visual feature and the second training text feature; determining a authenticity detection result of the training image based on the third probability and the fourth probability; Adjusting parameters of the initial visual feature adaptation module and the initial text feature adaptation module based on the authenticity detection result of the training image and the authenticity annotation data of the training image to obtain a trained visual feature adaptation module and a trained text feature adaptation module; During the training process, the parameters of the visual encoder and the parameters of the text encoder remain unchanged.

7. The image detection method according to claim 1, wherein: Determining a first probability that the image to be tested is identified as a real image based on the first visual feature and the first text feature, and determining a second probability that the image to be tested is identified as a forged image based on the first visual feature and the second text feature, includes: Calculating a first similarity between the first visual feature and the first text feature; calculating a second similarity between the first visual feature and the second text feature; The first similarity is normalized using a normalized exponential function to obtain the first probability, and the second similarity is normalized using a normalized exponential function to obtain the second probability: Determining the authenticity detection result of the image to be tested according to the first probability and the second probability includes: If the first probability is greater than or equal to the second probability, the authenticity detection result of the image to be tested indicates that the image to be tested is a real image; if the first probability is less than the second probability, the authenticity detection result of the image to be tested indicates that the image to be tested is a forged image.

8. The image detection method according to any one of claims 1 to 7, characterized in that: The image to be tested is a face image containing a face area; Determining an authenticity detection result of the image to be tested according to the first probability and the second probability includes: An authenticity detection result of the facial image is determined based on the first probability and the second probability, wherein the authenticity detection result is used to indicate whether the facial region included in the facial image is a real face.

9. An electronic device comprising: A processor and a memory, characterized in that The memory stores computer program instructions, which are used by the processor to execute the image detection method according to any one of claims 1 to 8 when the processor is running the computer program instructions.

10. A storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, they are used to execute the image detection method according to any one of claims 1 to 8.