Image authenticity detection method and device, storage medium and electronic equipment
By using pre-stored real text features and authenticity recognition models of fake text features in deep fake face detection, the problems of weak generalization ability and high training time cost in the prior art are solved, and high accuracy and fast response image authenticity detection is achieved.
Patent Information
- Application Number
- CN202311553842.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
The existing technology has weak generalization ability in deep fake face detection, resulting in low detection accuracy, and model training requires a lot of time and cannot meet the needs of emergency online.
The authenticity and false recognition model that pre-stores real text features and forged text features is adopted, and the images to be detected are input to the model to obtain authenticity and false detection results. The generalization ability of the model is improved through fine-grained semantic information, and the authenticity and false detection of zero samples is performed without training.
It improves the accuracy of image authenticity detection, improves the generalization ability of the model, and can meet the needs of deep forgery detection functions in emergencies.
Smart Images

Figure CN120020912A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, this application relates to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for detecting the authenticity of images. Background Art
[0002] The rapid development of deepfake face technology has brought about entertainment and convenience while also posing huge security risks.
[0003] Existing detection methods use digital labels as supervision information, where the digital label indicates whether the face in the corresponding sample image is real or fake. The training model classifies images based on the labels, but the generalization ability of the model trained based on the existing method is weak, resulting in a low detection accuracy. More importantly, since the training of the model requires a large amount of time cost, when an application needs to urgently launch a deepfake detection function, the related technology cannot meet this requirement. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for detecting the authenticity of images, which can solve the above problems of the existing technology. The technical solutions are as follows:
[0005] According to the first aspect of the embodiments of this application, a method for detecting the authenticity of images is provided. The method includes:
[0006] Input the image to be detected into the authenticity recognition model, and obtain the authenticity detection result of the image to be detected output by the authenticity recognition model;
[0007] Wherein, the authenticity recognition model pre-stores real text features and forged text features. The real text features are determined according to each real description information in a pre-determined description information set, and the forged text features are determined according to each forged description information in the description information set;
[0008] Any one of the real description information and the forged description information is used to describe an image and is obtained through alignment processing with the image. The description information includes a status field, an object field, and a template field. The status field is used to indicate the authenticity of the object shown in the image, the object field is used to describe the category of the object shown in the image, the template field is used to describe the scene type shown in the corresponding image according to the target description method. The object field is embedded in the status field, and the status field is embedded in the template field;
[0009] The real description information status field is used to indicate that the object shown in the corresponding image is real, and the status field of the forged description information is used to indicate that the object shown in the corresponding image is forged.
[0010] According to a second aspect of the embodiments of the present application, there is provided an image authenticity detection device, including:
[0011] A model processing module, configured to input the image to be detected into the authenticity recognition model, and obtain the authenticity detection result of the image to be detected output by the authenticity recognition model;
[0012] Wherein, the authenticity recognition model pre-stores real text features and forged text features. The real text features are determined according to each real description information in a pre-determined description information set, and the forged text features are determined according to each forged description information in the description information set;
[0013] Any one of the real description information and the forged description information is used to describe an image, and is obtained through alignment processing with the image. The description information includes a status field, an object field, and a template field. The status field is used to indicate the authenticity of the object shown in the image, the object field is used to describe the category of the object shown in the image, the template field is used to describe the scene type of the corresponding image shown in a target description manner, the object field is embedded in the status field, and the status field is embedded in the template field;
[0014] The real description information status field is used to indicate that the object shown in the corresponding image is real, and the status field of the forged description information is used to indicate that the object shown in the corresponding image is forged.
[0015] In some optional embodiments, the model processing module is specifically configured to:
[0016] Input the image to be detected into the authenticity recognition model to obtain the image features of the image to be detected;
[0017] Determine the similarities between the image features and the pre-determined real text features and forged text features respectively, and determine the authenticity detection result according to the magnitude relationship between the two similarities.
[0018] In some optional embodiments, the authenticity recognition model is a contrastive language-image pre-training CLIP model.
[0019] In some optional embodiments, the image authenticity detection device further includes a description information set acquisition module:
[0020] A field set acquisition unit for obtaining a field set, where the field set includes multiple status fields, multiple object fields, and multiple template fields, and randomly generating multiple description messages according to the field set;
[0021] A classification unit for classifying the multiple description messages according to the template fields in each description message to obtain multiple description message clusters, and each description message cluster includes description messages with the same scenario type;
[0022] An index calculation unit for, for each description message cluster, determining the text features of each description message in the description message cluster, and for each description message, determining the average value of the similarity of the text features of the description message to the text features of other description messages with the same status field, and the average value of the similarity of the text features of the description message to the text features of other description messages with different status fields, and taking the difference between the two average values as the index value of the description message;
[0023] A sorting and screening unit for, for each description message cluster, sorting the description messages in the description message cluster from largest to smallest according to the index values of the description messages, and retaining a preset number of description messages with the top sorting results to form a description message set.
[0024] In some alternative embodiments, the authenticity recognition model is used to recognize the authenticity of a face image, the image to be detected is a face image, and the real description message further includes facial feature fields for describing the facial features in the corresponding image.
[0025] In some alternative embodiments, the authenticity recognition model is used to recognize the authenticity of a face image, the image to be detected is a face image, and the object field describes that the category of the object shown in the corresponding image includes non-face.
[0026] According to the third aspect of the embodiments of the present application, an electronic device is provided, and the electronic device includes a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method provided in the first aspect above.
[0027] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method provided in the first aspect above are implemented.
[0028] According to the fifth aspect of the embodiments of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the method provided in the first aspect above are implemented.
[0029] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are:
[0030] By inputting the image to be detected into the authenticity recognition model, the authenticity detection result of the image to be detected output by the authenticity recognition model is obtained. The authenticity recognition model of the embodiment of the present application pre-stores real text features and forged text features. The real text features are determined according to each real description information in the pre-determined description information set, and the forged text features are determined according to each forged description information in the description information set. It can be seen that the real text features and the forged text features can accurately describe the authenticity of the object shown in the image corresponding to a description information (i.e., text) from a semantic perspective. And all description information in the present application includes a status field, an object field, and a template field. The object field is used to describe the category of the object shown in the corresponding image, and the template field is used to describe the scene type of the corresponding image display according to the target description method; the status field of the real description information is used to describe that the object shown in the corresponding image is real, and the status field of the forged description information is used to describe that the object shown in the corresponding image is forged, so that the authenticity recognition model learns fine-grained semantic information, improves the generalization ability of the model, and further improves the accuracy of image authenticity detection. Moreover, for the problem that the zero-shot classification ability of the CLIP model is poor because the two categories of real and forged in the deepfake detection task cannot be clearly defined by text, the embodiment of the present application designs a large number of description texts related to the authenticity detection task to make up for the above problems of the CLIP model. And because there are various types of target objects in the pre-training data of the CLIP model, and there is a good alignment between images and texts in the feature space, therefore, when the embodiment of the present application uses the CLIP model as the authenticity recognition model, zero-shot authenticity detection can be performed without training, and no additional training is required for the image authenticity recognition scenario. Based on this, the embodiment of the present application can meet the requirement of urgently launching the deepfake detection function. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description in the embodiments of the present application.
[0032] Figure 1 It is a schematic diagram of the system architecture for implementing the method for detecting the authenticity of an image provided by an embodiment of the present application;
[0033] Figure 2 It is a schematic flowchart of a method for detecting the authenticity of an image provided by an embodiment of the present application;
[0034] Figure 3 It is a schematic flowchart of a method for detecting the authenticity of an image provided by another embodiment of the present application;
[0035] Figure 4 It is a schematic flowchart of a method for detecting the authenticity of an image provided by an embodiment of the present application;
[0036] Figure 5 A schematic flowchart of a method for detecting the authenticity of an image provided in another embodiment of this application;
[0037] Figure 6 An application schematic diagram of a video review scenario provided in an embodiment of this application;
[0038] Figure 7 A schematic structural diagram of a device for detecting the authenticity of an image provided in an embodiment of this application;
[0039] Figure 8 A schematic structural diagram of an electronic device provided in an embodiment of this application. Detailed implementation manners
[0040] The embodiments of this application will be described below with reference to the accompanying drawings in this application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute limitations on the technical solutions of the embodiments of this application.
[0041] Those skilled in the art of this technology can understand that unless specifically stated otherwise, the singular forms "a", "an", and "the" used here may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of this application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components, and / or their combinations supported by this technical field. It should be understood that when we say an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to the other element, or it can mean that this element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by this term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0042] To make the purpose, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below in conjunction with the accompanying drawings.
[0043] First, several terms related to this application will be introduced and explained:
[0044] The Intelligent Traffic System (ITS), also known as the Intelligent Transportation System, effectively integrates advanced scientific and technological means (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing, strengthening the connection among vehicles, roads, and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, enhances the environment, and saves energy.
[0045] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0046] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, and mechatronics. Artificial intelligence technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0047] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as target recognition and measurement in machine vision, and further performing image processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0048] Deepfake technology is an artificial intelligence technology that uses a machine learning model called a generative adversarial network to merge and superimpose pictures or videos onto the source pictures or videos, and uses neural network technology for large-sample learning to splice and synthesize false content such as personal voices, facial expressions, and body movements.
[0049] The most common way of deepfake is AI face-swapping technology. In addition, it also includes voice simulation, face synthesis, video generation, etc. Its emergence makes it possible to tamper with or generate highly realistic and difficult-to-discriminate audio and video content, and ultimately observers cannot distinguish the authenticity with the naked eye.
[0050] Existing detection methods for deepfake images are mainly divided into two categories: simulating forgery traces and network structure engineering. The former enhances the generalization forgery features that are easily ignored by the network by artificially creating forgery samples, and the latter focuses on specific network structure designs to improve the generalization ability of the model. For example, using high-frequency data streams to reduce texture bias, using contrastive learning to maintain consistency between forgery instances, or using unsupervised methods to learn mask labels to learn forgery traces.
[0051] Through analysis, the embodiments of this application found that the above methods all use digital labels and ignore the fine-grained semantic information, and the semantic information of these texts can help the model improve its generalization ability. Therefore, the inventive concept of the embodiments of this application is to introduce fine-grained semantic information into the training process of the model, thereby improving the generalization ability of the model.
[0052] The Contrastive Language-Image Pre-Training (CLIP) model is a pre-trained neural network model for matching images and texts released by OpenAI at the beginning of 2021. The CLIP model demonstrates powerful zero-shot classification capabilities in various classification tasks. However, through analysis, the embodiments of this application found that compared with other classification tasks, in the deepfake detection task, the two categories of real and forged cannot be clearly defined using text, and due to the inability to fully utilize the knowledge in the image-text feature space aligned by the CLIP model, the zero-shot classification ability of the CLIP model is poor.
[0053] That is to say, on the one hand, although the existing deepfake detection methods using digital labels have achieved certain good performance, since the supervision information of digital labels does not contain semantic information, the model cannot learn the true meanings of real and forged; on the other hand, although the CLIP model has powerful zero-shot classification capabilities in various classification tasks, due to the fact that the objects of deepfake pictures are all human faces and there are forgery traces only in the local areas of the human face, the categories of real and forged cannot be clearly defined using text, resulting in the CLIP model not being directly applicable to deepfake detection.
[0054] The image authenticity detection method, device, electronic device, computer-readable storage medium, and computer program product provided by this application aim to solve the above technical problems in the prior art.
[0055] The technical solutions of the embodiments of this application and the technical effects produced by the technical solutions of this application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0056] Figure 1 It is a schematic diagram of the system architecture for implementing the image authenticity detection method provided by the embodiments of this application. The system can include a terminal 100 and a server 200.
[0057] The terminal 100 can be an electronic device such as a PC (Personal Computer), tablet computer, mobile phone, medical device, etc. A client for running a target application can be installed in the terminal 100. The target application can be an autonomous driving application or other applications with image processing functions, such as chat applications, sports and health applications, life service applications, etc. This application does not make any limitations in this regard. In addition, this application does not limit the form of the target application, including but not limited to an App (Application) installed in the terminal 100, a small program, etc., and can also be in the form of a web page. The terminal 100 can also be an in-vehicle terminal. Optionally, the in-vehicle terminal is used to collect the driver's face image. Optionally, the in-vehicle terminal can be used to collect the face image and simultaneously perform the processing of detecting the authenticity of the face image. In one example, the in-vehicle terminal realizes the processing of detecting the authenticity of the face image by establishing a connection with the server 200. In another example, the in-vehicle terminal can realize the processing of detecting the authenticity of the face image by itself.
[0058] The server 200 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The server 200 can be the background server of the above target application, used to provide background services for the client of the target application.
[0059] The terminal 100 and the server 200 can communicate with each other through a network, such as a wired or wireless network.
[0060] In the image authenticity detection method provided by the embodiments of the present application, the execution entity of each step can be a computer device, which refers to an electronic device with data calculation, processing, and storage capabilities. Taking Figure 1 the solution implementation environment shown as an example, the model training method can be executed by the terminal 100 (for example, the client of the target application installed and running in the terminal 100 executes the model training method), or the server 200 can execute the model training method, or the terminal 100 and the server 200 can cooperate with each other to execute. The present application does not make any limitations in this regard. For the convenience of description, in the following method embodiments, only the execution entity of each step of the image authenticity detection method being a computer device will be introduced and described.
[0061] Optionally, the technical solution provided by the present application can be applied in the intelligent transportation scenario. Exemplarily, the image acquisition device is connected to the in-vehicle terminal, and the in-vehicle terminal is connected to the server. Optionally, the connection between the image acquisition device and the in-vehicle terminal and the connection between the in-vehicle terminal and the server can communicate through a wireless connection. The present application does not make any limitations in this regard. The server executes the model training method to obtain a trained authenticity recognition model, and deploys the trained authenticity recognition model in the server. The image acquisition device acquires the driver's image and sends the image to the in-vehicle terminal, and the in-vehicle terminal then uploads the image to the server, and the server performs authenticity detection on the image. Optionally, the trained authenticity recognition model can also be deployed in the in-vehicle terminal. After receiving the driver's image sent by the image acquisition device, the in-vehicle terminal directly performs authenticity detection on the image, and then directly determines whether to perform keyless start according to the recognition result.
[0062] In the embodiments of the present application, an image authenticity detection method is provided, as Figure 2 shown. The method includes:
[0063] S101. Input the image to be detected into the authenticity recognition model, and obtain the authenticity detection result of the image to be detected output by the authenticity recognition model.
[0064] The authenticity recognition model in the embodiments of the present application pre-stores real text features and forged text features.
[0065] The real text features are determined according to each real description information in the pre-determined description information set, representing the semantic commonality of all real description information; while the forged text features are determined according to each forged description information in the description information set, representing the semantic commonality of all forged description information.
[0066] The description information set of the embodiments of the present application includes multiple pieces of description information, and each piece of description information has multiple fields: a status field, an object field, and a template field. Among them, the status field is used to indicate the authenticity of the object shown in the image, the object field is used to describe the category of the object shown in the image, and the template field is used to describe the scene type of the corresponding image shown in a target description manner.
[0067] That is to say, the description information of the embodiments of the present application not only needs to describe whether the object shown in the image is real or forged, but also needs to describe the content in the image from two dimensions of object and template, so as to improve the interpretation of the picture.
[0068] In some embodiments, the status field is used to describe the authenticity of the object shown in the corresponding image. Specifically, the status field can be real, natural, original, forged, false, etc.
[0069] The object field is used to describe the category of the object shown in the corresponding image. After analysis, it is found that there may be forged pictures obtained by exchanging animals and human faces in the pre-training data of the CLIP model. Therefore, non-human face text descriptions can be introduced to utilize this part of the pre-training knowledge. On the other hand, in order to better distinguish real and forged pictures, the model should not only focus on the alignment relationship between face images and text. Therefore, non-human text descriptions are added to the description information, such as object fields like "cat", "dog", "horse", etc.
[0070] The template field of the embodiments of the present application is used to describe the scene type of the image shown in a target description manner, such as a photo with an empty scene, a photo of a house scene, a photo of an office scene, etc.
[0071] In some embodiments, in order to enable the model to more easily understand the content described by each field in a piece of description information, taking the description information including a status field, an object field, and a template field as an example, the nested relationship between the three fields is defined: the object field is embedded in the status field, and the status field is embedded in the template field. If the object field is represented as [c], the status field is represented as [t], and the template field is represented as [t], then the description information can be represented as [t[s[c]]].
[0072]
[0073]
[0074] Table 1 Field Table
[0075] Please refer to Table 1, which exemplarily shows the field tables of three types of fields provided by the embodiments of the present application. Each field in Table 1 shows multiple examples, and it can be seen from the table that each status field and template field has a {} symbol, which represents the nesting position. An object field can be embedded in the {} of a template field. Taking the status field'real{}' and the object field 'face' as an example, embedding 'face' into'real{}' will obtain the status field'real{'face'}' after embedding the object field. Similarly, embedding'real{'face'}' into a template field, such as 'A dark photoof a{}', will obtain a description text 'A dark photo of a{'real'{'face'}}'.
[0076] It should be noted that current deepfake methods usually perform some forgery operations in the facial feature regions of the face, resulting in forgery artifacts in the facial feature parts of the face, and the image features of the facial feature regions may be different from the image features of the entire face. The facial feature regions of a real face are more real and clear. Therefore, since forgery of a forged face usually appears in the eyes, nose, and mouth regions, and the eye, nose, and mouth features of a real face are more aligned with the corresponding text, the embodiments of the present application design to add an additional facial feature field in the real description text compared to the forged description information. This facial feature field describes the facial features in the image, such as a text description like "with eyes, mouth and nose", so as to improve the classification ability of the model for real faces.
[0077] The embodiments of the present application construct a set of description information, and different description information is used to describe the content displayed in different images, so as to cover the image content of all images to be detected in the text description.
[0078] It should be understood that the real object in the embodiments of the present application is an object that has not been processed by AI or is not AI-generated, and the forged object is an object that has been processed by AI or is AI-generated.
[0079] In some embodiments, the embodiments of the present application can pre-collect multiple status fields, object fields, and template fields, and then generate multiple pieces of description information by means of free combination. It should be understood that the number of fields in all pieces of description information is the same. For example, taking the three types of fields shown in Table 1 as an example, since there are 17 status fields, 11 template fields, and 4 object fields, 17 * 11 * 4, a total of 748 pieces of description information can be generated.
[0080] There are two types of authenticity detection results in the embodiments of the present application. One is the first result, which describes that the object shown in the corresponding image is real, and the other is the second result, which describes that the object shown in the corresponding image is forged.
[0081] The method for detecting the authenticity of an image in the embodiments of the present application obtains the authenticity detection result of the image to be detected output by the authenticity recognition model. The authenticity recognition model in the embodiments of the present application pre-stores real text features and forged text features. The real text features are determined according to each real description information in the pre-determined description information set, and the forged text features are determined according to each forged description information in the description information set. Thus, it can be seen that the real text features and the forged text features can accurately describe the authenticity of the object shown in the image corresponding to a description information (i.e., text) from the semantic perspective. And all the description information in the present application includes a status field, an object field, and a template field. The object field is used to describe the category of the object shown in the corresponding image, and the template field is used to describe the scene type of the corresponding image shown in the target description manner; the status field of the real description information is used to describe that the object shown in the corresponding image is real, and the status field of the forged description information is used to describe that the object shown in the corresponding image is forged, so that the authenticity recognition model can learn fine-grained semantic information, improve the generalization ability of the model, and further improve the accuracy of image authenticity detection.
[0082] Moreover, in view of the problem that the zero-shot classification ability of the CLIP model is poor because the two categories of real and forged in the deepfake detection task cannot be clearly defined using text, the embodiments of the present application design a large number of description texts related to the authenticity detection task to make up for the above problems of the CLIP model. And since there are various types of target objects in the pre-training data of the CLIP model and there is already a good alignment between images and texts in the feature space, when the CLIP model is used as the authenticity recognition model in the embodiments of the present application, zero-shot authenticity detection can be performed without training, and no additional training is required for the image authenticity recognition scenario. Based on this, the embodiments of the present application can meet the need for urgently launching the deepfake detection function.
[0083] Based on the above embodiments, as an optional embodiment, inputting the image to be detected into a pre-trained authenticity recognition model to obtain the authenticity detection result of the image to be detected output by the authenticity recognition model includes:
[0084] S201. Input the image to be detected into the authenticity recognition model to obtain the image features of the image to be detected;
[0085] S202. Determine the similarities between the image features and the pre-determined real text features and forged text features respectively, and determine the authenticity detection result according to the magnitude relationship between the two similarities.
[0086] In the embodiments of the present application, by obtaining the image features of the image to be detected, calculating the similarity between the image features and the real text features and the similarity between the image features and the forged text features respectively, and determining the authenticity detection result according to the magnitude relationship between the two similarities. Since the real text features and the forged text features are both pre-determined by the model, the computational complexity of the model for determining the authenticity detection result is significantly reduced, and the detection efficiency is improved.
[0087] In some embodiments, determining the authenticity detection result according to the magnitude relationship between the two similarities may include:
[0088] If the first similarity is higher than the second similarity, the authenticity detection result is the first result; if the first similarity is not higher than the second similarity, the authenticity detection result is the second result.
[0089] The real text features in the embodiments of the present application are determined according to each real description information in the description information set. The status field of the real description information is used to describe that the object shown in the corresponding image is a real object. That is, the real text features can accurately summarize the semantic information of each real description information. The status field of the forged description information is used to describe that the object shown in the corresponding image is a forged object. The forged text features are determined according to each forged description information in the description information set. That is, the forged text features can accurately summarize the semantic information of each forged description information. If the image features of the image to be detected are more similar to the real text features, then the image is considered a real image; otherwise, the image to be detected is considered a forged image.
[0090] Please refer to Figure 3 , which exemplarily shows a schematic flowchart of a method for detecting the authenticity of an image provided by another embodiment of the present application. The authenticity recognition model in the embodiments of the present application is a CLIP model. The CLIP model includes an image encoding module. As shown in the figure, the image to be detected is input into the image encoding module in the authenticity recognition model, and the image encoding module outputs the image features I of the image to be detected. Further, the similarity ITr between the image features I and the pre-determined real text features Tr, and the similarity ITf between the image features I and the pre-determined forged text features Tf are calculated. If ITr is greater than ITf, it is determined that the image to be detected is a real image; if ITr is not greater than ITf, it is determined that the image to be detected is a forged image.
[0091] Through analysis, it is found in the embodiments of the present application that if the status fields, object fields, and scenario fields are simply combined freely to generate a description information set, when clustering the description information set, it will be found that there is a phenomenon in some description information clusters: there are a small number of forged description information in the clusters mainly composed of real description information, and there are a small number of real description information in the clusters mainly composed of forged description information. Therefore, the embodiments of the present application also provide a scheme for denoising the description information set. Specifically, the method for obtaining the description information set includes:
[0092] S301. Obtain a field set, where the field set includes multiple status fields, multiple object fields, and multiple template fields, and randomly generate multiple description information according to the field set;
[0093] S302. Classify the multiple description information according to the template fields in each description information to obtain multiple description information clusters, and each description information cluster includes description information with the same scenario type;
[0094] S303. For each description information cluster, determine the text features of each description information in the description information cluster;
[0095] S304. For each description information in each description information cluster, determine the first average value of the similarity of the text features between this description information and other description information with the same status fields, and the second average value of the similarity of the text features between this description information and other description information with different status fields, and use the difference between the first average value and the second average value as the index value of this description information;
[0096] S305. For each description information cluster, sort the description information in the description information cluster from large to small according to the index values of each description information, and retain the preset number of description information with the top ranking results to form a description information set.
[0097] Specifically, in step S304 of the embodiments of the present application, the first text features of each real description information and the second text features of each forged description information in the description information cluster are determined respectively;
[0098] For each real description information, determine the first average value of the similarity between the real description information and the first text features of each real description information in the description information cluster, determine the second average value of the similarity between the real description information and the second text features of each forged description information in the description information cluster, and use the difference between the first average value and the second average value as the index value of the real description information;
[0099] For each forged description information, determine the third mean of the similarities between the forged description information and the second text features of each forged description information in the description information cluster, determine the fourth mean of the similarities between the forged description information and the first text features of each true description information in the description information cluster, and use the difference between the third mean and the fourth mean as the index value of the forged description information.
[0100] It should be understood that the higher the index value of a description information, the closer the relationship between this description information and the description information of other same status fields, and the more distant the relationship with the description information of other different status fields. The status field of this description information is also more accurate. Therefore, for each description information cluster in the embodiments of the present application, only the preset number of description information with the top ranking results is retained.
[0101] It should be noted that the reason why the embodiments of the present application use the scenario field as the basis for classification is that through analysis, it is found that the true and false probabilities corresponding to different scenarios are different. For example, there are differences in forgery traces in a very dark scenario and in a very bright scenario. Therefore, different scenarios should have different true and false representations. Therefore, clustering the scenarios and then finding the corresponding true and false representations in each scenario will be more accurate.
[0102] Based on the above embodiments, as an alternative embodiment, before step S404, regularization processing can also be performed on each text feature. Regularization can prevent overfitting, reduce the degree of overfitting of the model to the training data, improve the generalization ability of the model, and can also reduce the interference of noise and improve the robustness of the model.
[0103] Please refer to Figure 4 , which exemplarily shows a schematic flowchart of the method for detecting the authenticity of images provided by the embodiments of the present application. As shown in the figure, it includes:
[0104] S401. Obtain a field set, where the field set includes multiple status fields, multiple object fields, and multiple template fields, and randomly generate multiple description information according to the field set;
[0105] S402. Classify the multiple description information according to the template fields in each description information to obtain multiple description information clusters, and each description information cluster includes description information with the same template type;
[0106] S403. For each description information cluster, determine the text features of each description information in the description information cluster, and for each description information, determine the mean of the similarities between the text features of this description information and the text features of other description information with the same status field, and the mean of the similarities between the text features of this description information and the text features of other description information with different status fields, and use the difference between the two means as the index value of the description information;
[0107] S404. For each cluster of description information, sort the description information in the cluster in descending order according to the index values of the respective description information, and retain a preset number of description information with the top-ranked sorting results to form a description information set;
[0108] S405. Input the description information set into the authenticity recognition model to obtain real text features and forged text features, and input the image to be detected into the authenticity recognition model to obtain the image features of the image to be detected;
[0109] S406. Determine the similarities between the image features of the image to be detected and the real text features and the forged text features respectively, and determine the authenticity detection result of the sample image according to the magnitude relationship between the two similarities.
[0110] Please refer to Figure 5, which exemplarily shows a schematic flowchart of a method for detecting the authenticity of an image according to another embodiment of the present application. As shown in the figure, first, the embodiment of the present application obtains a field set, which includes multiple status fields, object fields, and template fields. By randomly combining the field set, multiple description information is obtained. Taking the figure as an example, 3×3×3, a total of 27 pieces of description information can be obtained. Then, it is necessary to denoise the multiple pieces of description information. According to the template fields in each piece of description information, the multiple pieces of description information are classified to obtain multiple description information clusters. Each description information cluster includes description information with the same scene type (for example, the template fields of all the description information in a description information cluster are "a darkphote of a{}". In the embodiment of the present application, there are also some description information with empty scene types described by the template fields, such as 'A photo of a{}.', 'A dark photo of a{}.', and 'A grey photo of a{}.'. These several template fields do not involve specific scenes. Therefore, the description information with empty scene types described by these descriptions is located in the same description information cluster). For each description information cluster, the text features of each piece of description information in the description information cluster are determined, and for each piece of description information, the mean value of the similarity of the text features between the description information and other description information with the same status field is determined, as well as the mean value of the similarity of the text features between the description information and other description information with different status fields. The difference between the two mean values is used as the index value of the description information. For each description information cluster, the description information in the description information cluster is sorted from large to small according to the index value, and the preset number of description information with the top-ranked results is retained to form a description information set. The description information set is input into the text encoding module of the CLIP model to obtain the text features of each piece of description information. The mean value of the text features of each piece of real description information is taken to obtain the real text features, and the mean value of the text features of each piece of forged description information is taken to obtain the forged text features. The image to be detected is input into the image encoding module of the CLIP model to obtain the image features of the image to be detected. Further, the similarities between the image features of the image to be detected and the real text features and the forged text features are determined respectively, and the predicted authenticity result of the image to be detected is determined according to the magnitude relationship between the two similarities.
[0111] The deep face forgery technology has promoted the emerging development of the entertainment and cultural exchange industries, but at the same time has brought huge potential threats to face security. The embodiment of the present application can be applied to forged face detection products, such as the upgrade of face biometric authentication technology, to improve the security of multiple services such as face payment and identity authentication. This technology helps to verify evidence and prevent the forgery of evidence using Deepfakes-related technologies. On multimedia platforms, the widespread dissemination of face-swapped videos has continuously reduced the credibility of the media and is likely to mislead users. Please refer to Figure 6, which exemplarily shows a schematic diagram of the application of the present application to the video screening scenario. As shown in the figure, the broadcaster creates a video through the first terminal 601 and uploads it to the multimedia platform 602. The multimedia platform 602 identifies the people in the video. If it is determined that the person belongs to a public figure, the video is marked as in the review state. The video marked as in the review state is not visible to the audience of the multimedia platform 602. The multimedia platform 602 sends the video to the review server 603 for authenticity detection, that is, to identify whether the people in the video are forged through the deep face forgery technology in the embodiments of the present application. If it is determined that the image in the video is a real image, a notice of passing the review is sent to the multimedia platform 602. The multimedia platform 602 sets the video to the public state, and the audience can use the second terminal 604 to search for and view the video. If it is determined that the image in the video is a forged image, a notice of not passing the review is sent to the multimedia platform 602, and the multimedia platform 602 sends an alarm to the first terminal.
[0112] The embodiments of the present application can help the platform screen videos, add significant marks to the detected forged videos, such as "produced by Deepfakes", ensure the credibility of the video content, and guarantee social credibility. Overall, the embodiments of the present application can be used in products such as face identity verification, judicial verification tools, and picture and video authenticity verification.
[0113] The embodiments of the present application provide a device for detecting the authenticity of an image, as Figure 7 shown, including a model processing module 701. Specifically:
[0114] The model processing module 701 is configured to input the image to be detected into the authenticity recognition model, and obtain the authenticity detection result of the image to be detected output by the authenticity recognition model;
[0115] Among them, the authenticity recognition model pre-stores real text features and forged text features. The real text features are determined according to each real description information in the pre-determined description information set, and the forged text features are determined according to each forged description information in the description information set;
[0116] Any one of the real description information and the forged description information is used to describe an image and is obtained through alignment processing with the image. The description information includes a status field, an object field, and a template field. The status field is used to indicate the authenticity of the object shown in the image, the object field is used to describe the category of the object shown in the image, the template field is used to describe the scene type of the corresponding image shown in the target description manner, the object field is embedded in the status field, and the status field is embedded in the template field;
[0117] The real description information status field is used to indicate that the object shown in the corresponding image is real, and the status field of the forged description information is used to indicate that the object shown in the corresponding image is forged.
[0118] In some alternative embodiments, the model processing module is specifically configured to:
[0119] Input the image to be detected into the authenticity recognition model to obtain the image features of the image to be detected;
[0120] Determine the similarities between the image features and the pre-determined real text features and forged text features respectively, and determine the authenticity detection result according to the magnitude relationship between the two similarities.
[0121] In some alternative embodiments, the authenticity recognition model is a contrastive language-image pre-training CLIP model.
[0122] In some alternative embodiments, the image authenticity detection device further includes a description information set acquisition module:
[0123] The field set acquisition unit is configured to obtain a field set, where the field set includes multiple status fields, multiple object fields, and multiple template fields, and randomly generate multiple description information according to the field set;
[0124] The classification unit is configured to classify the multiple description information according to the template fields in each description information to obtain multiple description information clusters, and each description information cluster includes description information with the same scene type;
[0125] The metric calculation unit is configured to, for each description information cluster, determine the text features of each description information in the description information cluster, and for each description information, determine the mean of the similarities between the text features of the description information and the text features of other description information with the same status field, and the mean of the similarities between the text features of the description information and the text features of other description information with different status fields, and use the difference between the two means as the metric value of the description information;
[0126] The sorting and screening unit is configured to, for each description information cluster, sort the description information in the description information cluster from largest to smallest according to the metric values of the description information, and retain a preset number of description information with the top-ranked sorting results to form a description information set.
[0127] In some alternative embodiments, the authenticity recognition model is used to recognize the authenticity of a face image, the image to be detected is a face image, and the real description information further includes facial feature fields, and the facial feature fields are used to describe the facial features in the corresponding image.
[0128] In some alternative embodiments, the authenticity identification model is used to identify the authenticity of a face image. The image to be detected is a face image, and the object field describes that the category of the object shown in the corresponding image includes non-face.
[0129] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions executed by each module in the device according to the embodiments of the present application correspond to the steps in the methods according to the embodiments of the present application. For the detailed function descriptions of each module of the device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details are not described herein again.
[0130] An electronic device is provided in an embodiment of the present application, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the method for detecting the authenticity of an image or the method for model training. Compared with the related art, it can be achieved that: by inputting the image to be detected into a pre-trained authenticity identification model, the authenticity detection result of the image to be detected output by the authenticity identification model is obtained. Since the authenticity identification model according to the embodiment of the present application, in addition to the sample image set and the authenticity detection results of each sample image in the sample image set, further requires a description information set during training, the description information set includes a plurality of description information, the description information includes a plurality of fields, the plurality of fields includes a status field for describing the authenticity of the object shown in the corresponding image and at least one second field for describing the content in the corresponding dimension in the corresponding image, so that the authenticity identification model learns fine-grained semantic information, improves the generalization ability of the model, and further improves the accuracy of image authenticity detection. Moreover, for the problem that the zero-shot classification ability of the CLIP model is poor because the two categories of real and forged in the deep fake detection task cannot be clearly defined using text, a large number of description texts related to the authenticity detection task are designed in the embodiment of the present application to make up for the above problem of the CLIP model. And since there are various types of target objects in the pre-trained data of the CLIP model and there is already a good alignment between images and texts in the feature space, in the embodiment of the present application, when using the CLIP model as the authenticity identification model, zero-shot authenticity detection can be performed without training, and no additional training is required for the image authenticity identification scenario. Based on this, the embodiment of the present application can meet the requirement of quickly launching the deep fake detection function.
[0131] In an alternative embodiment, an electronic device is provided, as Figure 8 shown Figure 8The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as being connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data transmission and / or data reception, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.
[0132] The processor 4001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present application. The processor 4001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0133] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 8 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0134] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited here.
[0135] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0136] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding content of the foregoing method embodiments can be implemented.
[0137] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding content of the foregoing method embodiments can be implemented.
[0138] Terms such as "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims, and the above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than the illustrated or textually described order.
[0139] It should be understood that although the flowcharts of the embodiments of the present application indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0140] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.
Claims
1. A method for detecting the authenticity of an image, characterized in that: include: Inputting the image to be detected into the authenticity recognition model to obtain the authenticity detection result of the image to be detected output by the authenticity recognition model; The authenticity identification model pre-stores real text features and forged text features, wherein the real text features are determined based on each real description information in a predetermined description information set, and the forged text features are determined based on each forged description information in the description information set; Any one of the true description information and the forged description information is used to describe an image and is obtained by alignment with the image. The description information includes a status field, an object field and a template field. The status field is used to indicate the authenticity of the object displayed by the image. The object field is used to describe the category of the object displayed by the image. The template field is used to describe the scene type displayed by the image in a target description manner. The object field is embedded in the status field, and the status field is embedded in the template field. The status field of the real description information is used to indicate that the object displayed by the image is real, and the status field of the forged description information is used to indicate that the object displayed by the image is forged.
2. The method according to claim 1, characterized in that The step of inputting the image to be detected into a pre-trained authenticity recognition model to obtain an authenticity detection result of the image to be detected output by the authenticity recognition model includes: Inputting the image to be detected into the authenticity recognition model to obtain image features of the image to be detected; Determine the similarities between the image feature and the real text feature and the forged text feature, and determine the authenticity detection result according to the magnitude relationship between the two similarities.
3. The method according to claim 1 or 2, characterized in that the The authenticity identification model is a comparative language image pre-training CLIP model.
4. The method according to claim 1, characterized in that: The method for obtaining the description information set includes: Obtaining a field set, the field set including multiple status fields, multiple object fields, and multiple template fields, and randomly generating multiple description information according to the field set; Classifying the plurality of description information according to the template fields in the respective description information to obtain a plurality of description information clusters, each description information cluster including description information of the same scene type; For each description information cluster, determine the text features of each description information in the description information cluster, and for each description information, determine the accuracy of clustering the description information into the description information cluster according to the similarity between the text features of the description information and other description information with the same status field and the similarity between the text features of the description information and other description information with different status fields; For each description information cluster, the description information in the description information cluster is sorted from large to small according to the accuracy, and a preset number of description information with top sorting results is retained to form the description information set.
5. The method according to claim 4, characterized in that Determining the accuracy of clustering the description information into the description information cluster according to the similarity between the text features of the description information and other description information with the same status field and the similarity between the text features of the description information and other description information with different status fields comprises: Determine a first mean value of similarities between the text features of the description information and other description information having the same status field, and a second mean value of similarities between the text features of the description information and other description information having different status fields; The difference between the first mean and the second mean is used as the accuracy of clustering the description information to the description information cluster.
6. The method according to claim 1, characterized in that The real text features and the forged text features are determined in the following way: Inputting the description information set into the authenticity recognition model to obtain text features of each description information; Taking the average of the text features of each real description information to obtain the real text features; The text features of each forged description information are averaged to obtain the forged text features.
7. The method according to claim 1, characterized in that The authenticity recognition model is used to identify the authenticity of a face image, the image to be detected is a face image, and the real description information includes a facial features field, and the facial features field is used to describe the facial features in the image.
8. The method according to claim 1, characterized in that The authenticity recognition model is used to identify the authenticity of a face image, the image to be detected is a face image, and the object field describes the category of the object displayed in the corresponding image, including non-face.
9. A device for detecting the authenticity of an image, characterized in that: include: A model processing module, used to input the image to be detected into the authenticity recognition model, and obtain the authenticity detection result of the image to be detected output by the authenticity recognition model; The authenticity identification model pre-stores real text features and forged text features, wherein the real text features are determined based on each real description information in a predetermined description information set, and the forged text features are determined based on each forged description information in the description information set; Any one of the true description information and the forged description information is used to describe an image and is obtained by alignment with the image. The description information includes a status field, an object field and a template field. The status field is used to indicate the authenticity of the object displayed by the image. The object field is used to describe the category of the object displayed by the image. The template field is used to describe the scene type displayed by the corresponding image in a target description manner. The object field is embedded in the status field, and the status field is embedded in the template field. The status field of the real description information is used to indicate that the object displayed by the corresponding image is real, and the status field of the forged description information is used to indicate that the object displayed by the corresponding image is forged.
10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method for detecting the authenticity of an image according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting the authenticity of an image according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for detecting the authenticity of an image according to any one of claims 1 to 8 is implemented.