Image processing method and apparatus, storage medium, electronic device, and program product

By using the description information set in deep forged image detection, the generalization ability and detection accuracy of the model are improved, the problems of low detection accuracy and long training time in the prior art are solved, and the ability to go online is realized.

WO2025107773A1PCT designated stage expired Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/114209
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-08-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing deep fake image detection methods use digital labels, resulting in weak generalization capabilities of the model, low detection accuracy, and long training time, which cannot meet the needs of emergency online.

Method used

By pre-determining the description information set, including real description information and fake description information, the authenticity and false identification model is used to learn fine-grained semantic information, and the generalization ability of the model is improved.

Benefits of technology

It improves the accuracy of image authenticity detection, shortens training time, meets the needs of emergency online, and improves detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024114209_30052025_PF_FP_ABST
    Figure CN2024114209_30052025_PF_FP_ABST
Patent Text Reader

Abstract

An image processing method, comprising: pre-determining a description information set, the description information set comprising a plurality of pieces of real description information and a plurality of pieces of forged description information; separately inputting the plurality of pieces of real description information and the plurality of pieces of forged description information into an authenticity recognition model, and obtaining a plurality of real text features and a plurality of forged text features. Any piece of description information among the real description information and the forged description information comprises a state field, an object field, and a template field, and the state field is used for indicating whether an object displayed in an image is real or forged; the object field is used for indicating the category of the object; the template field is used for indicating the type of a scenario displayed in the image; the object field is embedded in the state field, and the state field is embedded in template field. Further disclosed are an image processing apparatus, a storage medium, an electronic device, and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, device, storage medium, electronic device and program product

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 20, 2023, with application number 202311553842.5 and application name “Image authenticity detection method, device, storage medium and electronic device”. Technical Field

[0002] The present application relates to the field of artificial intelligence technology. Specifically, the present application relates to an image processing method, device, electronic device, computer-readable storage medium and computer program product.

[0003] Background of the Invention

[0004] The rapid development of deep fake facial technology has brought entertainment and convenience, but also created huge security risks.

[0005] Existing detection methods use digital labels as supervisory information to identify whether faces in sample images are real or fake. This label-based image classification results in poor generalization of the trained models, which in turn reduces detection accuracy. Furthermore, because model training is time-consuming, these technologies are unable to meet the urgent need to implement deepfake detection functionality in an application.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide an image processing method, apparatus, electronic device, computer-readable storage medium, and computer program product that can solve the above-mentioned problems of the prior art. The technical solution is as follows:

[0008] In one aspect, an embodiment of the present application provides an image processing method, the method comprising:

[0009] The image to be detected is input into the authenticity recognition model to identify whether the image to be detected is a real image or a forged image; wherein,

[0010] Predetermining a description information set, wherein the description information set includes a plurality of real description information and a plurality of forged description information;

[0011] Inputting the plurality of authentic description information and the plurality of forged description information into the authenticity recognition model respectively to obtain a plurality of authentic text features and a plurality of forged text features;

[0012] Any one of the real description information and the forged description information includes a status field, an object field, and a template field, wherein:

[0013] The status field is used to indicate whether the object shown in the image is real or fake;

[0014] The object field is used to indicate the category of the object;

[0015] The template field is used to indicate the type of scene shown by the image; and

[0016] The object field is embedded in the status field, and the status field is embedded in the template field.

[0017] On the other hand, an embodiment of the present application provides an image processing device, including:

[0018] The model processing module is used to input the image to be detected into the authenticity recognition model to identify whether the image to be detected is a real image or a forged image; wherein,

[0019] Predetermining a description information set, wherein the description information set includes a plurality of real description information and a plurality of forged description information;

[0020] Inputting the plurality of authentic description information and the plurality of forged description information into the authenticity recognition model respectively to obtain a plurality of authentic text features and a plurality of forged text features;

[0021] Any one of the real description information and the forged description information includes a status field, an object field, and a template field, wherein:

[0022] The status field is used to indicate whether the object shown in the image is real or fake;

[0023] The object field is used to indicate the category of the object;

[0024] The template field is used to indicate the type of scene shown by the image; and

[0025] The object field is embedded in the status field, and the status field is embedded in the template field.

[0026] On the other hand, an embodiment of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the steps of the above-mentioned image processing method.

[0027] On the other hand, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned image processing method when the computer program is executed by a processor.

[0028] On the other hand, an embodiment of the present application provides a computer program product, including a computer program, which implements the steps of the above-mentioned image processing method when executed by a processor.

[0029] BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.

[0031] FIG1 is a schematic diagram of a system architecture for implementing an image processing method according to an embodiment of the present application;

[0032] FIG2 is a schematic diagram of a flow chart of an image processing method provided in an embodiment of the present application;

[0033] FIG3 is a schematic diagram of the structure of the authenticity identification model provided in an embodiment of the present application;

[0034] FIG4 is a schematic diagram of a flow chart of an image processing method provided in an embodiment of the present application;

[0035] FIG5 is a flow chart of an image processing method provided in another embodiment of the present application;

[0036] FIG6 is a schematic diagram of an application of a video review scenario provided by an embodiment of the present application;

[0037] FIG7 is a schematic structural diagram of an image processing device provided in an embodiment of the present application;

[0038] FIG8 is a schematic structural diagram of an electronic device provided in an embodiment of the present application.

[0039] Implementation Method

[0040] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0041] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an" and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to that the element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or as "B", or as "A and B".

[0042] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0043] First, several terms involved in this application are introduced and explained:

[0044] Computer vision technology (CV) refers to the use of cameras and computers to replace the human eye to identify and measure targets, and further perform graphic processing to obtain images that are more suitable for human eye observation or transmission to instrument detection.

[0045] Deepfake technology uses machine learning models, such as generative adversarial networks, to merge and overlay images or videos onto source images or videos. Using neural network technology for large-scale learning, it splices together an individual's voice, facial expressions, and body movements to create fake content. For example, deepfakes enable artificial intelligence (AI) face-swapping, voice simulation, face synthesis, and video generation. Its emergence makes it possible to manipulate or generate highly realistic and difficult-to-distinguish audio and video content, ultimately making it impossible for observers to distinguish authenticity with the naked eye.

[0046] In related technologies, the detection methods for deep fake images all use digital labels, ignoring fine-grained semantic information. This fine-grained semantic information can help the model improve its generalization ability. Therefore, the inventive concept of the embodiment of this application is to introduce fine-grained semantic information into the training process of the model, thereby improving the generalization ability of the model.

[0047] The Contrastive Language-Image Pre-Training (CLIP) model is a pre-trained neural network model for matching images and text, with zero-shot classification capabilities.

[0048] However, compared to other classification tasks, the two categories of real and fake in the deepfake detection task cannot be clearly defined using text. Therefore, the feature information of images and text matched by the CLIP model cannot be fully utilized, resulting in poor zero-shot classification ability of the CLIP model.

[0049] In other words, on the one hand, existing deepfake detection methods use digital labels, and the supervision information of digital labels does not contain semantic information, so the model cannot learn the true meaning of real and fake. On the other hand, because the objects in deepfake images are all faces, only the traces of forgery are found in local areas of the face. The categories of real and fake cannot be clearly defined using text, resulting in the CLIP model not being directly applicable to deepfake detection.

[0050] The image processing method, device, electronic device, computer-readable storage medium, and computer program product provided in this application are intended to solve the above technical problems in the prior art.

[0051] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.

[0052] FIG1 is a schematic diagram of a system architecture for implementing an image processing method provided in an embodiment of the present application. The system may include a terminal 100 and a server 200 .

[0053] The terminal 100 may be an electronic device such as a PC (Personal Computer), a tablet computer, a mobile phone, or a medical device. A client for running a target application may be installed in the terminal 100. The target application may be an autonomous driving application or other application with image processing capabilities, such as a chat application, a sports and health application, or a life service application, but this application does not limit this.

[0054] In addition, the present application does not limit the form of the target application, including but not limited to an App (Application) or a mini-program installed in the terminal 100, or a web page. The terminal 100 may also be a vehicle-mounted terminal.

[0055] Optionally, the vehicle-mounted terminal is used to collect the driver's facial image. Optionally, the vehicle-mounted terminal can be used to collect the facial image and simultaneously perform a process of detecting the authenticity of the facial image.

[0056] In one example, the vehicle-mounted terminal establishes a connection with the server 200 to detect the authenticity of the facial image.

[0057] In another example, the vehicle-mounted terminal can detect the authenticity of facial images by itself.

[0058] The server 200 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The server 200 may be a background server of the target application, used to provide background services for the client of the target application.

[0059] The terminal 100 and the server 200 may communicate with each other via a network, such as a wired or wireless network.

[0060] In the image processing method provided in the embodiment of the present application, the execution subject of each step can be a computer device, and the computer device refers to an electronic device with data calculation, processing and storage capabilities. Taking the solution implementation environment shown in Figure 1 as an example, the image processing method can be executed by the terminal 100, for example, the client is installed in the terminal 100, the target application is run, and the image processing method is executed. The image processing method can also be executed by the server 200, or the terminal 100 and the server 200 interact and cooperate to execute it. This application does not limit this. For ease of explanation, in the following method embodiments, only the execution subject of each step of the image processing method is introduced as a computer device.

[0061] Optionally, the technical solution provided by this application can be applied in smart transportation scenarios. For example, the image acquisition device is connected to the vehicle terminal, and the vehicle terminal is connected to the server.

[0062] Optionally, the image acquisition device is connected to the vehicle-mounted terminal, and the vehicle-mounted terminal is connected to the server and can communicate via a wireless connection, which is not limited in this application.

[0063] Optionally, the server executes the image processing method to obtain a trained authenticity recognition model, which is then deployed on the server. The image acquisition device captures an image of the driver and transmits the image to the vehicle terminal, which then uploads the image to the server, which then performs authenticity verification on the image.

[0064] Optionally, the trained authenticity recognition model can also be deployed in the vehicle terminal. After receiving the driver's image sent by the image acquisition device, the vehicle terminal directly performs authenticity detection on the driver's image, and then determines whether to perform keyless start based on the recognition result.

[0065] An embodiment of the present application provides an image processing method, as shown in FIG2 , comprising:

[0066] S101: Input the image to be detected into the authenticity recognition model to identify whether the image to be detected is a real image or a forged image.

[0067] In this step, the authenticity recognition model is used to output the authenticity detection result of the image to be detected, that is, whether it is a real image or a forged image.

[0068] In the embodiment of the present application, the authenticity recognition model pre-stores real text features and forged text features.

[0069] The real text feature is determined based on each real description information in a predetermined description information set, and represents the semantic commonality of all real description information; while the forged text feature is determined based on each forged description information in the description information set, and represents the semantic commonality of all forged description information.

[0070] In one embodiment, the authenticity identification model includes a text encoding module, and features of genuine text and forged text are determined in the following manner:

[0071] Obtaining, by a text encoding module, a text feature of each of the plurality of true description information and a text feature of each of the plurality of forged description information;

[0072] Take the average of the text features of each real description information to obtain the real text features;

[0073] The text features of each forged description information are averaged to obtain the forged text features.

[0074] In the embodiment of the present application, the description information set includes multiple description information, specifically, multiple real description information and multiple forged description information.

[0075] Each description information corresponds to an image and is used to describe various information of the image. Specifically, each description information has multiple fields: a status field, an object field, and a template field. Each field is obtained by aligning (or matching) with the image. Specifically,

[0076] (1) The status field is used to indicate the authenticity of the object displayed in the image, that is, whether the object is real or forged, which corresponds to whether the description information is real description information or forged description information.

[0077] Specifically, the information included in the status field may be: real, natural, native, forged, false, etc.

[0078] (2) The object field is used to indicate the category of the object displayed in the image.

[0079] Because some fake images are created by swapping human faces with animals, for example, the pre-training data for the CLIP model includes this type of fake images. In this case, the human faces are no longer human faces, but replaced with animal faces. Therefore, in one embodiment, the category indicated by the object field includes non-human faces. This allows the pre-trained knowledge of these fake images to be utilized through text descriptions of non-human faces.

[0080] In another embodiment, in order to better distinguish between real and fake pictures, the model should not only focus on the matching relationship (or alignment relationship) between the face image and the text, so non-human text descriptions can also be added to the description information, that is, the category indicated by the object field includes non-humans, such as object fields such as "cat", "dog", and "horse".

[0081] (3) The template field is used to indicate the type of scene displayed by the image.

[0082] The target description method can be pre-set to indicate the type of scene, including keywords, word order, etc. An example is shown in Table 1 below.

[0083] The scene type can be empty scene, house scene, office scene, etc.

[0084] It can be seen that the description information of the embodiment of the present application not only needs to describe whether the object displayed in the image is real or forged, but also needs to describe the content in the image from two dimensions: object and template, thereby enhancing the interpretability of the authenticity recognition model for real and forged images.

[0085] In some embodiments, to make it easier for the model to understand the content described in each field of a description, for example, a description including a status field, an object field, and a template field is defined as a nested relationship between the three fields: the object field is embedded in the status field, and the status field is embedded in the template field. If the object field is represented as [c], the status field as [t], and the template field as [t], then the description can be represented as [t[s[c]]].

[0086] Table 1 Field table

[0087] Please refer to Table 1, which exemplarily shows a field table of three fields provided in an embodiment of the present application. Each field in Table 1 shows multiple examples, and it can be seen from the table that each status field and template field has a {} symbol, which indicates a nested position. An object field can be embedded in the {} of a template field. Taking the status field 'real{}' and the object field 'face' as an example, 'face' is embedded in 'real{}', that is, the status field 'real{'face'}' after the embedded object field is obtained. Similarly, 'real{'face'} is embedded in a template field, for example, 'A dark photo of a{}', and a description text 'A dark photo of a{'real'{'face'}}' can be obtained.

[0088] It's important to note that current deepfake methods typically perform some manipulation on the facial features, resulting in artifacts. The features of these features may differ significantly from those of the entire face. The facial features of a real face are more realistic and clear.

[0089] Therefore, since the forgery of a fake face usually occurs in the eyes, nose, and mouth areas, and the features of the eyes, nose, and mouth of a real face are more closely matched to the features of the corresponding text, in an embodiment of the present application, in the real description information, an additional facial features field is added compared to the fake description information. The facial features field is used to describe the facial features information in the image, for example, it can be a text description of "with eyes, mouth and nose", thereby improving the model's ability to classify real faces.

[0090] In the embodiment of the present application, by constructing a description information set, different description information is used to describe the various contents of objects in different images, thereby covering the image contents of all images to be detected in the text description.

[0091] In the embodiments of the present application, a real object refers to an object that has not been processed by AI or is not generated by AI, and a forged object refers to an object that has been processed by AI or generated by AI.

[0092] In some embodiments, multiple status fields, multiple object fields, and multiple template fields can be preselected and collected, and then multiple description information can be generated through a predetermined combination. The number of fields in each description information is consistent, for example, 3. Taking the three fields shown in Table 1 as an example, since there are 17 status fields, 11 template fields, and 4 object fields, 17 * 11 * 4 = 748 description information can be generated.

[0093] In the embodiment of the present application, there are two kinds of authenticity detection results. The first result indicates that the object displayed in the image is real, and the second result indicates that the object displayed in the image is forged.

[0094] The image processing method in the embodiment of the present application inputs the image to be detected into the authenticity recognition model to identify whether the image to be detected is a real image or a forged image, wherein the authenticity recognition model pre-stores real text features and forged text features, the real text features are determined based on each real description information in a predetermined description information set, and the forged text features are determined based on each forged description information in the description information set. It can be seen that the real text features and the forged text features can accurately characterize the authenticity of the object displayed in the image corresponding to a description information (i.e., text) from a semantic perspective.

[0095] In addition, the description information in the embodiments of the present application includes a status field, an object field, and a template field. The status field is used to indicate whether the object displayed in the image is real or fake; the object field is used to indicate the category of the object; and the template field is used to indicate the type of scene displayed in the image. Among them, the status field of the real description information is used to indicate that the object displayed in the image is real, and the status field of the fake description information is used to indicate that the object displayed in the image is fake. This allows the authenticity recognition model to learn fine-grained semantic information, improve the generalization ability of the model, and thus improve the accuracy of image authenticity detection.

[0096] In addition, to address the problem that the CLIP model has poor zero-sample classification ability because the two categories of real and fake in the deep fake detection task cannot be clearly defined using text, the embodiment of the present application designs a large number of descriptive texts related to the authenticity detection task to compensate for the above-mentioned problem of the CLIP model. Moreover, since there are various types of target objects in the pre-training data of the CLIP model, the image and text have been matched in the feature space. Therefore, when the authenticity recognition model in the embodiment of the present application is constructed based on the CLIP model, zero-sample authenticity detection can be performed without training, and no additional training is required for the image authenticity recognition scenario. Based on this, the embodiment of the present application can meet the needs of urgently launching the deep fake detection function.

[0097] Based on the above embodiments, as an optional embodiment, the image to be detected is input into a pre-trained authenticity recognition model, and the authenticity recognition result of the image to be detected output by the authenticity recognition model is obtained, including:

[0098] S201, inputting the image to be detected into the authenticity recognition model to obtain image features of the image to be detected;

[0099] S202, determining a first similarity between the image feature and the real text feature;

[0100] S203, determining a second similarity between the image feature and the forged text feature;

[0101] S204: Determine whether the image to be detected is a real image or a forged image based on the first similarity and the second similarity.

[0102] In an embodiment of the present application, by obtaining the image features of the image to be detected, the similarity between the image features and the real text features and the similarity between the image features and the forged text features are calculated respectively, and the authenticity detection result is determined based on the size relationship of the two similarities. Since the real text features and the forged text features are both predetermined by the model, the amount of calculation required by the model to determine the authenticity detection result is significantly reduced, thereby improving the detection efficiency.

[0103] In some embodiments, determining the authenticity detection result based on the magnitude relationship between the two similarities may include:

[0104] When the first similarity is higher than the second similarity, determining that the authenticity detection result is the first result, that is, the image to be detected is a real image;

[0105] When the first similarity is not higher than the second similarity, the authenticity detection result is determined to be the second result, that is, the image to be detected is a forged image.

[0106] In an embodiment of the present application, the real text feature is determined based on each real description information in the description information set, and the status field of the real description information is used to indicate that the object displayed in the image is real. It can be seen that the real text feature can accurately summarize the semantic information of each real description information; the forged text feature is determined based on each forged description information in the description information set, and the status field of the forged description information is used to indicate that the object displayed in the image is forged. It can be seen that the forged text feature can accurately summarize the semantic information of each forged description information. If the image feature of the image to be detected is more similar to the real text feature, then the image is considered to be a real image. Otherwise, the image to be detected is considered to be a forged image.

[0107] Please refer to Figure 3, which exemplarily shows a structural diagram of an authenticity recognition model provided in an embodiment of the present application. The authenticity recognition model 300 includes an image encoding module 302.

[0108] As shown in Figure 3, the image to be detected 301 is input into the image encoding module 302, and the image encoding module 302 outputs the image feature I of the image to be detected 301, and further calculates the similarity ITr between the image feature I and the predetermined real text feature Tr, as well as the similarity ITf between the image feature I and the predetermined forged text feature Tf.

[0109] In step 303 , it is determined whether ITr is greater than ITf. If ITr is greater than ITf, the image to be detected 301 is determined to be a real image 304 . If ITr is not greater than ITf, the image to be detected 301 is determined to be a forged image 305 .

[0110] In addition, if the status field, object field, and template field are simply freely combined to generate a description information set, and when the description information set is clustered, it is found that a phenomenon exists in some description information clusters: a small amount of forged description information exists in a cluster with mostly real description information, and a small amount of real description information exists in a cluster with mostly forged description information. Therefore, the embodiment of the present application also provides a solution for denoising the description information set. Specifically, the method for obtaining the description information set includes:

[0111] S301: Obtain a field set, where the field set includes multiple status fields, multiple object fields, and multiple template fields.

[0112] S302: Combine multiple status fields, multiple object fields, and multiple template fields to generate multiple description information.

[0113] The combination method can be random combination; or, traversal in a certain order, combination and so on.

[0114] S303: Classify the plurality of description information according to the template fields in the respective description information to obtain a plurality of description information clusters, each description information cluster including a plurality of description information of the same scene type.

[0115] S304: For each description information cluster, perform the following processing:

[0116] Determining text features of each descriptive information in the descriptive information cluster;

[0117] For each description information in the description information cluster,

[0118] determining at least a third similarity between the text feature of the description information and text features of other description information having the same status field;

[0119] determining at least one fourth similarity between the text feature of the description information and text features of other description information having different status fields;

[0120] determining, based on at least one third similarity and at least one fourth similarity, an accuracy with which the description information is clustered into the description information cluster;

[0121] The accuracy of each piece of description information is sorted from large to small, and a preset number of description information with the highest sorting results are retained to form the description information set.

[0122] Specifically, determining the accuracy of clustering the description information into the description information cluster based on at least one third similarity and at least one fourth similarity includes: calculating the difference between the mean of at least one third similarity and the mean of at least one fourth similarity; and using the difference as the accuracy.

[0123] Depending on whether the description information is true or fabricated, the above-mentioned method of determining accuracy specifically includes:

[0124] respectively determining a first text feature of each authentic description information and a second text feature of each forged description information in the description information cluster;

[0125] For each true description information, determining a first mean of the third similarities between the true description information and the first text features of each true description information in the description information cluster, determining a second mean of the fourth similarities between the true description information and the second text features of each forged description information in the description information cluster, and using the difference between the first mean and the second mean as the accuracy of the true description information;

[0126] For each forged description information, determine the third mean of the third similarity between the forged description information and the second text feature of each forged description information in the description information cluster, determine the fourth mean of the fourth similarity between the forged description information and the first text feature of each real description information in the description information cluster, and use the difference between the third mean and the fourth mean as the accuracy of the forged description information.

[0127] It should be understood that the higher the accuracy of a description information, the closer the relationship between the description information and other description information of the same status field, and the more distant the relationship between the description information and other description information of different status fields, the more accurate the status field of the description information. Therefore, the embodiment of the present application only retains a preset number of description information at the top of the sorting results for each description information cluster.

[0128] It should be noted that in this embodiment, the template field is used as the basis for classification because the corresponding authenticity probabilities vary in different scenarios. For example, there is a clear difference between forgery traces in a very dark scene and in a very bright scene, so different scenes have different authenticity representations. Therefore, clustering the scenes and then determining the authenticity representation for each scene will be more accurate.

[0129] Based on the above embodiments, as an optional embodiment, each text feature can also be regularized before step S404. Regularization can prevent overfitting, reduce the degree of overfitting of the model to the training data, improve the generalization ability of the model, and also reduce the interference of noise and improve the robustness of the model.

[0130] Please refer to FIG4 , which exemplarily shows a flowchart of the image processing method provided in an embodiment of the present application, as shown in the figure, including:

[0131] S401: Obtain a field set, where the field set includes multiple status fields, multiple object fields, and multiple template fields.

[0132] S402: Combine multiple status fields, multiple object fields, and multiple template fields to generate multiple description information.

[0133] S403: Classify the plurality of description information according to the template fields in each description information to obtain a plurality of description information clusters, each description information cluster including a plurality of description information of the same scene type.

[0134] S404: For each description information cluster, perform the following processing:

[0135] Determining text features of each descriptive information in the descriptive information cluster;

[0136] For each description information in the description information cluster,

[0137] determining at least a third similarity between the text feature of the description information and text features of other description information having the same status field;

[0138] determining at least one fourth similarity between the text feature of the description information and text features of other description information having different status fields;

[0139] determining, based on at least one third similarity and at least one fourth similarity, an accuracy with which the description information is clustered into the description information cluster;

[0140] The accuracy of each piece of description information is sorted from large to small, and a preset number of description information with the highest sorting results are retained to form the description information set.

[0141] S405: Input the description information set into the authenticity recognition model, and obtain multiple real text features and multiple forged text features through the text encoding module.

[0142] S406: Input the image to be detected into the authenticity recognition model, and obtain the image features of the image to be detected through the image encoding module.

[0143] S407: Determine a first similarity between the image feature and the real text feature, determine a second similarity between the image feature and the forged text feature, and determine whether the image to be detected is a real image or a forged image based on the first similarity and the second similarity.

[0144] Please refer to Figure 5, which illustrates a flowchart of an image processing method according to another embodiment of the present application. As shown in Figure 5, first, a field set 501 is obtained. Field set 501 includes multiple object fields 5011, status fields 5012, and template fields 5013. These fields are randomly combined to obtain multiple description information 5021, forming description information set 502. For example, if there are three object fields 5011, three status fields 5012, and three template fields 5013, 3*3*3 = 27 description information 5021 can be obtained.

[0145] Then, the plurality of description information 5021 is subjected to denoising 503, as described in steps S301-S304 above. Specifically, the denoising 503 includes:

[0146] Based on the content of the template field 5013 in each description information, the multiple description information are clustered to obtain multiple description information clusters 5041. Each description information cluster 5041 includes description information of the same scene type. For example, the template field of all description information in a description information cluster is "a dark photo of a{}".

[0147] In addition, in the embodiment of the present application, there are some description information template fields in which the scene types indicated are empty, such as 'A photo of a{}.', 'A dark photo of a{}.' and 'A grey photo of a{}.'. These template fields do not involve specific scenes, so the description information with these scene types being empty is located in the same description information cluster.

[0148] In step 505, for each description information cluster, the text features of each description information in the cluster are determined. For each description information, the average of at least one third similarity between the text features of the description information and the text features of other description information with the same status field, and the average of at least one fourth similarity between the text features of the description information and other description information with different status fields, are determined. The difference between the two averages is used as the accuracy of the description information. Furthermore, for each description information cluster, the accuracy of each description information in the cluster is sorted from highest to lowest. A preset number of description information with the highest ranking results are retained, and all retained description information constitutes a description information set.

[0149] The authenticity identification model 500 includes a text encoding module 506 and an image encoding module 510 .

[0150] The filtered description information is taken as a description information set and input into the text encoding module 506 to obtain the text features of the description information.

[0151] Furthermore, the text features 5071 of each real description information are averaged to obtain the real text feature 5081; the text features 5072 of each forged description information are averaged to obtain the forged text feature 5082.

[0152] The image to be detected 509 is input into the image encoding module 510 to obtain the image features 510 of the image to be detected 509, and the similarities between the image features 510 and the real text features 5081 and the forged text features 5082 are further determined. According to the size relationship between the two similarities, the authenticity detection result 512 of the image to be detected is determined.

[0153] Deepfake technology has driven the development of the entertainment and cultural exchange industries, but it also poses a significant potential threat to human facial security. The widespread dissemination of face-swapped videos on multimedia platforms has led to a decline in media credibility and can easily mislead users.

[0154] The embodiments of the present application can be applied to forged face detection products, such as upgrading facial authentication technology to improve the security of face-based payment, identity authentication, and other services. The embodiments of the present application can help verify the authenticity of evidence and prevent the use of Deepfakes-related technologies to forge evidence.

[0155] Please refer to Figure 6, which exemplarily shows a schematic diagram of the present application applied to a video screening scenario. As shown in the figure, the anchor produces a video through a first terminal 601 and uploads it to a multimedia platform 602.

[0156] The multimedia platform 602 identifies the person designed in the video and determines whether the person is a public figure.

[0157] When it is determined that the person is a public figure, the video is marked as under review. The video marked as under review is not visible to viewers of the multimedia platform, such as viewers using the second terminal 604. The multimedia platform 602 sends the video marked as under review to the review server 603 for authenticity detection. That is, through the method described in the embodiments of the present application, it is identified whether the person in the video has been forged using deep face forgery technology, and an authenticity detection result is generated, that is, whether the image in the video is a real image or a forged image.

[0158] If the audit server 603 determines that the image in the video is a real image, it sends a notification of audit approval to the multimedia platform 602, and the multimedia platform 602 sets the video to a public state, and the audience can use the second terminal 604 to search and browse the video.

[0159] If the audit server 603 determines that the image in the video is a forged image, it sends a notification of audit failure to the multimedia platform 602, and the multimedia platform 602 sends an alarm to the first terminal 601.

[0160] The embodiments of the present application can help the platform screen videos and add prominent marks to the detected fake videos, such as "produced by Deepfakes", to ensure the credibility of the video content and guarantee social credibility.

[0161] The embodiments of the present application can also be used in products such as face verification, judicial verification tools, and image and video authentication.

[0162] The embodiment of the present application provides an image processing device 700, as shown in FIG7 , including a model processing module 701, specifically:

[0163] Model processing module 701, used for

[0164] The image to be detected is input into the authenticity recognition model to identify whether the image to be detected is a real image or a forged image; wherein,

[0165] Predetermining a description information set, wherein the description information set includes a plurality of real description information and a plurality of forged description information;

[0166] Inputting the plurality of authentic description information and the plurality of forged description information into the authenticity recognition model respectively to obtain a plurality of authentic text features and a plurality of forged text features;

[0167] Any one of the real description information and the forged description information includes a status field, an object field, and a template field, wherein:

[0168] The status field is used to indicate whether the object shown in the image is real or fake;

[0169] The object field is used to indicate the category of the object;

[0170] The template field is used to indicate the type of scene shown by the image; and

[0171] The object field is embedded in the status field, and the status field is embedded in the template field.

[0172] In some optional implementations, the model processing module 701 is specifically configured to:

[0173] Inputting the image to be detected into the authenticity recognition model to obtain image features of the image to be detected;

[0174] Determining a first similarity between the image feature and the real text feature;

[0175] determining a second similarity between the image feature and the forged text feature;

[0176] It is determined whether the image to be detected is a real image or a forged image according to the first similarity and the second similarity.

[0177] In some optional implementations, the image processing apparatus further includes a description information set acquisition module 702, comprising:

[0178] A field set acquisition unit is configured to obtain a field set, the field set including a plurality of status fields, a plurality of object fields, and a plurality of template fields; and combine the plurality of status fields, the plurality of object fields, and the plurality of template fields to generate a plurality of description information;

[0179] a classification unit, configured to classify the plurality of description information according to the template field in each description information to obtain a plurality of description information clusters, each description information cluster including a plurality of description information having the same scene type;

[0180] The indicator calculation unit is used to perform the following processing for each description information cluster:

[0181] Determining text features of each descriptive information in the descriptive information cluster;

[0182] For each description information in the description information cluster,

[0183] determining at least a third similarity between the text feature of the description information and text features of other description information having the same status field;

[0184] determining at least one fourth similarity between the text feature of the description information and text features of other description information having different status fields;

[0185] determining, based on the at least one third similarity and the at least one fourth similarity, an accuracy with which the description information is clustered into the description information cluster;

[0186] The sorting and screening unit is used to sort the accuracy of each description information from large to small for each description information cluster, and retain a preset number of description information with top sorting results to form the description information set.

[0187] In some optional implementations, the indicator calculation unit is configured to calculate a difference between an average of the at least one third similarity and an average of the at least one fourth similarity; and use the difference as the accuracy.

[0188] In some optional embodiments, the authenticity identification model includes a text encoding module, and the real text features and the forged text features are determined in the following manner: obtaining the text features of each real description information in the multiple real description information and the text features of each forged description information in the multiple forged description information through the text encoding module; taking the average of the text features of each real description information to obtain the real text features; taking the average of the text features of each forged description information to obtain the forged text features.

[0189] In some optional embodiments, the authenticity recognition model is used to identify the authenticity of a facial image, the image to be detected is a facial image, and the real description information includes a facial features field, which is used to describe the facial features information in the face.

[0190] In some optional implementations, the authenticity recognition model is used to identify the authenticity of a facial image, the image to be detected is a facial image, and the category indicated by the object field includes non-face.

[0191] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, and will not be repeated here.

[0192] In an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory. The processor executes the above-mentioned computer program to implement the steps of the image processing method. Compared with the related art, it can achieve: by inputting the image to be detected into the authenticity recognition model, the authenticity detection result of the image to be detected output by the authenticity recognition model is obtained.

[0193] In an optional embodiment, an electronic device is provided, as shown in FIG8 , and the electronic device 4000 shown in FIG8 includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.

[0194] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0195] Bus 4002 may include a path for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. Bus 4002 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, FIG8 shows a single thick line, but this does not imply that there is only one bus or only one type of bus.

[0196] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.

[0197] The memory 4003 is used to store the computer program for executing the embodiment of the present application, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiment.

[0198] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.

[0199] An embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.

[0200] The terms "first," "second," "third," "fourth," "1," "2," and the like (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than that shown or described in the drawings.

[0201] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.

[0202] The above are only optional implementation methods for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, other similar implementation methods based on the technical ideas of this application also fall within the protection scope of the embodiments of this application.

Claims

1. An image processing method, performed by an electronic device, comprising: The image to be detected is input into the authenticity recognition model to identify whether the image to be detected is a real image or a forged image; wherein, Predetermining a description information set, wherein the description information set includes a plurality of real description information and a plurality of forged description information; Inputting the plurality of real description information and the plurality of forged description information into the authenticity identification model respectively to obtain a plurality of real text features and a plurality of forged text features; Any one of the real description information and the forged description information includes a status field, an object field and a template field, wherein: The status field is used to indicate whether the object shown in the image is real or fake; The object field is used to indicate the category of the object; The template field is used to indicate the type of scene shown by the image; and, The object field is embedded in the status field, and the status field is embedded in the template field.

2. The method according to claim 1, wherein: The step of inputting the image to be detected into the authenticity recognition model to identify whether the image to be detected is a real image or a forged image includes: Inputting the image to be detected into the authenticity recognition model to obtain image features of the image to be detected; Determining a first similarity between the image feature and the real text feature; determining a second similarity between the image feature and the forged text feature; It is determined whether the image to be detected is a real image or a forged image according to the first similarity and the second similarity.

3. The method according to claim 1 or 2, wherein: The status field of the real description information is used to indicate that the object is real, and the status field of the forged description information is used to indicate that the object is forged.

4. The method according to any one of claims 1 to 3, wherein: The predetermined description information set includes: Obtaining a field set, the field set comprising a plurality of status fields, a plurality of object fields, and a plurality of template fields; Combining the multiple status fields, the multiple object fields, and the multiple template fields to generate multiple description information; Classifying the plurality of description information according to the template fields in the respective description information to obtain a plurality of description information clusters, each description information cluster including a plurality of description information having the same scene type; For each description information cluster, perform the following processing: Determine the text features of each description information in the description information cluster; For each description information in the description information cluster, Determining at least one third similarity between the text feature of the description information and text features of other description information having the same status field; determining at least one fourth similarity between the text feature of the description information and text features of other description information having different status fields; Determining, according to the at least one third similarity and the at least one fourth similarity, the accuracy with which the description information is clustered into the description information cluster; The accuracy of each piece of description information is sorted from large to small, and a preset number of description information with top sorting results are retained to form the description information set.

5. The method according to claim 4, wherein: The determining, according to the at least one third similarity and the at least one fourth similarity, the accuracy of clustering the description information into the description information cluster comprises: calculating a difference between a mean of the at least one third similarity and a mean of the at least one fourth similarity; The difference is taken as the accuracy.

6. The method according to any one of claims 1 to 5, wherein: The authenticity identification model includes a text encoding module, and the authentic text features and the forged text features are determined in the following manner: Obtaining, by means of the text encoding module, a text feature of each of the plurality of authentic description information and a text feature of each of the plurality of forged description information; Taking the average of the text features of each real description information to obtain the real text features; The text features of each forged description information are averaged to obtain the forged text features.

7. The method according to any one of claims 1 to 6, wherein: The authenticity recognition model is used to identify the authenticity of a face image, the image to be detected is a face image, and the real description information includes a facial features field, and the facial features field is used to describe the facial features information of the face.

8. The method according to any one of claims 1 to 7, wherein: The authenticity recognition model is used to identify the authenticity of a face image, the image to be detected is a face image, and the category indicated by the object field includes non-face.

9. An image processing device, comprising: The model processing module is used to input the image to be detected into the authenticity recognition model to identify whether the image to be detected is a real image or a forged image; wherein, Predetermining a description information set, wherein the description information set includes a plurality of real description information and a plurality of forged description information; Inputting the plurality of real description information and the plurality of forged description information into the authenticity identification model respectively to obtain a plurality of real text features and a plurality of forged text features; Any one of the real description information and the forged description information includes a status field, an object field and a template field, wherein: The status field is used to indicate whether the object shown in the image is real or fake; The object field is used to indicate the category of the object; The template field is used to indicate the type of scene shown by the image; and, The object field is embedded in the status field, and the status field is embedded in the template field.

10. The device according to claim 9, wherein: The model processing module is used to input the image to be detected into the authenticity recognition model to obtain image features of the image to be detected; determine a first similarity between the image features and the real text features; determine a second similarity between the image features and the forged text features; and determine whether the image to be detected is a real image or a forged image based on the first similarity and the second similarity.

11. The device according to claim 9 or 10, wherein: The status field of the real description information is used to indicate that the object is real, and the status field of the forged description information is used to indicate that the object is forged.

12. The device according to any one of claims 9 to 11, further comprising: Describe the information set acquisition module; The description information set acquisition module includes: A field set acquisition unit is used to obtain a field set, wherein the field set includes a plurality of status fields, a plurality of object fields, and a plurality of template fields; the plurality of status fields, the plurality of object fields, and the plurality of template fields are combined to generate a plurality of description information; A classification unit, configured to classify the plurality of description information according to the template field in each description information to obtain a plurality of description information clusters, each description information cluster including a plurality of description information having the same scene type; The indicator calculation unit is used to perform the following processing for each description information cluster: Determine the text features of each description information in the description information cluster; For each description information in the description information cluster, Determining at least one third similarity between the text feature of the description information and text features of other description information having the same status field; determining at least one fourth similarity between the text feature of the description information and text features of other description information having different status fields; Determining, according to the at least one third similarity and the at least one fourth similarity, the accuracy with which the description information is clustered into the description information cluster; The sorting and screening unit is used to sort the accuracy of each description information from large to small for each description information cluster, and retain a preset number of description information with top sorting results to form the description information set.

13. The device according to claim 12, wherein: The indicator calculation unit is used to calculate the difference between the mean of the at least one third similarity and the mean of the at least one fourth similarity; and use the difference as the accuracy.

14. The device according to any one of claims 9 to 13, wherein: The authenticity identification model includes a text encoding module, and the real text features and the forged text features are determined in the following manner: obtaining the text features of each of the multiple real description information and the text features of each of the multiple forged description information through the text encoding module; taking the average of the text features of each real description information to obtain the real text features; taking the average of the text features of each forged description information to obtain the forged text features.

15. An electronic device comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the image processing method according to any one of claims 1 to 8.

16. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the image processing method according to any one of claims 1 to 8.

17. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the image processing method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Commodity identification method and system based on multi-modal data

    CN115601582A

  • Electrical equipment electric spark image generation method based on two-stage generative adversarial network

    CN116563410A

  • Image authenticity detection method and device, computer equipment and storage medium

    CN117037182A

  • Dynamic and automatic generation of interactive text related objects

    US20170286390A1