Image editing trace recognition method, image editing trace recognition model training method, computer storage medium and program product
By acquiring the editing trace features of an image and training it using a preset editing pool and an image editing trace recognition model, the problem of image editing trace recognition was solved, and the authenticity of the image was determined.
Patent Information
- Application Number
- PCT/CN2025/081129
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-03-06
- Publication Date
- 2026-01-15
AI Technical Summary
Existing technologies struggle to effectively identify editing traces in images, making it difficult to determine the authenticity and credibility of images.
By acquiring the editing trace features of the image to be identified, including the editing trace features inside and outside the image acquisition device, generating image training samples using a preset editing pool, training the image editing trace recognition model, and combining the trace feature extractor and the text feature extractor for comparative learning and classification, the editing traces of the image can be identified.
It enables effective identification of image editing traces, determines whether an image has been edited and the type of editing, and provides evidence of image authenticity.
Smart Images

Figure CN2025081129_15012026_PF_FP_ABST
Abstract
Description
Image editing trace recognition and its model training methods, computer storage media and program products
[0001] This application claims priority to Chinese Patent Application No. 202410909997.6, filed on July 8, 2024, entitled "Image Editing Trace Recognition and Model Training Method Thereof, Computer Storage Medium and Program Product", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence technology, and in particular to an image editing trace recognition method, an image editing trace recognition model training method, a computer storage medium, and a computer program product. Background Technology
[0003] With the development of image editing and AI (Artificial Intelligence) technologies, many images have been created that have been edited and modified from the original images, or generated or synthesized purely through software. This has led to doubts about the authenticity and credibility of these images.
[0004] In this context, determining the authenticity of an image and effectively identifying editing traces becomes an urgent problem to be solved. Summary of the Invention
[0005] In view of this, embodiments of this application provide an image editing trace recognition scheme to at least partially solve the above-mentioned problems.
[0006] According to a first aspect of the embodiments of this application, an image editing trace recognition method is provided, comprising: acquiring an image to be recognized; extracting editing trace features from the image to be recognized to obtain corresponding editing trace features, wherein the editing trace features include editing trace features inside an image acquisition device and editing trace features outside an image acquisition device; and recognizing editing traces in the image based on the editing trace features.
[0007] According to a second aspect of the embodiments of this application, a training method for an image editing trace recognition model is provided, comprising: performing editing operations on original images in an original image set based on multiple image editing operation sequences in a preset editing pool to obtain image training samples; extracting editing trace features from the image training samples using a trace feature extractor in the image editing trace recognition model to obtain corresponding editing trace features; and extracting operation features from text labels corresponding to the image editing operation sequences using a text feature extractor in the image editing trace recognition model to obtain corresponding editing operation features; wherein the editing trace features include editing trace features inside the image acquisition device and editing trace features outside the image acquisition device, and the editing operation features include editing operation features inside the image acquisition device and editing operation features outside the image acquisition device; calculating a loss based on the editing trace features and the editing operation features; and training the image editing trace recognition model based on the loss.
[0008] According to a third aspect of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is used to store at least one executable instruction, wherein the executable instruction causes the processor to perform an operation corresponding to the method described in the first or second aspect.
[0009] According to a fourth aspect of the embodiments of this application, a computer storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the method as described in the first or second aspect.
[0010] According to a fifth aspect of the embodiments of this application, a computer program product is provided, including computer instructions that instruct a computing device to perform an operation corresponding to the method described in the first or second aspect.
[0011] According to the scheme provided in this application embodiment, feature extraction is performed on the image to be identified. Unlike traditional methods that extract semantic information from images, the scheme in this application embodiment extracts editing trace features from the image to obtain editing trace features. These editing trace features include both internal and external editing trace features of the image acquisition device. Internal editing trace features can characterize whether the image originates from the same type of image acquisition device, while external editing trace features can effectively characterize whether the image has undergone editing processing during transmission and what kind of editing processing it has undergone. In some scenarios, the editing processing of the image during transmission can indicate the image's source or even its source path. Based on this, the editing traces of the image to be identified are determined using the editing trace features of the image to be identified.
[0012] Therefore, the solution of this application embodiment can effectively identify the editing traces of an image by using the editing trace features of the image, so as to determine whether the image has been edited and what kind of editing has been done, and thus provide a basis for determining the authenticity of the image. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0014] Figure 1 is a schematic diagram of an exemplary system to which the embodiments of this application are applied;
[0015] Figure 2A is a flowchart of the steps of a training method for an image editing trace recognition model according to an embodiment of this application;
[0016] Figure 2B is a comparative schematic diagram of different ways of generating image training samples in the embodiment shown in Figure 2A;
[0017] Figure 2C is a schematic diagram of the model structure of an image editing trace recognition model in the embodiment shown in Figure 2A;
[0018] Figure 2D is a schematic diagram of the model structure of another image editing trace recognition model in the embodiment shown in Figure 2A;
[0019] Figure 3A is a flowchart of the steps of an image editing trace recognition method according to an embodiment of this application;
[0020] Figure 3B is a schematic diagram of an application of editing trace recognition in the embodiment shown in Figure 3A;
[0021] Figure 3C is a schematic diagram of another application of editing trace recognition in the embodiment shown in Figure 3A;
[0022] Figure 4 is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0023] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.
[0024] The specific implementation of the embodiments of this application will be further described below with reference to the accompanying drawings.
[0025] Figure 1 illustrates an exemplary system for a check code generation method applicable to embodiments of this application. As shown in Figure 1, the system 100 may include a cloud server 102, a communication network 104, and / or one or more user devices 106, with multiple user devices being an example in Figure 1.
[0026] The cloud server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 102 can perform any suitable function. For example, in some embodiments, the cloud server 102 can be used to perform image editing trace recognition. As an optional example, in some embodiments, the cloud server 102 can extract editing trace features from the image to be recognized to obtain corresponding editing trace features, and perform image editing trace recognition based on these features. The extracted editing trace features include editing trace features inside the image acquisition device and editing trace features outside the image acquisition device. As an optional example, in some embodiments, the cloud server 102 is equipped with a trace feature extractor, which allows the cloud server 102 to extract editing trace features from the image to be recognized. As another example, in some embodiments, the cloud server 102 can receive a recognition request sent by the user device 106, which carries information about the image to be recognized; after receiving the recognition request, the cloud server 102 obtains the image to be recognized based on the information to perform the aforementioned editing trace recognition. In addition, as an optional example, in some embodiments, the cloud server 102 may also send information about the identified editing traces to the user device 106.
[0027] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user equipment 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud server 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user equipment 106 and the cloud server 102, such as a network link, a dial-up link, a wireless link, a hardwired link, any other suitable communication link, or any suitable combination of such links.
[0028] User device 106 may include any one or more user devices suitable for presenting images and information and capable of interacting with a user. As an optional example, in some embodiments, user device 106 may send an identification request to cloud server 102, carrying information about the image to be identified, including but not limited to the image itself or its identifier or location information. In some embodiments, user device 106 may receive information about edit marks on an image returned by cloud server 102 and display it to the user. As an optional example, in some embodiments, user device 106 may include any suitable type of device. For example, in some embodiments, user device 106 may include mobile devices, tablet computers, laptop computers, desktop computers, wearable computers, game consoles, media players, vehicle entertainment systems, and / or any other suitable type of user device.
[0029] Based on the above system, this application provides an image editing trace recognition scheme, which will be described below through embodiments. To facilitate understanding of the scheme provided by this application, the structure and training process of the image editing trace recognition model involved in the scheme will be described first, and then the process of performing image editing trace recognition using the image editing trace recognition model will be described based on this.
[0030] Referring to FIG2A, a flowchart of the steps of a training method for an image editing trace recognition model according to an embodiment of the present application is shown.
[0031] The training method for the image editing trace recognition model in this embodiment includes the following steps:
[0032] Step S202: Based on multiple image editing operation sequences in the preset editing pool, perform editing operations on the original images in the original image set to obtain image training samples.
[0033] Here, "raw image" refers to an image in RAW format, which is raw image data captured and saved directly from the image sensor (such as CMOS or CCD) of an image acquisition device such as a digital camera or digital camcorder, without any form of processing or compression. The raw image collection includes a large number of raw images.
[0034] In this embodiment, because the image editing trace recognition model can extract image editing traces after training, the positive sample portion of the image training samples used for model training needs to be edited and carry corresponding editing operation information. Traditionally, taking a camera as an example, these positive samples can be obtained by acquiring a large number of image samples with known camera labels. However, this method requires a large amount of label data, making the acquisition of training samples costly and unable to cover future camera models. Furthermore, even for labeled image samples, traditional methods only consider editing traces within the camera, i.e., traces inside the camera, without considering traces outside the camera. This limits trace recognition to tasks strongly related to traces inside the camera (such as camera classification, image forgery evidence collection, etc.), and fails to address tasks related to traces outside the camera (such as social network transmission trace identification, synthetic image trace identification, etc.).
[0035] To avoid the aforementioned issues and ensure that the image training samples meet the training requirements of the image editing trace recognition model, this embodiment provides a preset editing pool. Optionally, this preset editing pool can be an online editing pool to meet the need for timely generation of image training samples. The preset editing pool contains various image editing operation sequences. By performing editing operations on the original images in the original image set based on these sequences, the corresponding image training samples can be obtained. Optionally, in addition to the images processed by the editing operations, the image training samples may also contain text labels corresponding to the editing operations (described below).
[0036] It should be noted that, in this embodiment, the editing operations on the original image include not only internal editing operations of the image acquisition device but also external editing operations. Accordingly, the preset editing pool includes a set of internal image acquisition device editing operations and a set of external image acquisition device editing operations. The set of internal image acquisition device editing operations includes various internal image acquisition device editing operations, and the set of external image acquisition device editing operations includes various external image acquisition device editing operations. Internal image acquisition device editing operations are used to simulate the operation of converting light signals into digital signals by the image acquisition device; external image acquisition device editing operations are used to simulate the editing processing operations that the image undergoes during propagation. Therefore, it can cover most image editing situations in practical applications, making the image editing operations more comprehensive and not dependent on specific image acquisition devices, hardware, software, or platforms.
[0037] For example, internal editing operations of the image acquisition device include, but are not limited to, de-mosaic operations, white balance operations, and tone mapping operations. External editing operations of the image acquisition device include, but are not limited to, JPEG compression operations, WEBP compression operations, scaling operations, blurring operations, ISO adjustments, and noise reduction operations. These operations encompass most of the operations that images may undergo in practical applications. Since the same operation may produce different editing traces, this application embodiment also introduces more detailed parameters to describe the operations. For example, a quality factor QF variable is introduced for JPEG compression and WEBP compression, with a range of 50 to 99; four interpolation types are introduced for scaling operations: Linear (bilinear interpolation), Nearest (nearest neighbor interpolation), Area (region interpolation), and Cubic (bicubic interpolation); different kernel sizes, from 3 to 9, are introduced for blurring operations such as Gaussian blurring. An exemplary editing operation and its optional parameters are shown in Table 1 below:
[0038] Table 1
[0039] It should be noted that, in practical applications, those skilled in the art can modify, add, or delete the operations and parameters in Table 1 according to actual needs. Because this preset editing pool is an online editing pool, this update operation can be performed in real time as needed.
[0040] To better simulate the complex editing process in the real world, instead of applying a single operation each time, this application embodiment uses multiple operations to form an operation sequence, and uses one operation sequence to process the original image each time. However, those skilled in the art should understand that the method of applying a single operation each time is also applicable to the solution of this application embodiment.
[0041] Based on this, in one feasible approach, the image editing operation sequence in the preset editing pool can be generated in the following way: arranging the internal editing operations of various image acquisition devices to obtain a corresponding first arrangement set, the first arrangement set including N1 first arrangements; and arranging the external editing operations of various image acquisition devices to obtain a corresponding second arrangement set, the second arrangement set including N2 second arrangements; and obtaining N1*N2 image editing operation sequences based on the first and second arrangement sets, where each of the N1*N2 image editing operation sequences is an operation sequence where the first arrangement precedes the second arrangement.
[0042] Because various internal and external editing operations of image acquisition devices are explicitly defined in the preset editing pool, image editing operations based on this pool can take into account both traces inside the image acquisition device, such as the camera, and traces outside the camera. Furthermore, since the preset editing pool can generate a massive number of editing operations, a small number of training samples can be greatly expanded without introducing a large number of training samples (e.g., only 10 images are needed for model training) or a large amount of label data. Moreover, because both internal and external editing operations of the image acquisition device are considered, the image editing trace recognition model trained on the image training samples can be widely applied to a wider range of image editing trace forensics tasks, such as social media platform source detection, synthetic image detection, image forgery detection, and camera model recognition.
[0043] It should be noted that the preset editing pool can also be configured to automatically generate corresponding text labels based on the image editing operation sequence. For example, if each editing operation corresponds to a unique text, a text sequence with the same order as the operations in the image editing operation sequence can be generated based on the text corresponding to each editing operation in the image editing operation sequence. The text in the text sequence is separated by preset delimiters, such as "-", spaces, etc. This text sequence is the text label of the image editing operation sequence.
[0044] In one example, let's assume an edit pool. Include For each image editing operation (including internal and external editing operations of the image acquisition device), the image editing operation sequence C of length m can be defined as: C = P′1→P′2→…→P′ m ,in,
[0045] set up The trace feature extractor E contains all possible operation sequences. θ The goal is to train samples C for different instances, i.e., images. i (X) Identify different edit marks Ii This is achieved without concern for the image semantics of the original image X. Given X1, X2, C1, and C2, when using the same operation sequence C1, it is expected that the edit traces of C1(X1) and C1(X2) corresponding to the original images X1 and X2 will be very similar. However, even when using different operation sequences for the same original image X1, it is expected that the edit traces of C1(X1) and C2(X1) will be very dissimilar. For each C... i Generate text label T according to the order of its operations. i For example, for an operation sequence consisting of AHD depixelation, automatic white balance, tonal mapping with pixel value scaling, and JPEG and WEBP with QF values of 75 and 80 respectively, such as operation sequence #1, the corresponding text label T1 is “DM-AHD WB-Auto Tone-Scale JPEG-75 WEBP-80”. In this sequence, the operation names and parameters of each editing operation are first connected by hyphens such as “-”, and then by spaces. However, this is not the only possibility; in practical applications, those skilled in the art can set other connection symbols according to actual needs.
[0046] As can be seen, by using the image editing operation sequences in the preset editing pool, the original images can be processed according to each image editing operation sequence to obtain the corresponding image training samples. Although in one feasible implementation, editing operations can be performed on a single image basis to obtain the corresponding single image training sample, in order to improve the training speed and efficiency of the image editing trace recognition model, another feasible approach is to first construct original image pairs based on the original images in the original image set; then, using the original image pairs as units, editing operations can be performed on the original image pairs based on various image editing operation sequences in the preset editing pool to obtain the corresponding positive example sample pairs; based on these positive example sample pairs, image training samples are obtained, that is, the positive example sample pair is used as an image training sample. Since the original images are all in RAW format, after two original images in the same original image pair are edited according to a certain image editing operation sequence, the two original images will have the same editing trace features.
[0047] For ease of explanation, the process of generating image training samples described above will be illustrated below with a simple example, referring to Figure 2B.
[0048] In Figure 2B, the interface to the left of the dashed line shows the processing for a single original image, while the interface to the right of the dashed line shows the processing for pairs of original images. In this example, the preset editing pool includes N1*N2 editing operation sequences. Editing operation sequence #1 is DM-WB-Tone-JPEG-WEBP, editing operation sequence #2 is DM-WB-ISO-JPEG, and other editing operation sequences are different from editing operation sequences #1 and #2, such as DM-WB-Tone-WEBP-Blur, etc., which are not listed here.
[0049] When processing a single original image, as shown on the left side of Figure 2B, assuming that the original image X1 is processed using the editing operation sequence #1 to obtain image training sample C(X1); and the original image X2 is processed using the editing operation sequence #2 to obtain image training sample C(X2), it can be seen that C(X1) and C(X2) have different editing operation traces.
[0050] When processing is done on a pair of original images, as shown on the right side of Figure 2B, the original images X1 and X2 are first constructed into a pair of original images, denoted as X1'. They are then processed using the edit operation sequence #1 to obtain the image training sample C(X1'). It can be seen that the two parts contained in C(X1'), namely X1 and X2, have the same edit operation traces.
[0051] The following sections will explain the image training samples generated in these two scenarios.
[0052] Step S204: Extract editing trace features from the image training samples using the trace feature extractor in the image editing trace recognition model to obtain the corresponding editing trace features. Also, extract operation features from the image editing operation sequence using the text feature extractor in the image editing trace recognition model to obtain the corresponding editing operation features.
[0053] Among them, editing trace features include editing trace features inside the image acquisition device and editing trace features outside the image acquisition device; editing operation features include editing operation features inside the image acquisition device and editing operation features outside the image acquisition device.
[0054] To facilitate understanding of the solutions in the embodiments of this application, the structure of the image editing trace recognition model in the embodiments of this application will be described below with reference to FIG2C and FIG2D. FIG2C and FIG2D respectively illustrate the model structure of two exemplary image editing trace recognition models.
[0055] The image editing trace recognition model includes a trace feature extractor, a text feature extractor, and a classifier (optional). In Figure 2C, the trace feature extractor is specifically implemented as an image encoder, the text feature extractor is specifically implemented as a text encoder, and the classifier is specifically implemented as a multilayer perceptron (MLP).
[0056] The image editing features are extracted using a trace feature extractor; the text feature extractor extracts operation features from the text labels corresponding to the image editing operation sequence; and the classifier classifies the edit trace features. In this embodiment, the model is trained using a joint image-text comparison learning approach based on both edit trace features and editing operation features.
[0057] To improve the training efficiency of the image editing trace recognition model, this application embodiment also employs a batch processing method. In this case, multiple image training samples can be obtained according to the sample quantity indicated by a single batch processing. That is, in the input to the image editing trace recognition model, one batch of image training samples is input each time. The specific number of images in one batch can be appropriately set by those skilled in the art based on the amount of input data acceptable to the model, and this application embodiment does not impose any limitations on this. In Figure 2C, a simple example of the number of image training samples input in one batch is 4 images. When generating image training samples on a single original image basis, these 4 image training samples can be based on the same or different 4 original images X, selecting 4 different image editing operation sequences C from a preset editing pool. i Image editing operations are performed on these four original images to obtain four corresponding training images, as shown in Figure 2C. i (X), where i represents the number of the image editing operation sequence, which can be from 1 to N1*N2, for example.
[0058] For a trace feature extractor, it can be represented as an image encoder E with trainable parameters θ. θ The expectation is that it can train on any input image sample C. i (X) Extract its editing trace feature I, i.e., I = E θ (X). Generally, the image training samples C i The dimension of (X) is Feature I is Where H and W are the length and width of the image training sample (e.g., 1024*768), respectively, and d is the length of the edit trace feature (e.g., 1024). For a batch of image training samples, the image encoder can process them in parallel, extracting edit trace features from multiple image training samples in that batch to obtain multiple corresponding edit trace features.
[0059] It should be noted that the specific implementation of the image encoder is not limited in the embodiments of this application. Those skilled in the art can select an appropriate image encoder, such as EfficientNet, ResNet, XceptionNet, ViT, MiT, etc., according to actual needs and the hardware performance of the device where the model is located.
[0060] On the other hand, since each batch of multiple image training samples has text labels corresponding to the image editing operation sequence, that is, each batch of multiple image training samples also corresponds to multiple text labels, in this case, we can first determine the text labels corresponding to the multiple image editing operation sequences corresponding to the multiple image training samples, and then use a text encoder to extract operation features from these multiple text labels to obtain the corresponding multiple editing operation features. For example, in Figure 2C, four image editing operation sequences #1, #2, #3, and #4 are used to process the original image respectively, obtaining four corresponding image training samples. These four image training samples will be used as a batch and input into the image editing trace recognition model. At the same time, these four image editing operation sequences also correspond to four text labels, namely text labels #1, #2, #3, and #4. The text encoder then extracts features from these four text labels to obtain the editing operation features corresponding to these four image training samples, which are illustrated as T1, T2, T3, and T4 in Figure 2C. The text encoder can also be implemented as a structure for extracting other editable features (a type of text feature), such as ALBERT, DistilBERT, etc.
[0061] Based on this, subsequent comparative learning can be carried out.
[0062] Unlike Figure 2C, in Figure 2D, since the image training samples are the results of processing the original image pairs according to the image editing operation sequence (i.e., positive sample pairs), the editing trace features extracted by the image encoder for a single image training sample include two parts, as illustrated in the figure. Where j represents the number of the image editing operation sequence, and i represents the number of the two parts of the image included in an image training sample. When the image training sample is obtained based on the original image pair, i = 1, 2; j = 1, 2, ..., N1*N2.
[0063] In this scenario, the image editing trace extractor, implemented as an image encoder, extracts editing trace features from positive sample pairs to obtain a first trace feature and a second trace feature corresponding to each positive sample pair. Based on the first and second trace features, it obtains the editing trace features for each positive sample pair. Adaptively, a self-built branch is provided in Figure 2D, which can: generate editing trace features for positive sample pairs based on the first and second trace features; extract editing operation features using a text editor; and perform comparative learning based on the editing trace features and editing operation features.
[0064] For example, in Figure 2D, after four image training samples are input into the image encoder, the image encoder processes them to obtain four sets of edit trace features, as illustrated in the figure. It includes: the first set of edit trace features corresponding to the first image training sample. and (i.e., the first and second trace features corresponding to the first image training sample), corresponding to the second set of edit trace features of the second image training sample. and (i.e., the first and second trace features corresponding to the second image training sample), corresponding to the third set of edit trace features of the third image training sample. and (i.e., the first and second trace features corresponding to the third image training sample), corresponding to the fourth set of edit trace features of the fourth image training sample. and (That is, the first and second trace features corresponding to the fourth image training sample). Based on this, for subsequent comparative learning, each set of trace features can be merged by creating a custom branch to form an edit trace feature corresponding to each image training sample.
[0065] In one feasible approach, this merging can be achieved by averaging, for example, for... and Calculate the average to obtain the final edit trace feature I1 corresponding to the first image training sample; for and Calculate the average to obtain the final edit trace feature I2 corresponding to the second image training sample; for and Calculate the average to obtain the final edit trace feature I3 corresponding to the third image training sample; for and The average is calculated to obtain the final edit trace feature I4 corresponding to the fourth image training sample.
[0066] However, this is not the only option. In another feasible approach, a fuzzy mapping (FM) strategy can be used to obtain the I corresponding to each image training sample. The FM strategy typically refers to a strategy that defines and describes membership (i.e., the degree to which an element belongs to a set) by introducing fuzzy sets and fuzzy logic when dealing with problems involving fuzziness or uncertainty. This strategy allows elements to belong to a set with a certain degree of membership, rather than the either-or (belong or not) relationship in traditional set theory. Specifically, in this example, based on the FM strategy, the membership matrix U = [u...] is set... ij Based on this membership matrix, aggregation is performed through weighted summation. This generates a new editing trace feature I j Each weight coefficient u ij express with I j The degree of correlation. Among them, for u ij The constraints are: Where i ranges from 1 to the number of original images contained in an image training sample, which is 2 in the case of original image pairs; j ranges from 1 to the number N of edit trace features to be generated, which generally corresponds to the number of image training samples in a batch, for example, 4 in Figure 2D. In Figure 2D, u ij Also expressed as
[0067] Optimization of U involves minimizing and I j The distance between them can be expressed as in ∑ j U .j Under the condition that = 1 Where i and j take values in the same range as above; ρ is the fuzzy coefficient, ρ∈[1,+∞); d(,) represents the distance function (such as l2 distance, cosine distance, etc.); U .j Let represent the vector in the j-th column of U.
[0068] The extraction of operational features using a text encoder is similar to the method shown in Figure 2C. Since each image training sample is generated from two original images, it undergoes the same image editing operation, thus corresponding to a text label. In this case, we can first determine the text labels corresponding to the multiple image editing operation sequences for multiple image training samples, and then extract operational features from these text labels using a text encoder to obtain the corresponding multiple editing operation features. For example, in Figure 2D, four image editing operation sequences #1, #2, #3, and #4 are used to process the original image pairs, resulting in four corresponding image training samples. These four image training samples will be used as a batch and input into the image editing trace recognition model. Simultaneously, these four image editing operation sequences also correspond to four text labels, namely text labels #1, #2, #3, and #4. The text encoder then extracts operational features from these four text labels (indicated as Text{T} in Figure 2D). j Feature extraction is performed on the four training images to obtain the editing operation features corresponding to each of the four images, which are also shown as T1, T2, T3, and T4 in Figure 2D.
[0069] Once the above-mentioned editing trace characteristics and editing operation characteristics are obtained, subsequent comparative learning can be carried out.
[0070] Step S206: Calculate the loss based on the editing trace features and editing operation features, and train the image editing trace recognition model based on the loss.
[0071] In one feasible approach, the image editing trace recognition model of this application embodiment employs a contrastive learning training method. Accordingly, this step can be implemented as follows: based on multiple editing trace features and multiple editing operation features, perform contrastive learning to obtain a first loss for the multiple editing trace features and multiple editing operation features; and train the image editing trace recognition model based on the first loss. Employing contrastive learning enables the model to learn richer and more general feature representations, reduces the need for labeled data, and enhances the model's generalization ability.
[0072] For example, suppose that from the preset edit pool Randomly selected A sequence of image editing operations C i and its corresponding text label T i This forms the input data for the current batch, i.e., the image training samples, represented as... First, edit trace features in the visual image are extracted using an image encoder and a text encoder, respectively. i =E θ (C i ) and text editing operation features Ti = in, Let I represent a text encoder with trainable parameters φ, and I i and T i They all belong to the same feature dimension R d The objective function for contrastive learning is used to maximize the cosine similarity between the edit trace features and edit operation features of matched pairings (grey boxes on the diagonal of the squares in Figures 2C and 2D), while minimizing the non-matching pairings (white boxes outside the diagonal of the squares in Figures 2C and 2D). This objective function can jointly optimize E. θ and Formally, the objective function along the visual image axis for:
[0073] Similarly, the objective function along the text axis for:
[0074] Where τ is a learnable temperature parameter.
[0075] Therefore, the overall objective function for contrastive learning can be...
[0076] While contrastive learning can train an image editing trace recognition model, it can lead to poor recognition performance in certain situations, such as when editing trace features are very similar. This is because contrastive learning has limitations in constraining the distance between different features, resulting in insufficient margins at the feature boundaries learned by the model. Therefore, this embodiment introduces additional classification learning to supervise the distance between different image training samples, thereby improving decision margins. In one feasible approach, this classification learning can be implemented using a classifier such as an MLP. However, it is not limited to this; other structures with classification capabilities, such as fully connected layers, can also be applied to the scheme of this embodiment to improve the recognition performance of trace feature extractors such as image encoders.
[0077] It should be noted that, in the scheme of this application embodiment, the classifier P ψ The aim is to distinguish edit operations within the current batch as effectively as possible, rather than performing a global classification. This approach makes the model easier to optimize because it is a... The classification of categories (the categories corresponding to a batch) (e.g., If we perform a global classification, then |C|≈10 7 category.
[0078] Based on this, optionally, the training of the image editing trace recognition model can be implemented as follows: by using a classifier, the classification results corresponding to multiple editing trace features are obtained, and the feature soft labels are determined based on the self-similarity between multiple editing operation features obtained by self-distillation of multiple editing operation features; based on the classification results and feature soft labels, the second loss of multiple editing trace features and multiple editing operation features is obtained; and the image editing trace recognition model is trained according to the first loss and the second loss.
[0079] For example, a multilayer perceptron (MLP) is denoted as P. ψ As a classifier, its trainable parameter is ψ. The training samples C from each batch of input images are used. i (X) is converted to logits Let C in the current batch i The index of (X) is used as the single label y i Therefore, classification supervision can be implemented using the following cross-entropy loss:
[0080] in, express The kth term, It is an indicator function, if y i If the value is k, then take 1; otherwise, take 0.
[0081] If y in cross-entropy i Using one-hot encoding implies equidistant orthogonality between different editing operation traces, which may reduce the efficiency of model optimization. Therefore, alternatively, embodiments of this application employ feature soft labeling to relax the optimization objective. To obtain suitable feature soft labels, in one feasible approach, a self-distillation technique can be used to compute them from the text encoder. The self-similarity of editing operation features learned in training. Self-distillation refers to a model acting as both a teacher and a student model at different stages of training or under different training settings, optimizing itself by learning from its outputs in earlier stages or under different configurations. Self-distillation avoids overfitting during model training; it eliminates the need to maintain an additional large teacher model as required by traditional knowledge distillation, reducing resource and computational costs; and it simplifies the training process and improves the model's generalization ability and performance because it eliminates the need to alternate between training teacher and student models.
[0082] Specifically, in the embodiments of this application, given text features The soft label y of the i-th editing operation i It can be represented as:
[0083] Where S(·,·) denotes the cosine similarity function:
[0084] After obtaining the feature soft labels, the cross-entropy loss can be converted into:
[0085] Accordingly, the target loss of classification supervision can be expressed as:
[0086] Therefore, the final optimization of the image editing trace recognition model is for... and The joint optimization can be expressed as:
[0087] Although the above process can effectively train the model, since the classifier's goal is only to distinguish edit trace features in the current batch, it may encounter unstable update issues when processing different batches.
[0088] Therefore, in one feasible approach, the parameters of the classifier can be initialized before each batch of training; the parameters of the classifier obtained by training the model based on multiple image training samples from the previous batch can be transferred to the initialized classifier.
[0089] In another feasible approach, adaptive constraints can be used to penalize the loss between previous batches and the current batch, ensuring that the classifier can effectively guide the classification supervision of the current batch. Based on this, as shown in Figure 2D, the portion of this embodiment that includes the adaptive constraints, MLP processing, and soft label acquisition through the self-distillation process of the text encoder is referred to as the adaptive transfer branch.
[0090] For this adaptive constraint, a K-Lipschitz discriminant function f can be learned, which expects to evaluate the previous features. Give a high score, but for the current feature A low score is given. Then, the batch loss can be measured by the 1-Wasserstein distance between the two feature distributions, i.e.
[0091] Among them, ||·|| L Let K denote the Lipschitz seminorm, where K is the K in the aforementioned K-Lipschitz. Expressing expectation. Minimizing. This can naturally reduce batch loss, thereby promoting category-level correlation between previous and current batches. In one feasible approach, this can be achieved by using ||P| in the previous batch. ψThe kernel norm of the classifier trained on (·)||* is used as the discriminant function f, where Therefore, the aforementioned batch loss can be transformed into:
[0092] Among them, W N yes abbreviation, This represents the visual and textual features corresponding to the image training sample set for the current batch of samples, namely, edit trace features and edit operation features. The visual and textual features represent the image training sample set corresponding to the previous batch of samples; sup represents the support set, which refers to the set of input values that, under given conditions, make the value of Equation 9 non-zero or meaningful, i.e., let P ψ The set of Lipschitz seminorms that satisfy the condition ≤ K. Then, KW can be estimated by maximizing the loss. N ,as follows:
[0093] Thus, the trace feature extractor E θ And classifier P ψ The joint training can be achieved through min-max optimization, that is:
[0094] The adaptive loss constraint equation Cross-entropy loss equation based on feature soft labels Combined, the overall loss used to optimize classification supervision can be expressed as:
[0095] Among them, the aforementioned The change is as follows:
[0096] Unlike the cross-entropy loss equation shown in Formula 3 above, in the equation... And in the above formula: This represents the visual and textual features of the current batch, namely, edit trace features and edit operation features. Represents the visual and textual features of previous batches; N is the number of image training samples in the current batch, and P is the classifier. ψ Will I j Convert to predicted probability This explicitly increases the discriminative distance between features of different categories. j The self-distance between text features (editing operation features) learned by the text encoder is called the feature soft label.
[0097] Based on this, in one feasible approach, training the image editing trace recognition model according to the first loss and the second loss can be achieved as follows: for multiple image training samples in each batch, first fix the parameters of the trace feature extractor, such as the image encoder, and update the parameters of the classifier according to the first loss and the second loss; after completing the update of the classifier parameters, fix the parameters of the classifier again, and update the parameters of the trace feature extractor, such as the image encoder, according to the first loss and the second loss, until the training termination condition is reached.
[0098] For example, for each batch, the classifier's parameters ψ are first initialized; then, iterations of u are performed, focusing on updating ψ; finally, a gradient backpropagation is performed to update the image encoder's parameters θ. This forms a "start-up-backpropagation" approach, solving the problem of unstable classifier updates.
[0099] The training of the image editing trace recognition model can be performed iteratively until the training termination condition is met, such as reaching a preset number of training iterations, or the loss value meeting a preset threshold, etc.
[0100] After training, during the inference phase, the image editing trace recognition model can use only the trace feature extractor, such as the image encoder, to perform image editing trace recognition. The preset editing pool, text feature extractor, and classifier are only used during the model training phase.
[0101] This embodiment achieves effective training of an image editing trace recognition model. Furthermore, by using a preset editing pool, it can generate a massive number of image editing operation types, greatly expanding the scope of a small number of original images. This generates a rich and abundant set of image training samples with diverse editing operations without requiring a large number of training samples. Positive example pairs effectively improve the model's training efficiency and convergence speed. The classifier enhances the recognition performance of the trace feature extractor, and so on. Therefore, the trained image editing trace recognition model can be widely applied to a wider range of image forensics tasks, such as social media platform source detection, synthetic image detection, image forgery detection, and camera model recognition.
[0102] The image editing trace recognition method provided in this application embodiment will be described below based on the image editing trace recognition model obtained through training.
[0103] Referring to 3A, a flowchart of steps for an image editing trace recognition method according to an embodiment of this application is shown.
[0104] The image editing trace recognition method in this embodiment includes the following steps:
[0105] Step S302: Obtain the image to be identified.
[0106] In this embodiment of the application, the image to be identified may be an image captured by any image acquisition device such as a camera or video camera, or an image after these images have undergone external editing operations, or an image generated or synthesized entirely by software, etc. This embodiment of the application does not limit the specific source and acquisition method of the image to be identified.
[0107] Step S304: Extract editing trace features from the image to be identified to obtain the corresponding editing trace features.
[0108] The editing trace features include both internal editing trace features of the image acquisition device and external editing trace features of the image acquisition device for the image to be identified.
[0109] Among them, the internal editing traces of the image acquisition device are the traces of the image after internal editing operations. These internal editing operations are used to simulate the operation of the image acquisition device to convert light signals into digital signals, including but not limited to: de-mosaic operation, white balance operation, tone mapping operation, etc.
[0110] The external editing traces of an image acquisition device are the traces of an image that have undergone external editing operations. These external editing operations are used to simulate the editing processes that an image undergoes during transmission, including but not limited to: JPEG compression, WEBP compression, scaling, blurring, and ISO noise reduction.
[0111] When extracting editing trace features based on the image editing trace recognition model obtained through the aforementioned training, it can be implemented using a trace feature extractor such as an image encoder. For example, the trace feature extractor can be used to extract editing trace features from the image to be identified. The trace feature extractor is trained on image training samples, which are obtained by editing multiple image editing operation sequences from a preset editing pool. Each of these multiple image editing operation sequences includes an internal editing operation sequence from the image acquisition device and an external editing operation sequence from the image acquisition device. The specific implementation of this process can be referred to the process of extracting editing trace features from image training samples using the trace feature extractor in the aforementioned embodiment, and will not be elaborated further here.
[0112] Step S306: Based on the editing trace features, perform editing trace recognition on the image to be identified.
[0113] Since the editing trace features can characterize which image editing operations the image to be identified has undergone, editing traces can be identified based on these features.
[0114] In one feasible approach, information about editing operations performed on the image to be identified can be obtained based on the editing trace characteristics. This information about editing operations can effectively and accurately characterize the editing traces of the image.
[0115] For example, the editing trace features of the comparison image can be obtained; based on the editing trace features of the image to be identified and the editing trace features of the comparison image, information about the editing operations performed on the image to be identified can be determined; or, based on the editing trace features of the image to be identified and the editing trace features of the comparison image, it can be determined whether the image to be identified and the comparison image have undergone the same editing operation.
[0116] In one feasible approach, determining the information about the editing operations performed on the image to be identified based on the editing trace features of the image to be identified and the editing trace features of the comparison image can be achieved by: determining, from a feature pool storing multiple editing trace standard features, standard features similar to the editing trace features of the image to be identified; and determining the information about the editing operations performed on the image to be identified based on the information of the image editing operation sequence corresponding to the similar standard features.
[0117] In one example, the image encoder in the trained image editing trace recognition model can be used for two scenarios: 1) open set verification, which aims to determine whether two images have undergone the same imaging process; and 2) closed set classification, which aims to identify the source imaging process of the image to be identified from a limited imaging process.
[0118] As shown in Figure 3B, in the open set validation scenario, given two images X1 and X2, one of which is the image to be identified (X1 in this example), firstly, edit trace features are extracted using a trace feature extractor such as an image encoder, resulting in E1 = E θ (X1) and E2 = E θ (X2). Then, the distance d(E1,E2) is used to evaluate whether the two images originated from the same imaging process, that is, whether they have undergone the same image editing operation.
[0119] For the case of closed set classification, given a set of text labels y... i The image X ∈{1,2,…,Y} i Trace feature extractors, such as image encoders, first extract their edit trace features E. i By assigning the same text label y i =k average E of the graph i Establish a feature pool Each element in the feature pool represents the average edit trace feature of a specific imaging process, also known as the edit trace standard feature. For the image X to be identified... t By identifying the feature pool containing its edit trace feature E t =E θ (X t The element with the smallest distance is used to predict its text label y. t ,Right now:
[0120] Furthermore, since text labels can represent the sequence of image editing operations, the image X to be identified can be determined. t The corresponding operation sequence is used to determine the editing operations performed on it.
[0121] In another feasible approach, after obtaining the editing trace features, the downstream tasks can be fine-tuned based on these features. The downstream tasks can be any suitable task, including but not limited to image editing trace forensics tasks. This is because, although the trained image editing trace recognition model, especially its trace feature extractor, already has the function of extracting editing trace features, the extracted editing trace features may have different downstream sub-applications (i.e., downstream tasks) in different scenarios. Unlike conventional use, which requires further fine-tuning of the model using training samples corresponding to the downstream sub-applications, in this embodiment, the images to be identified that the image editing trace recognition model actually processes can be used as training samples for fine-tuning, thereby achieving fine-tuning training of the image editing trace recognition model for downstream sub-applications.
[0122] In other words, in addition to the applications mentioned above, if strong prior knowledge is available, trace feature extractors can also be used to fine-tune them for downstream tasks, or their parameters can be used as forensic pre-training weights.
[0123] For example, as shown in Figure 3C, in a camera source recognition task, a batch of labeled data {X} is given. add ,y add}, where X add To represent specific data, y add X represents add For example, the corresponding label can be a one-hot encoded label. Therefore, the one-hot encoded label {y} can be... add} Convert to text label {T addThe parameters of the trace feature extractor are fine-tuned using language-guided contrastive supervision. Furthermore, the trace feature extractor parameters can also be used as forensic pre-training weights (similar to pre-training on ImageNet) and integrated into other frameworks for fine-tuning. This is because the trace feature extractor has learned highly discriminative prior knowledge from signal imaging and post-processing, which is beneficial for image forensics. Subsequent experiments will verify the effectiveness of using the trace feature extractor parameters as pre-training weights.
[0124] This embodiment extracts features from the image to be identified. Unlike traditional methods that extract semantic information from images, this embodiment extracts features of editing traces to obtain editing trace features. These editing trace features include both internal and external editing trace features of the image acquisition device. Internal editing trace features indicate whether the image comes from the same type of image acquisition device, while external editing trace features effectively indicate whether the image has undergone editing processing during transmission and what kind of editing processing it has undergone. In some scenarios, the editing processing of the image during transmission can indicate the image's source or even its source path. Based on this, the editing traces of the image to be identified are determined using these editing trace features.
[0125] Therefore, the solution in this embodiment can effectively identify editing traces in an image by using the editing trace features, in order to determine whether the image has been edited and what kind of editing has been done, and thus provide a basis for determining the authenticity of the image.
[0126] The following examples, using multiple application scenarios, illustrate the above process.
[0127] Scenario Example 1
[0128] This scenario example demonstrates the application of the solution in this application to detect the social media platform source of an image.
[0129] Today, social networks are ubiquitous, leading to the widespread dissemination of massive amounts of multimedia information, including digital images, on these platforms. Therefore, verifying the dissemination pathways of digital images on social media platforms has become a crucial issue in image forensics. Because different social media platforms often employ different processing mechanisms for uploaded images, they can produce varying degrees of image editing traces.
[0130] Based on this, the image editing trace recognition model in the embodiments of this application can be used to detect the social platform source of an image. One feasible approach is to use open set verification. For example, three representative datasets can be selected, such as VISION, FODB, and SDR, which contain 3, 6, and 10 social platforms, respectively. The data in VISION includes no transmission and transmissions via social platform 1 and social platform 2, while FODB, in addition to no transmission, transmissions via social platform 1 and social platform 2, also includes transmissions via social platform 3, social platform 4, and social platform 5. The SDR dataset, in addition to the above 6 transmissions, also includes transmissions from 4 other social platforms, simply represented as transmissions from social platform 6, social platform 7, social platform 8, and social platform 9. Here, social platforms 1-9 represent different social platforms.
[0131] When using the open set verification method, the editing trace features of the image to be identified extracted by the image encoder can be compared with the editing trace features of the images in the three datasets extracted by the image encoder to determine the closest feature. Then, based on the feature and its corresponding social platform transmission chain, the social platform transmission path of the image to be identified can be determined, and the social platform from which it originates can be determined.
[0132] If a closed-set classification method is used, a certain number of images, such as 125, can be randomly sampled from each dataset. A feature pool is then constructed based on these images and their corresponding text labels. For the image to be identified, its predicted label is obtained by matching its edit trace features with the features in the feature pool. Based on this predicted label, its social media platform transmission path is determined, thereby identifying the social media platform from which it originated.
[0133] Scenario Example 2
[0134] This scenario example demonstrates a scenario where the solution of this application embodiment is used for synthetic image detection. The synthetic image in this scenario is an image generated using a Generative Adversarial Network (GAN).
[0135] Because images generated by GANs often contain special artifacts, such as checkerboard texture artifacts produced by upsampling, while natural images captured by a camera do not contain such textures, the image editing trace recognition model in the embodiments of this application can be used to distinguish between images generated by GANs and natural images.
[0136] For example, five representative GANs were selected from the ForenSyth dataset: StarGAN for style transfer, CRN for photographic road synthesis, SITD for low-light enhancement, IMLE for conditional transfer model, and SAN for image super-resolution model.
[0137] Editing traces in the image to be identified can still be detected using open set verification and closed set classification methods, respectively.
[0138] When using the open set verification method, the editing trace features of the image to be identified extracted by the image encoder can be compared with the editing trace features of the images generated by the five GANs extracted by the image encoder to determine the closest feature. Then, based on the closest feature, it can be determined which GAN generated the image to be identified.
[0139] If a closed-set classification method is used, a certain number of images can be randomly sampled from each GAN-generated image, and a feature pool can be constructed based on these images and their corresponding text labels. For the image to be identified, its predicted label is obtained by matching its edit trace features with the features in the feature pool, and then the corresponding GAN is determined based on the predicted label, thereby determining which GAN generated the image to be identified.
[0140] Scenario Example 3
[0141] This scenario example illustrates the application of the solution in this application for image forgery localization.
[0142] Forged images are often created through splicing, copying, repairing, and Photoshop manipulation. Therefore, the forgery can be located by taking advantage of the inconsistencies in features between the real and tampered areas in the forged image.
[0143] For example, the given image can first be divided into blocks, and then for each block, its editing trace features can be extracted using an image editing trace recognition model. The dot product of the editing trace features of adjacent block pairs can be used to calculate the affinity matrix. Finally, based on the affinity matrix, a clustering algorithm can be used to separate real and fake regions, thereby achieving fake location.
[0144] However, this is not the only option. In another feasible approach, the trace feature extractor in the image editing trace recognition model can be used to pre-train the model. For example, the parameters of the trace feature extractor can be used as the weights of the pre-trained model to obtain a downstream task model that can better assist the downstream supervised task in achieving better performance.
[0145] Scenario Example 4
[0146] This scenario example demonstrates a case where camera model recognition is performed using the solution described in this application.
[0147] For example, four datasets can be selected: VISION, FODB, SDR, and Daxing, which contain 29, 25, 21, and 22 different camera models, respectively.
[0148] Editing traces in the image to be identified can still be detected using open set verification and closed set classification methods, respectively.
[0149] When using the open set verification method, the edit trace features of the image to be identified extracted by the image encoder can be compared with the edit trace features of the images in the above four datasets extracted by the image encoder to determine the closest feature. Then, based on the closest feature, the camera model of the image to be identified can be determined.
[0150] If a closed-set classification method is used, a certain number of images can be randomly sampled from the images in each dataset, and a feature pool can be constructed based on these images and their corresponding text labels. For the image to be identified, its predicted label is obtained by matching its edit trace features with the features in the feature pool, and then its corresponding camera model is determined based on the predicted label.
[0151] As can be seen from the above examples, the image trace recognition scheme provided in this application can be applied to various image forensics tasks to achieve effective forensics results.
[0152] Referring to Figure 4, a schematic diagram of an electronic device according to Embodiment 5 of this application is shown. The specific embodiments of this application do not limit the specific implementation of the electronic device.
[0153] As shown in Figure 4, the electronic device may include: a processor 402, a communications interface 404, a memory 406, and a communications bus 408.
[0154] in:
[0155] The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408.
[0156] Communication interface 404 is used to communicate with other electronic devices or servers.
[0157] The processor 402 is used to execute program 410, specifically to execute the relevant steps of any of the above-described method embodiments.
[0158] Specifically, program 410 may include program code that includes computer operation instructions.
[0159] Processor 402 may be a CPU, a GPU (Graphics Processing Unit), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.
[0160] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0161] Program 410 may include multiple computer instructions. Specifically, program 410 may use multiple computer instructions to cause processor 402 to perform the operation corresponding to any of the methods described in the foregoing multiple method embodiments.
[0162] The specific implementation of each step in procedure 410 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.
[0163] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in any of the foregoing method embodiments. The computer storage medium includes, but is not limited to, compact disc read-only memory (CD-ROM), random access memory (RAM), floppy disk, hard disk, or magneto-optical disk.
[0164] This application also provides a computer program product, including computer instructions that instruct a computing device to perform an operation corresponding to any of the methods in the above-described multiple method embodiments.
[0165] Furthermore, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used for training the model, image data used for identification, data used for analysis, stored data, and displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0166] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.
[0167] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., Random Access Memory (RAM), Read-Only Memory (ROM), Flash Memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0168] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for specific applications, but such implementations should not be considered beyond the scope of the embodiments of this application.
[0169] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.
Claims
1. A method for recognizing image editing traces, comprising: Acquire the image to be identified; Editing trace features are extracted from the image to be identified to obtain corresponding editing trace features, which include editing trace features inside the image acquisition device and editing trace features outside the image acquisition device. Editing traces are identified in the image based on the editing trace features.
2. The method according to claim 1, wherein, The step of identifying editing traces in the image based on the editing trace features includes: Based on the editing trace features, information on the editing operations performed on the image to be identified is obtained.
3. The method according to claim 2, wherein, The step of obtaining information on editing operations performed on the image to be identified based on the editing trace features includes: Obtain the editing trace features of the comparison images; Based on the editing trace features of the image to be identified and the editing trace features of the comparison image, information about the editing operations performed on the image to be identified is determined; or, based on the editing trace features of the image to be identified and the editing trace features of the comparison image, it is determined whether the image to be identified and the comparison image are images that have undergone the same editing operations.
4. The method according to claim 2, wherein, The step of obtaining information on editing operations performed on the image to be identified based on the editing trace features includes: Based on the edit trace features of the image to be identified, standard features similar to the edit trace features of the image to be identified are determined from a feature pool that stores multiple standard features of edit traces; Based on the information of the image editing operation sequence corresponding to the similar standard features, information on the editing operations performed on the image to be identified is determined.
5. The method according to any one of claims 1-4, wherein, The step of extracting editing trace features from the image to be identified includes: Editing trace features are extracted from the image to be identified using a trace feature extractor. The trace feature extractor is trained based on image training samples. The image training samples are obtained by editing multiple image editing operation sequences in a preset editing pool. Each of the multiple image editing operation sequences includes an internal editing operation sequence of the image acquisition device and an external editing operation sequence of the image acquisition device.
6. The method according to claim 5, wherein, The method further includes: Based on the aforementioned trace feature extractor, fine-tuning is performed for downstream tasks.
7. A training method for an image editing trace recognition model, comprising: Based on a series of image editing operation sequences in a preset editing pool, the original images in the original image set are processed by editing operations to obtain image training samples; Editing trace features are extracted from the image training samples using the trace feature extractor in the image editing trace recognition model to obtain corresponding editing trace features. Furthermore, operation features are extracted from the text labels corresponding to the image editing operation sequence using the text feature extractor in the image editing trace recognition model to obtain corresponding editing operation features. The editing trace features include editing trace features inside the image acquisition device and editing trace features outside the image acquisition device. The editing operation features include editing operation features inside the image acquisition device and editing operation features outside the image acquisition device. The loss is calculated based on the edit trace features and the edit operation features, and the image edit trace recognition model is trained based on the loss.
8. The method according to claim 7, wherein, The preset editing pool includes a set of internal editing operations of the image acquisition device and a set of external editing operations of the image acquisition device. The set of internal editing operations of the image acquisition device includes a variety of internal editing operations of the image acquisition device, and the set of external editing operations of the image acquisition device includes a variety of external editing operations of the image acquisition device. The method further includes: The internal editing operations of the various image acquisition devices are arranged to obtain a corresponding first arrangement set, which includes N1 first arrangements; and the external editing operations of the various image acquisition devices are arranged to obtain a corresponding second arrangement set, which includes N2 second arrangements. Based on the first permutation set and the second permutation set, N1*N2 image editing operation sequences are obtained, wherein each of the N1*N2 image editing operation sequences is an operation sequence in which the first permutation precedes the second permutation.
9. The method according to claim 8, wherein, The internal editing operation of the image acquisition device is used to simulate the operation of converting light signals into digital signals; the external editing operation of the image acquisition device is used to simulate the editing processing operation that the image undergoes during propagation.
10. The method according to any one of claims 7-9, wherein, The process of obtaining image training samples includes: obtaining multiple image training samples according to the number of samples indicated in a single batch processing. The step of extracting editing trace features from the image training samples using the trace feature extractor in the image editing trace recognition model to obtain corresponding editing trace features includes: extracting editing trace features from the multiple image training samples using the trace feature extractor in the image editing trace recognition model to obtain multiple corresponding editing trace features; The step of extracting operation features from the text labels corresponding to the image editing operation sequence using the text feature extractor in the image editing trace recognition model to obtain corresponding editing operation features includes: determining the text labels corresponding to the multiple image editing operation sequences corresponding to the multiple image training samples; and extracting operation features from the multiple text labels corresponding to the multiple image editing operation sequences using the text feature extractor in the image editing trace recognition model to obtain corresponding multiple editing operation features.
11. The method according to claim 10, wherein, The step of calculating a loss based on the edit trace features and the edit operation features, and training the image edit trace recognition model based on the loss, includes: Based on the multiple edit trace features and the multiple edit operation features, comparative learning is performed to obtain a first loss for the multiple edit trace features and the multiple edit operation features; The image editing trace recognition model is trained based on the first loss.
12. The method according to claim 11, wherein, The step of training the image editing trace recognition model based on the first loss includes: The classification results corresponding to the multiple editing trace features are obtained through a classifier, and the feature soft labels are determined based on the self-similarity between the multiple editing operation features obtained by self-distillation of the multiple editing operation features. Based on the classification results and the feature soft labels, a second loss is obtained for the multiple edit trace features and the multiple edit operation features; The image editing trace recognition model is trained based on the first loss and the second loss.
13. The method according to claim 12, wherein, The step of training the image editing trace recognition model based on the first loss and the second loss includes: For each batch of multiple image training samples, the parameters of the trace feature extractor are first fixed, and the parameters of the classifier are updated according to the first loss and the second loss. After updating the parameters of the classifier, the parameters of the classifier are fixed, and the parameters of the trace feature extractor are updated according to the first loss and the second loss until the training termination condition is met.
14. The method according to claim 13, wherein, Before obtaining the classification results corresponding to the multiple edit trace features through the classifier, the method further includes: Initialize the parameters of the classifier; The parameters of the classifier obtained by training the model based on multiple image training samples from the previous batch are transferred to the initialized classifier.
15. The method according to any one of claims 7-9, wherein, The method involves performing editing operations on original images in the original image set based on multiple image editing operation sequences from a preset editing pool to obtain image training samples, including: Construct the original image pairs; Using the original image pairs as units, and based on a variety of image editing operation sequences in a preset editing pool, the original image pairs are processed by editing operations to obtain the corresponding positive sample pairs; Based on the positive sample pairs, image training samples are obtained.
16. The method according to claim 15, wherein, The step of extracting editing trace features from the image training samples using the trace feature extractor in the image editing trace recognition model to obtain the corresponding editing trace features includes: The image editing trace recognition model extracts editing trace features from the positive sample to obtain the first trace feature and the second trace feature corresponding to each positive sample pair. Based on the first trace feature and the second trace feature, the edit trace feature of each positive sample pair is obtained.
17. An electronic device comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation corresponding to the method as described in any one of claims 1-16.
18. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of claims 1-16.
19. A computer program product comprising computer instructions that instruct a computing device to perform an operation corresponding to any one of the methods described in claims 1-16.
Citation Information
Patent Citations
Image authenticity identification method and application thereof in certificate identification
CN111445454A
Deep learning image tamper-proofing system and method
CN114936986A
Business license identification and verification method and device, electronic equipment and storage medium
CN115273120A
Image detection model training method and device and image detection method and device
CN116189167A
Anti-fraud method for intelligently identifying picture modification trace
CN117523214A