Image tampering detection large model training method and electronic equipment

By training the image tamper detection large models in stages, using components such as large language models and visual encoders, the problem of low accuracy of face tamper recognition in single-frame face images is solved, and more accurate tamper recognition and positioning is achieved.

CN120164087AActive Publication Date: 2025-06-17HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510646270.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The prior art is not very accurate and poorly interpretable when identifying face tampering in a single-frame face image.

Method used

A training method for image tamper detection large models is adopted, including large language models, visual encoders, segmentation decoders, visual text mappers, classification layers and positioning structures. Through phased training, the parameters of different modules are frozen, and the model parameters are fine-tuned using sample data with tampered with classification information and tampered with area data.

Benefits of technology

It realizes more accurate identification and positioning of face tampering in single-frame images, reduces dependence on sample data, and improves the interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164087A_ABST
    Figure CN120164087A_ABST
Patent Text Reader

Abstract

The invention discloses an image tampering detection large model training method and electronic equipment. An image tampering detection large model comprises a large language model, a visual encoder, a segmentation decoder, a visual text mapper, a classification layer and a positioning structure. In the pre-training stage, parameters of a large language model, a visual encoder and a visual text mapper are frozen, and parameters of a segmentation decoder, a classification layer and a positioning structure are trained through sample data with tampering classification information and tampering positioning results; in the classification training stage, parameters of a large language model and a visual encoder are frozen, sample data are converted into text information with tampering classification results and tampering positioning features, and parameters of a segmentation decoder, a visual text mapper, a classification layer and a positioning structure are trained; in the task training stage, the global structure of the image tampering detection large model is unfrozen, and all parameters of the image tampering detection large model are finely adjusted through sample data with a tampering classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular, to a method for training a large model for image tampering detection and an electronic device. Background Art

[0002] With the rapid development of computer vision and image processing technologies, people's demand for privacy and security protection of human faces is increasing. The technology for defending against human face tampering is a popular application direction in the human face security business. Although large language models have been applied in many scenarios, directly using large language models to identify human face tampering still has problems such as low accuracy and poor interpretability.

[0003] How to accurately identify the situation of human face tampering through a single-frame human face image is a technical problem that urgently needs to be solved in the human face security business. Summary of the Invention

[0004] The purpose of the present application is to provide a method for training a large model for image tampering detection and an electronic device, so as to solve the problems of low accuracy and poor interpretability of the method for identifying human face tampering in a single-frame image.

[0005] In a first aspect, the present application provides a method for training a large model for image tampering detection. The large model for image tampering detection includes: a large language model, a visual encoder, a segmentation decoder, a visual text mapper, a classification layer, and a localization structure; In the pre-training stage, freeze the parameters of the large language model, the visual encoder, and the visual text mapper, and train the parameters of the segmentation decoder, the classification layer, and the localization structure with sample data with tampering classification information and tampering localization results; In the classification training stage, freeze the parameters of the large language model and the visual encoder, convert the sample data into text information with tampering classification results and tampering localization features, and train the parameters of the segmentation decoder, the visual text mapper, the classification layer, and the localization structure; In the task training stage, unfreeze the global structure of the large model for image tampering detection, and fine-tune all parameters of the large model for image tampering detection with sample data with tampering classification results.

[0006] Optionally, converting the sample data into text information with tampering classification results and tampering localization features and training the parameters of the segmentation decoder, the visual text mapper, the classification layer, and the localization structure includes: Through the visual encoder, convert the tampered human face image in the sample data into a feature map with tampering area data, and convert the feature map into visual features through the segmentation decoder; Convert the visual features into visual text features through the visual text mapper; Encode and map the tampering classification information in the sample data into a tampering classification result through the classification layer using a prompt template; Obtain tampering localization features corresponding to the tampering localization results in the sample data through the localization structure; Convert the tampering localization features into tampering localization text through the visual text mapper; Train the parameters of the segmentation decoder, the visual text mapper, the classification layer, and the localization structure based on the visual text features, the tampering classification result, and the tampering localization text.

[0007] Optionally, before the pre-training stage, the method further includes: Obtain an initial face image, and replace the initial face image based on a face swapping algorithm to obtain a first tampered face image; Calculate the pixel difference between the key points of the first tampered face image and the key points of the initial face image, where the pixel difference includes at least one of: color histogram difference, structural difference, clarity difference, boundary difference, texture difference; For a first region in the first tampered face image with an overly large pixel difference, replace the corresponding region of the initial face image with the first region to obtain a second tampered face image and the tampering information of the second tampered face image, where the tampering information includes: tampering region data representing the tampered region of the second tampered face image and tampering attribute information representing the pixel difference; Construct a VQA visual question-answer sample with tampering information based on the second tampered face image, the tampering information, and a prompt template.

[0008] Optionally, the method further includes: When the pixel difference is a structural difference, the structural difference is obtained by comparing luminance similarity, contrast similarity, and structural similarity; When the pixel difference is a clarity difference, the clarity difference is obtained through convolution calculation based on the Laplacian operator; When the pixel difference is a boundary difference, based on at least two of the gradient change, edge transformation, and frequency domain change of the key region boundary, determine whether there is a boundary difference in the key region boundary; When the pixel difference is a texture difference, the texture difference is obtained based on the gray-level co-occurrence matrix (GLCM) contrast.

[0009] Optionally, the positioning structure includes: a linear mapping, an activation function, and a split linear mapping, and the classification layer includes: a linear mapping, a pooling layer, a classification linear mapping, and an activation function; In the pre-training stage, training the parameters of the classification layer and the positioning structure with sample data with tampering classification information and tampering positioning results includes: Through the visual encoder, converting the tampered face image in the sample data into a first feature map with tampering area data, and converting the first feature map into a tampering feature through the segmentation decoder; Converting the tampering feature into a segmentation feature through a linear mapping and an activation function, and converting the segmentation feature into a tampering positioning result through a split linear mapping and an activation function; Based on the tampering positioning result, adjusting the parameters of the positioning structure through a positioning loss function; Converting the tampering feature into a tampering classification result through a linear mapping, a pooling layer, a classification linear mapping, and an activation function, and adjusting the parameters of the classification layer based on the tampering classification result through a classification loss function.

[0010] Optionally, in the pre-training stage, freezing the language model, the visual encoder, and the visual-text mapper, and training the parameters of the classification layer and the positioning structure with sample data with tampering classification information and tampering area data includes: In the pre-training stage, using the initial face image as a positive sample, the second tampered face image as a hard example sample, and the first tampered face image as a negative sample; Training the classification layer and the positioning structure based on the positive sample, hard positive sample, and negative sample through a contrast loss function, where the contrast loss function is implemented using the following formula:

[0011] where, represents the contrast loss, P represents the positive sample, N represents the negative sample, and H represents the hard example sample.

[0012] Optionally, the visual-text mapper includes a first visual-text mapper and a second visual-text mapper. In the classification training stage: Through the visual encoder, converting the tampered face image in the sample data into a feature map with tampering area data, converting the tampering area data in the feature map into a tampering feature through the segmentation decoder, and converting the tampering feature into a segmentation feature through a linear mapping and an activation function; Converting the segmentation feature into a segmentation prompt through the first visual-text mapper; The feature map output by the visual encoder is converted into a visual prompt through the second visual text mapper; The question in the sample data is converted into a question prompt through the large language model; The tampering classification result is multiplied by the text features of the preset prompt words to obtain a classification prompt; The segmentation features, the visual prompt, the question prompt, and the classification prompt are respectively input into the large language model to inject tampering knowledge into the large language model.

[0013] Optionally, in the task training stage, the global structure of the large model for image tampering detection is unfrozen, and all parameters of the large model for image tampering detection are fine-tuned through the sample data with tampering classification results, including: For each linear mapping layer of the large language model: A global LORA structure, multiple attribute expert LORA structures, and an expert selection structure are added to each linear mapping layer; According to the face image attributes in the tampering attribute information of the sample and the text features input into the current linear mapping layer, the corresponding attribute expert LORA structure is selected from the multiple attribute expert LORA structures through the expert selection structure; Based on the text features input into the current linear mapping layer, low-rank features are generated through the global LORA structure and the selected attribute expert LORA structure, and the low-rank features and the text features processed by the current linear mapping layer are fused with probability weights to realize the inference of knowledge associated with face image attributes.

[0014] In a second aspect, an embodiment of the present application provides an image tampering detection method, where a face image to be recognized and a tampering detection instruction are input into a large model for image tampering detection, and the large model for image tampering detection is trained by the above method; The large model for image tampering detection determines whether the face image to be recognized is tampered with to obtain a tampering detection result, The tampering detection result includes: a tampering classification result, tampering localization data, and tampering attribute information. Among them, the tampering classification result includes a fake label indicating that the face image to be recognized is a tampered image, and a real label indicating that the face image to be recognized is not a tampered image; the tampering localization data represents the tampering area when the face image to be recognized is a tampered image, and the tampering attribute information represents the pixel difference in the tampering area, where the pixel difference includes at least one of color histogram difference, structural difference, clarity difference, boundary difference, and texture difference.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, including: at least one memory and at least one processor, The at least one memory stores executable code, and the at least one processor is configured to execute the executable code in the at least one memory to implement a method for training an image restoration model and / or a method for image restoration.

[0016] Optionally, the above electronic device further includes a camera and a display. The camera is configured to capture an initial face image. The display is configured to display the image tampering detection result output by the image tampering detection large model.

[0017] In the embodiments of the present application, by using sample data with tampering classification information and tampering region data to train the large language model in stages, the sample data can be converted into text with tampering localization features, which can reduce the dependence of the large language model on the sample data. By freezing different modules of the large language model in stages, the concentration of different capabilities can be achieved, so as to obtain more accurate recognition results. It does not need to rely on the correlation between the front and back frames of the video, and can identify the tampering type and locate the tampering region only through a single-frame image, thus realizing accurate face tampering localization. Description of the Drawings

[0018] Figure 1 It is a schematic flowchart of the method for constructing difficult example samples provided by the embodiments of the present application. Figure 2 It is a schematic structural diagram of an image tampering detection large model provided by the embodiments of the present application. Figure 3 It is a schematic structural diagram of an image tampering detection large model provided by the embodiments of the present application. Figure 4 It is a schematic structural diagram of a linear layer of a large model provided by the embodiments of the present application. Detailed Embodiments

[0019] The following will describe the present application in detail with reference to the specific embodiments shown in the drawings, but these embodiments do not limit the present application. Any structural, method, or functional transformation made by those of ordinary skill in the art based on these embodiments is included in the protection scope of the present application.

[0020] For tampered images, the embodiments of the present application provide a method for training an image tampering detection large model. This method can adopt the LLAMA framework, for example, Lla MA\Mobile LLAMA\PHI, etc.

[0021] The image tampering detection large model may include: a large language model, a visual encoder, a segmentation decoder, a visual text mapper, a classification layer, and a localization structure.

[0022] In the pre-training stage, freeze the parameters of the large language model, visual encoder, and visual-text mapper, and train the parameters of the segmentation decoder, classification layer, and localization structure with sample data containing tampering classification information and tampering localization results; In the classification training stage, freeze the parameters of the large language model and visual encoder, convert the sample data into text information with tampering classification results and tampering localization features, and train the parameters of the segmentation decoder, visual-text mapper, classification layer, and localization structure; In the task training stage, unfreeze the global structure of the large model for image tampering detection, and fine-tune all the parameters of the large model for image tampering detection with sample data containing tampering classification results.

[0023] The above sample data can be constructed in the following way: Obtain an initial face image, and replace the initial face image based on a face-swapping algorithm to obtain a first tampered face image; Calculate the pixel differences between the key points of the first tampered face image and the key points of the initial face image, where the pixel differences include at least one of the following: color histogram difference, structural difference, sharpness difference, boundary difference, and texture difference; For the first region in the first tampered face image with excessive pixel differences, replace the corresponding region of the initial face image with the first region to obtain a second tampered face image and the tampering information of the second tampered face image. The tampering information includes: tampering region data representing the tampered region of the second tampered face image and tampering attribute information representing the pixel differences; Based on the second tampered face image, tampering information, and a prompt template, construct a VQA visual question-answer sample with tampering information.

[0024] Exemplarily, the initial face image can be a face image captured by a camera or a face image such as a captured ID photo. The face-swapping algorithm can be deepfakes, stable_diffusion, dalle, etc., which is not limited here.

[0025] Exemplarily, hard example samples can be constructed by the method as Figure 1 shown.

[0026] Step 101: Obtain an initial face image.

[0027] Step 102: Generate a first tampered face image.

[0028] Process the initial face image with a face-swapping algorithm to obtain a first tampered face image.

[0029] Step 103: Detect key points.

[0030] Specifically, step 103 includes detecting key points of the initial face image and the first tampered face image.

[0031] Step 104: Locate the current key area.

[0032] Specifically, step 104 can determine multiple key areas of the initial face image and the tampered face image according to the key points detected in step 103. The key areas can include the eye area, the mouth area, the nose area, the cheek area, etc.

[0033] Step 105: Calculate the pixel difference.

[0034] For each key area, calculate the pixel difference between the key area of the initial face image and the key area of the first tampered face image.

[0035] For example, for the eye area, calculate the pixel difference between the eye area of the initial face image and the eye area of the first tampered face image. The pixel difference can be the pixel mean difference.

[0036] Exemplarily, the pixel difference can include at least one of: color histogram difference, structure difference, sharpness difference, boundary difference, texture difference.

[0037] Step 106: Determine whether the pixel difference of the current key area is greater than the threshold. If so, go to step 107; if not, return to step 102 to generate a new first tampered face image.

[0038] Specifically, step 106 includes determining whether the pixel difference between the current key area of the initial face image and the current key area of the first tampered face image is greater than the threshold.

[0039] Step 107: Replace the current key area to obtain the second tampered face.

[0040] Specifically, the corresponding current key area of the initial face image can be replaced with the current key area of the first tampered face image to obtain a new second tampered face image as a hard example sample.

[0041] In the embodiments of the present application, the sample information may further include face image attributes, such as light intensity, light consistency, sharpness, visibility of facial features, skin integrity, etc., which are not limited herein.

[0042] In an alternative embodiment of the present application, step 106 can be implemented by at least one of the following 5 methods.

[0043] Method 1: For the j-th key region among multiple key regions, calculate the color histogram difference between the j-th key region of the initial face image and the j-th key region of the first tampered face image through the following formula: Formula (1) Wherein, represents the pixel value in the j-th key region of the initial face image, represents the pixel value in the j-th key region of the first tampered face image.

[0044] Assume that each face image has 4 key regions, namely the mouth region, the nose region, the eye region, and the cheek region.

[0045] It is possible to determine whether the mean value of the color histogram difference of the j-th key region is greater than the threshold through the following formula : Formula (2) represents the j-th key region in the initial face image, represents the sum of the pixels in the j-th key region of the initial face image.

[0046] If in the above formula, the mean value of the color histogram difference of the j-th key region is greater than the threshold , then, use the j-th key region of the first tampered face image to replace the j-th key region in the initial face image .

[0047] Optionally, in the embodiments of the present application, when the mean value of the color histogram difference of the j-th key region is greater than the threshold , add the identifier of the j-th key region to the candidate set. After calculating the color histogram differences of all key regions, select a key region L pick from the candidate set, and use L in the first tampered face image pick to replace L in the initial face image pick .

[0048] Method 2: For the j-th key region among multiple key regions, calculate the structural difference between the j-th key region of the initial face image and the j-th key region of the first tampered face image, which can be achieved through the following method: Assume that the above structural difference is characterized by the structural similarity index SSIM, and SSIM includes luminance similarity, contrast similarity, and structural similarity.

[0049] Assume that the initial face image is represented by x and the first tampered face image is represented by y.

[0050] (1) Calculate the luminance similarity of x and y through formula (3) Formula (3), wherein, represents the average luminance of the pixels in the j-th key region of the initial face image, represents the average luminance of the j-th key region in the first tampered face image, represents a constant.

[0051] (2) Calculate the contrast similarity of x and y through formula (4) , Formula (4) wherein, characterizes the standard deviation of the j-th key region in the initial face image, represents the standard deviation of the j-th key region in the first tampered face image, represents a constant.

[0052] (3) Calculate the structural similarity of x and y through formula (5)

[0053] Formula (5) wherein, characterizes the standard deviation of the j-th key region in the initial face image, characterizes the standard deviation of the j-th key region in the first tampered face image, characterizes the covariance of images x and y, represents a constant.

[0054] (4) Calculate the SSIM structural similarity index of x and y through formula (6)

[0055] Formula (6) wherein, x represents the initial face image, y represents the first tampered face image, wherein, α, β, γ are constants, and all three can be 1, l, c, s represent the weights for balancing the three types of similarity indices, and l, c, s can also be constants. For example, the sum of l, c, and s is 1.

[0056] It can be determined that the structural difference of the j-th key region is greater than the threshold when the above is greater than the threshold, and replace the j-th key region of the initial face image with the j-th key region of the first tampered face image.

[0057] Method 3: For the j-th key region among multiple key regions, calculate the sharpness difference between the j-th key region of the initial face image and the j-th key region of the first tampered face image in the following manner: Quantify the blurriness of the j-th key region of the first tampered face image using the Laplacian operator as the above-mentioned sharpness difference.

[0058] The Laplacian operator can be represented by the following convolution kernel: This convolution kernel can be used to calculate the second-order derivative of the j-th key region, thereby detecting the edges and details of this key region.

[0059] Perform a convolution operation between the Laplacian operator and the j-th key region of the initial face image as the Laplacian response map lp1; Perform a convolution operation between the Laplacian operator and the j-th key region of the first tampered face image to obtain the Laplacian response map lp2; Calculate the variance of lp1 and lp2 as the sharpness difference between the j-th key region of the initial face image and the j-th key region of the first tampered face image.

[0060] Optionally, the embodiment of the present application can use the following formula (7) to calculate the mean value of the Laplacian response map lp1 of the j-th key region in the initial face image, and calculate the variance between this mean value and the Laplacian response map lp2 of the j-th key region of the first tampered face image as the above-mentioned sharpness difference.

[0061] Formula (7) N represents the number of pixels in the key region, represents the Laplacian response value of the i-th pixel in the j-th key region of the first tampered face image, represents the mean value of the Laplacian response map lp2 of the j-th key region of the first tampered face image.

[0062] If the variance between the Laplacian response map lp2 of the j-th key region in the first tampered face image and the mean value of the Laplacian response map lp1 of the j-th key region in the initial face image exceeds the set threshold, then execute step 107.

[0063] Method 4: For the j-th key region among multiple key regions, calculate the boundary difference between the j-th key region of the initial face image and the j-th key region of the first tampered face image in the following manner: Evaluate by analyzing three indicators of the boundary of the j-th key region: gradient change, edge transition, and frequency domain change. If at least two of these indicators exceed their respective thresholds, it is determined that the key region has boundary differences.

[0064] First, the Sobel operator can be used to calculate the gradient magnitudes of the inner boundary (the inner boundary of the first tampered image key region) and the outer boundary (the outer boundary of the initial face image key region) of the key region, and calculate the gradient change of the boundary of the j-th key region:

[0065]

[0066] Among them, characterizes the inner boundary of the key region, represents the outer boundary of the key region; The gradient means of the above outer boundary and inner boundary can be calculated by the following two formulas:

[0067]

[0068] Among them, represents the number of pixel points of the outer boundary, represents the number of pixel points of the inner boundary.

[0069] Then, according to the gradient magnitudes of the inner boundary and the outer boundary, calculate the gradient discontinuity index :

[0070] If is greater than the threshold, it is considered that there is gradient discontinuity at the boundary of the key region.

[0071] When there is gradient discontinuity at the boundary of the key region, the boundary hybrid artifacts can be calculated in the following way: Use the CANNY operator to detect the edges of the inner boundary and the outer boundary:

[0072]

[0073] is greater than the threshold, it is considered that there are edge artifacts.

[0074] The above scheme calculates the edge gradient difference by the Sobel operator and the edge density difference by the CANNY operator.

[0075] If the above Greater than the threshold, it is necessary to perform DCT transformation on the inner and outer boundaries of the key area to detect whether there is an abnormality in the frequency of the boundary of the key area:

[0076] The boundary of the key area can be divided into high-frequency and low-frequency parts, and the ratio of high-frequency to low-frequency is calculated:

[0077] If is greater than the threshold, it is considered that there is a frequency abnormality in the boundary of the key area.

[0078] Method 5: For the j-th key area among multiple key areas, calculate the texture difference between the j-th key area of the initial face image and the j-th key area of the first tampered face image in the following way: 1. Use the gray-level co-occurrence matrix (GLCM) to measure the texture clarity of the two images of the initial face image and the first tampered face image. Among them, the GLCM calculation method is as follows: First, convert the two images from color images to grayscale images respectively:

[0079] where I represents the color picture, 、 、 represent the pixel values of red, green, and blue of the color picture I respectively, represents the gray value of the grayscale image.

[0080] Map the gray value of the grayscale image to the specified number of gray levels L (for example, 256 levels). Then, the mapping formula can be seen in the following formula, rounding down:

[0081] For the distance between the color picture and the grayscale picture in each direction, construct a gray-level co-occurrence matrix. The direction can be 0°, 45°, 90°, 135°, and the distance can be 1, 2, 3 pixels. Assume that the size of the gray-level co-occurrence matrix is L×L. Then, GLCM(i,j) represents the number of times that pixels with gray value i and pixels with gray value j appear simultaneously in the specified direction and distance.

[0082] Step 201: Initialize a zero matrix of L×L: Step 202: Traverse each pixel in the image. For each pixel (x,y), calculate its neighbor pixel (x',y') in the specified direction and distance; Step 203: The gray-level co-occurrence matrix of the gray values of the pixel (x,y) and the neighbor pixel (x',y') Increment by 1; Step 204: Repeat Steps 201 and 202 until all pixels in the image are traversed.

[0083] For directions 0°, 45°, 90°, and 135°, calculate the co-occurrence probability matrix for the current direction based on the gray-level co-occurrence matrix of each pixel

[0084] Formula (9) Where, represents the co-occurrence probability matrix of the gray-level co-occurrence matrix (GLCM) of the j-th key region of the image in the current direction, i represents the pixel of the j-th key region of the image, and j represents the serial number of the current key region.

[0085] Exemplarily, the contrast of the gray-level co-occurrence matrix (GLCM) of the j-th key region of the initial face image can be calculated based on the following formula (9): Formula (9) Where, represents the contrast of the gray-level co-occurrence matrix (GLCM) of the j-th key region of the initial face image, represents the gray-level co-occurrence matrix (GLCM) of the j-th key region of the initial face image, i represents the pixel of the j-th key region of the initial face image, and j represents the serial number of the current key region.

[0086] The contrast of the gray-level co-occurrence matrix (GLCM) of the j-th key region of the first tampered face image can be calculated based on the following formula (10): Formula (10) Where, represents the contrast of the gray-level co-occurrence matrix (GLCM) of the j-th key region of the first tampered face image, represents the gray-level co-occurrence matrix (GLCM) of the j-th key region of the first tampered face image, i represents the pixel of the j-th key region of the first tampered face image, and j represents the serial number of the current key region. N represents the number of key regions in each image.

[0087] If and The difference between them is greater than the preset threshold, then it is determined that the texture difference between the j-th key region of the initial face image and the j-th key region of the first tampered face image exceeds the preset threshold. At this time, it can be determined that the texture of the j-th key region of the first tampered face is abnormal.

[0088] In the embodiments of the present application, for the first region in the first tampered face image with excessive pixel differences, the corresponding region of the initial face image can be replaced with the first region to obtain a second tampered face image and the tampering information of the second tampered face image. The tampering information may include tampered region data and tampering attribute information. The tampered region data may include the position information of the first region, and the tampering attribute information may include the attributes of the pixel differences of the first region, such as color histogram difference, structural difference, clarity difference, boundary difference, texture difference, etc.

[0089] Based on the second tampered face image, the tampering information, and the prompt template, construct a VQA visual question-answer sample with tampering information.

[0090] Exemplarily, referring to Figure 2 , the positioning structure in the embodiments of the present application may include: linear mapping, activation function, and split linear mapping, and the classification layer in the embodiments of the present application may include: linear mapping, pooling layer, classification linear mapping, and activation function; In the pre-training stage, train the parameters of the classification layer and the positioning structure with the sample data with tampering classification information and tampering positioning results, including: Through the visual encoder, convert the tampered face image in the sample data into a first feature map with tampered region data, and convert the first feature map into a tampering feature through the segmentation decoder; Convert the tampering feature into a segmentation feature through linear mapping and activation function, and convert the segmentation feature into a tampering positioning result through split linear mapping and activation function; Based on the tampering positioning result, adjust the parameters of the positioning structure through the positioning loss function; Convert the tampering feature into a tampering classification result through linear mapping, pooling layer, classification linear mapping, and activation function, and adjust the parameters of the classification layer through the classification loss function based on the tampering classification result.

[0091] The above prompt template can be optimized in the following way.

[0092] Referring to Figure 3 As shown, through the visual encoder, convert the tampered face image in the sample data into a feature map with tampered region data, convert the tampered region data in the feature map into a tampering feature through the segmentation decoder, and convert the tampering feature into a segmentation feature through linear mapping and activation function; Convert the segmentation feature into a segmentation prompt through the first visual-text mapper; Convert the feature map output by the visual encoder into a visual prompt through the second visual-text mapper; Convert the question in the sample data into a question prompt through the large language model; Multiply the tampering classification result by the text features of the preset prompt words to obtain a classification prompt; Input the segmentation feature, visual prompt, question prompt, and classification prompt into the large language model respectively to inject tampering knowledge into the large language model.

[0093] In the embodiments of the present application, in addition to the question prompt and visual prompt, the segmentation feature and classification prompt are also fed into the large language model. In this way, the large language model can better learn the tampering attributes in the image and improve the accuracy of tampering detection.

[0094] In the tampering localization branch, the tampering localization result can be an image composed of pixel values 0 and 1, and the value of each pixel represents the probability that the large model for image tampering detection predicts that the pixel belongs to the tampered face.

[0095] The loss function of the tampering localization result can be as follows, which is a per-pixel cross-entropy loss function:

[0096] where N represents the number of pixels included in the tampering localization result, represents the predicted value of each pixel belonging to the real face, represents the predicted value of each pixel belonging to the tampered face, represents the label indicating whether the pixel belongs to the tampered pixel.

[0097] Through the above loss function of the tampering localization result, the ability of the tampering localization result predicted by the large model for image tampering detection can be improved.

[0098] In the tampering classification branch, the value of the tampering classification result represents the probability that the large model for image tampering detection predicts that the image belongs to the tampered face.

[0099] The loss function of the tampering classification result can be a per-pixel cross-entropy loss function:

[0100] where N represents the number of pixels included in the tampering localization result, represents the predicted value that the face image belongs to the real face, represents the predicted value that the face image belongs to the tampered face, represents the label indicating whether the face in the image belongs to the tampered face.

[0101] Through the above loss function, the attention of the classification layer of the large model for image tampering can be more concentrated on the tampering classification result, and the ability of the tampering classification result predicted by the large model for image tampering detection can be improved.

[0102] In the pre-training stage, the initial face image can be used as a positive sample, the second tampered face image can be used as a difficult sample (difficult sample), and the first tampered face image can be used as a negative sample; by comparing the loss function, the classification layer and positioning structure of the image tampering detection model are trained based on the positive sample, the difficult sample and the negative sample. Through such a training process, the parameters of the tampering module (classification layer and positioning layer) of the image tampering detection model can be fine-tuned to improve the model's ability to locate and classify tampering.

[0103] The above contrast loss calculation is implemented in the following way:

[0104] in, represents contrast loss, P represents positive samples (real samples, i.e., initial face images), N represents the first tampered face (negative samples), and H represents the second tampered face (difficult sample). exp(⋅) represents exponential function, sim(⋅) represents similarity function, and log(⋅) represents logarithmic function.

[0105] This contrast loss function can shorten the feature distance between positive samples, shorten the feature distance between negative samples, and shorten the feature distance between positive samples and difficult samples.

[0106] In the classification training phase, the parameters of the large language model and the visual encoder are frozen. The tampered face images in the sample data can be converted into feature maps with tampered area data through the visual encoder, and the feature maps can be converted into visual features through the segmentation decoder. Converting visual features into visual-text features through a visual-text mapper; Through the classification layer, the tampering classification information in the sample data is mapped into the tampering classification result through the prompt word template encoding; Obtaining, through the positioning structure, a tampering positioning feature corresponding to the tampering positioning result in the sample data; The tampering location features are converted into tampering location texts through a visual text mapper; Based on the visual-text features, tampered classification results and tampered localization text, the parameters of the segmentation decoder, visual-text mapper, classification layer and localization structure are trained.

[0107] During the task training phase, thaw the global structure of the large model for image tampering detection, and fine-tune all parameters of the large model for image tampering detection with sample data carrying tampering classification results. Its cross-loss function can adopt the following formula to improve the matching degree between the two types of texts, real and fake, and the sample labels, and enhance the model's attention to tampered faces. The cross-entropy loss for the difference between the two types of texts, real and fake, can be obtained through the following formula:

[0108] represents the tampering label of the sample, 0 represents the tampering label, and 1 represents the real label. (a K-dimensional vector) represents the predicted position of the "real / tampered" word in the expected output, where K represents the vocabulary size in the visual text mapper.

[0109] In the embodiments of the present application, the sample information may further include face image attributes, such as light intensity, light consistency, clarity, visibility of facial features, skin integrity, etc., which are not limited herein.

[0110] During the task training phase, thaw the global structure of the large model for image tampering detection, add a global LORA structure, multiple attribute expert LORA structures, and an expert selection structure to each linear mapping layer of the large language model (where the global LORA structure and the attribute expert LORA structures may include multiple low-rank matrices), and fine-tune all parameters of the large model for image tampering detection with sample data carrying tampering classification results, which can be achieved in the following ways: Add 1 global LORA structure, multiple attribute-specific expert LORA structures, and an expert selection structure to each linear mapping layer of the large model. Add a global LORA structure, multiple attribute expert LORA structures, and an expert selection structure to each linear mapping layer; according to the face image attributes in the sample and the text features input to the current linear mapping layer, select the corresponding attribute expert LORA structure from the multiple attribute expert LORA structures through the expert selection structure; based on the text features input to the current linear mapping layer, generate low-rank features through the global LORA structure and the selected attribute expert LORA structure, and fuse the low-rank features and the text features processed by the current linear mapping layer with probability weights to realize the reasoning of knowledge associated with face image attributes.

[0111] Exemplarily, each attribute expert LORA structure includes two low-rank matrices A and B.

[0112] See Figure 4 , each expert selection structure includes a quality embedding layer, a text feature linear layer, and a selection linear layer.

[0113] For each sample data, according to the face image attributes and the input of the current linear mapping layer, a specific attribute expert LORA structure is selected through an expert selection structure. The selection of the attribute expert LORA structure can be achieved in the following ways: Suppose 9 image quality metrics (i.e., 9 face image attributes) are input into the quality embedding layer of the expert selection structure, and the input of the current linear mapping layer is fed into the text feature linear layer of the expert selection structure. The outputs of the above quality embedding layer and text feature linear layer are summed and input into the selection linear layer of the expert selection structure. The activation value obtained by passing the output of the selection linear layer through the Softmax function is used as the result probability of the selected attribute expert LORA structure, and the attribute expert LORA structure with the highest probability is the selected attribute expert LORA structure.

[0114] Based on the text features input into the current linear mapping layer, low-rank features are generated through the global LORA structure and the selected attribute expert LORA structure, and the low-rank features and the text features processed by the current linear mapping layer are fused with probability weights to realize the inference of knowledge related to face image attributes. Specifically, it can be achieved in the following ways: According to the result probability p of the selected attribute expert LORA structure, the selected attribute expert LORA structure, the global LORA structure, and the original linear mapping layer of the large language model are used to process the text features, and weighted fusion output is performed. The output formula can be as follows:

[0115] Where A n B n respectively represent the low-rank matrices of the selected attribute expert LORA structure, A g B g respectively represent the low-rank matrices of the global LORA structure, linear represents the original linear mapping layer of the large language model, x i represents the input of this linear mapping layer, and x i+1 represents the output of this linear mapping layer.

[0116] See Figure 4 As shown, through this structure, the inference of knowledge related to face image attributes can be realized, thereby improving the learning ability of the large language model for face image attributes.

[0117] In the embodiments of the present application, by using sample data with tampering classification information and tampering area data to train a large language model in stages, the dependence of the large language model on the sample data can be reduced. By freezing different modules of the large language model in stages, the concentration of attention on different capabilities can be achieved, so as to obtain more accurate recognition results. It is not necessary to rely on the correlation between the front and rear frames of the video, and the tampering type can be recognized and the tampering area can be located only through a single-frame image, thereby achieving accurate face tampering positioning.

[0118] Based on the same inventive concept, the embodiments of the present application also provide an electronic device, including: at least one memory and at least one processor. The at least one memory stores executable code, and the at least one processor is configured to execute the executable code in the at least one memory to implement the above-mentioned method for training an image restoration model, and / or the method for image restoration.

[0119] The above-mentioned electronic device further includes a camera and a display. The camera is configured to capture an initial face image. The display is configured to display the image tampering detection result output by the image tampering detection large model.

[0120] The memory can be a random access memory, a read-only memory, a non-volatile, programmable ROM, an erasable PROM, an electrically erasable, flash memory, an optical memory, and a register, etc. The processor can be a general-purpose processor. The general-purpose processor can be a processor that executes specific steps and / or operations by reading and executing the computer program stored in the memory. The general-purpose processor may use the memory stored during the execution of the steps and / or operations. The general-purpose processor can be a central processing unit, an ASIC, and an FPGA, etc. During the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instruction in the form of software. The method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor.

[0121] Exemplarily, the input device includes, but is not limited to, at least one of a keyboard, a touch panel, a voice input device, and an image sensor.

[0122] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.

[0123] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.

[0124] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The above are only the preferred embodiments of the present application and are not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A large model training method for image tampering detection, characterized in that: include: The image tampering detection large model includes: a large language model, a visual encoder, a segmentation decoder, a visual text mapper, a classification layer and a positioning structure; In the pre-training stage, the parameters of the large language model, the visual encoder, and the visual text mapper are frozen, and the parameters of the segmentation decoder, the classification layer, and the localization structure are trained by using sample data with tampered classification information and tampered localization results; In the classification training phase, the parameters of the large language model and the visual encoder are frozen, the sample data is converted into text information with tampered classification results and tampered positioning features, and the parameters of the segmentation decoder, the visual text mapper, the classification layer and the positioning structure are trained; During the task training phase, the global structure of the image tampering detection model is unfrozen, and all parameters of the image tampering detection model are fine-tuned using sample data with tampering classification results.

2. The method according to claim 1, characterized in that The converting the sample data into text information with tampering classification results and tampering positioning features, and training parameters of the segmentation decoder, the visual text mapper, the classification layer, and the positioning structure, comprises: The tampered face image in the sample data is converted into a feature map with tampered area data by the visual encoder, and the feature map is converted into visual features by the segmentation decoder; converting the visual features into visual-text features by the visual-text mapper; By means of the classification layer, the tampering classification information in the sample data is mapped into a tampering classification result by means of a prompt word template encoding; Obtaining, through the positioning structure, a tampering positioning feature corresponding to the tampering positioning result in the sample data; converting the tampering location feature into tampering location text by the visual text mapper; Based on the visual text features, the tampering classification results and the tampering localization text, the parameters of the segmentation decoder, the visual text mapper, the classification layer and the localization structure are trained.

3. The method according to claim 1, characterized in that Before the pre-training stage, the method further includes: Acquire an initial face image, and replace the initial face image based on a face-changing algorithm to obtain a first tampered face image; Calculating pixel differences between key points of the first tampered face image and key points of the initial face image, wherein the pixel differences include at least one of: color histogram difference, structure difference, clarity difference, boundary difference, and texture difference; For a first area in the first tampered face image with a large pixel difference, replacing a corresponding area of ​​the initial face image with the first area to obtain a second tampered face image and tampering information of the second tampered face image, the tampering information comprising: tampering area data representing the tampered area of ​​the second tampered face image and tampering attribute information representing the pixel difference; Based on the second tampered face image, the tampered information and the prompt word template, a VQA visual question answer sample with tampered information is constructed.

4. The method according to claim 3, characterized in that The method further comprises: When the pixel difference is a structural difference, the structural difference is obtained by comparing brightness similarity, contrast similarity and structural similarity; When the pixel difference is a definition difference, the definition difference is obtained by convolution calculation based on a Labras operator; When the pixel difference is a boundary difference, judging whether there is a mixed artifact at the boundary of the key area based on at least two indicators of gradient change, edge conversion and frequency domain change at the boundary of the key area, and if so, calculating the high-frequency and low-frequency ratios of the inner boundary and the outer boundary of the key area based on the Sobel operator to determine the boundary difference; When the pixel difference is a texture difference, the texture difference is obtained based on a grayscale fairness matrix GLCM contrast.

5. The method according to claim 1, characterized in that The positioning structure includes: a linear mapping, an activation function and a segmentation linear mapping, and the classification layer includes: a linear mapping, a pooling layer, a classification linear mapping and an activation function; In the pre-training stage, the parameters of the classification layer and the positioning structure are trained by using sample data with tampered classification information and tampered positioning results, including: The tampered face image in the sample data is converted into a first feature map with tampered area data through the visual encoder, and the first feature map is converted into tampered features through the segmentation decoder; The tampering feature is converted into a segmentation feature through linear mapping and activation function, and the segmentation feature is converted into a tampering positioning result through segmentation linear mapping and activation function; Based on the tampered positioning result, adjusting the parameters of the positioning structure through a positioning loss function; The tampering feature is converted into a tampering classification result through linear mapping, pooling layer, classification linear mapping and activation function, and the parameters of the classification layer are adjusted based on the tampering classification result through a classification loss function.

6. The method according to claim 3, characterized in that In the pre-training stage, the language model, the visual encoder, and the visual text mapper are frozen, and the parameters of the classification layer and the positioning structure are trained by sample data with tampered classification information and tampered area data, including: In the pre-training stage, the initial face image is used as a positive sample, the second tampered face image is used as a hard sample, and the first tampered face image is used as a negative sample; By contrasting the loss function, the classification layer and the positioning structure are trained based on the positive samples, the difficult positive samples and the negative samples, The contrast loss function is implemented using the following formula: in, represents contrast loss, P represents the positive sample, N represents the negative sample, and H represents the difficult sample.

7. The method according to claim 1, characterized in that The visual text mapper includes a first visual text mapper and a second visual text mapper. In the classification training stage: The tampered face image in the sample data is converted into a feature map with tampered region data through the visual encoder, the tampered region data in the feature map is converted into tampered features through a segmentation decoder, and the tampered features are converted into segmentation features through linear mapping and activation functions; converting the segmentation features into segmentation hints by a first visual-text mapper; The feature map output by the visual encoder is converted into visual cues by a second visual-text mapper; Converting the questions in the sample data into question prompts by using the large language model; Multiplying the tampering classification result and the text feature of the preset prompt word to obtain a classification prompt; The segmentation features, the visual prompts, the question prompts and the classification prompts are respectively input into the large language model to perform tampering knowledge injection on the large language model.

8. The method according to claim 1, characterized in that In the task training phase, the global structure of the image tampering detection model is unfrozen, and all parameters of the image tampering detection model are fine-tuned using sample data with tampering classification results, including: For each linear mapping layer of the large language model: Add a global LORA structure, multiple attribute expert LORA structures, and an expert selection structure to each linear mapping layer; According to the attributes of the face image in the sample and the text features input into the current linear mapping layer, a corresponding attribute expert LORA structure is selected from the multiple attribute expert LORA structures through an expert selection structure; Based on the text features input into the current linear mapping layer, low-rank features are generated through the global LORA structure and the selected attribute expert LORA structure. The low-rank features and the text features processed by the current linear mapping layer are fused in combination with probability weights to achieve reasoning of knowledge associated with facial image attributes.

9. A method for detecting image tampering, characterized in that: Inputting the face image to be identified and the tampering detection instruction into a large image tampering detection model, wherein the large image tampering detection model is trained by the method according to any one of claims 1 to 8; The image tampering detection model determines whether the face image to be identified has been tampered with, and obtains a tampering detection result. The tampering detection result includes: a tampering classification result, tampering location data and tampering attribute information, wherein the tampering classification result includes a fake label representing that the face image to be identified is a tampered image, and a real label representing that the face image to be identified is not a tampered image; the tampering location data represents a tampered area when the face image to be identified is a tampered image, and the tampering attribute information represents a pixel difference in the tampered area, wherein the pixel difference includes: at least one of color histogram difference, structure difference, clarity difference, boundary difference, and texture difference.

10. An electronic device, characterized in that: include: at least one memory and at least one processor, The at least one memory stores executable codes, and the at least one processor is configured to execute the executable codes in the at least one memory to implement the method according to any one of claims 1 to 9.

11. The electronic device according to claim 10, characterized in that: Also includes a camera and a display, The camera is used to capture an initial face image; The display is used to display the image tampering detection result output by the image tampering detection large model.

Citation Information

Patent Citations

  • Image tampering recognition model training method and device and image tampering recognition method and device

    CN111368342A

  • Face tampering video detection method and system based on double-flow contrast learning model

    CN115116108A

  • Learning unpaired multi-modal feature matching for semi-supervised learning

    CN116685989A

  • Ophthalmology large model construction method based on ophthalmology vision model and instruction set fine tuning

    CN117671422A

  • Open vocabulary object detection based on frozen vision and language model

    CN119604903A

Cited By

  • Image forgery detection method and device based on CLIP model

    CN120997613A

  • Image forgery detection method based on multi-modal large language model

    CN121236571A